mhcflurry.cli package
mhcflurry command-line modules.
The mhcflurry console script dispatches to one of the registered
subcommands, and the historical mhcflurry-* console scripts point at
the same implementations in this package. Top-level *_command.py
modules remain as import shims for compatibility.
Submodules
mhcflurry.cli.affinity_candidate_figures module
Build common-cohort affinity figure inputs for shortlisted candidates.
- mhcflurry.cli.affinity_candidate_figures.make_parser(prog='mhcflurry eval affinity-candidate-figures')[source]
Build the shortlisted-candidate figure parser.
- mhcflurry.cli.affinity_candidate_figures.build_candidate_figure_inputs(factorial_dir, out_dir, conditions, public_predictor_name, external_predictions=())[source]
Write and return common-cohort predictions and score metadata.
mhcflurry.cli.calibrate_percentile_ranks_command module
Calibrate percentile ranks for models. Runs in-place.
- mhcflurry.cli.calibrate_percentile_ranks_command.run(argv=['-b', 'html', '-v', '-d', '_build/doctrees', '-W', '.', '_build/html'])[source]
- mhcflurry.cli.calibrate_percentile_ranks_command.requested_calibration_alleles(args, predictor)[source]
Return predictor-canonical alleles selected by CLI arguments.
User-supplied names (
--allele/--alleles-file) are mapped through the affinity predictor’scanonicalize_allele_nameso a canonicalizable- but-non-canonical entry resolves to the predictor’s own pseudosequence key (which is built no-alias-first) instead of being rejected as “unsupported” by the key lookup below. Presentation predictors delegate this through their embedded affinity predictor.predictor.supported_allelesare already keys, so the default path only needs the pseudogene/null filter.
- mhcflurry.cli.calibrate_percentile_ranks_command.missing_percent_rank_alleles(predictor, alleles)[source]
Return alleles lacking direct or sequence-equivalent calibration.
- mhcflurry.cli.calibrate_percentile_ranks_command.percent_rank_status_df(predictor, alleles)[source]
Return per-allele percentile-rank calibration status.
- mhcflurry.cli.calibrate_percentile_ranks_command.run_class1_affinity_percent_rank_status(args)[source]
Print percent-rank status for requested affinity alleles and exit.
- mhcflurry.cli.calibrate_percentile_ranks_command.run_class1_processing_predictor(args)[source]
Calibrate processing against explicit background contexts, not invented flanks.
- mhcflurry.cli.calibrate_percentile_ranks_command.run_class1_presentation_predictor(args, peptides)[source]
- mhcflurry.cli.calibrate_percentile_ranks_command.do_class1_presentation_percent_rank_scores(genotypes, chunk_num=None, constant_data={})[source]
- mhcflurry.cli.calibrate_percentile_ranks_command.run_class1_affinity_predictor(args, peptides)[source]
- mhcflurry.cli.calibrate_percentile_ranks_command.do_class1_affinity_calibrate_percentile_ranks(alleles, constant_data={})[source]
- mhcflurry.cli.calibrate_percentile_ranks_command.class1_affinity_calibrate_percentile_ranks_fast(alleles, predictor, peptides, motif_summary=False, summary_top_peptide_fractions=(0.001,), verbose=False, gpu_allele_batch_size='auto', gpu_peptide_batch_size='auto', num_workers_per_gpu=1, method=None, max_knots=128)[source]
Worker-side fast-path wrapper for the GPU-batched calibration path.
Returns the same
(transforms_dict, summary_results)tuple the per-allele wrapper produces so the surrounding result-aggregation code doesn’t need to know which path ran.
mhcflurry.cli.compare_models module
Compare model ensembles on the data_evaluation benchmarks.
Combines the three legacy scripts/training/compare_*.py tools into one
command. --a and --b may each be a training-run directory, the
literal public (resolves to the currently-installed public release),
or public:<release_name> (pin a non-default release). --b defaults
to public.
Runs whichever components are available on both sides:
training_stats— per-task wall-time, epoch-count, final-loss deltas from each side’smanifest.csv. Skipped when either side is public (no manifest).affinity— per-allele ROC-AUC / PR-AUC / PPV@N on the monoallelic hit/decoy benchmark.processing— per-sample + per-length metrics on the multiallelic hit/decoy benchmark for the requested processing flank variants.presentation— per-sample + per-length micro/macro metrics on the multiallelic hit/decoy benchmark, with-flanks and without-flanks.
Writes detailed CSV/JSON artifacts plus release-summary CSV/Markdown tables.
mhcflurry plot-model-comparison consumes the CSVs to render plots.
- mhcflurry.cli.compare_models.make_parser()[source]
Return a standalone parser for documentation tooling (autoprogram).
mhcflurry.cli.downloads_command module
Download MHCflurry released datasets and trained models.
Examples
- Fetch the default downloads:
$ mhcflurry-downloads fetch
- Fetch a specific download:
$ mhcflurry-downloads fetch models_class1_pan
- Get the path to a download:
$ mhcflurry-downloads path models_class1_pan
- Get the URL of a download:
$ mhcflurry-downloads url models_class1_pan
- Summarize available and fetched downloads:
$ mhcflurry-downloads info
- mhcflurry.cli.downloads_command.run(argv=['-b', 'html', '-v', '-d', '_build/doctrees', '-W', '.', '_build/html'])[source]
- mhcflurry.cli.downloads_command.mkdir_p(path)[source]
Make directories as needed, similar to mkdir -p in a shell.
From: http://stackoverflow.com/questions/600268/mkdir-p-functionality-in-python
- mhcflurry.cli.downloads_command.suspicious_tar_member(member)[source]
Return whether a tar member should not be extracted.
- class mhcflurry.cli.downloads_command.TqdmUpTo(*_, **__)[source]
Bases:
tqdmProvides
update_to(n)which usestqdm.update(delta_n).
- mhcflurry.cli.downloads_command.list_subcommand(args)[source]
Show a readable catalogue or machine-readable bundle records.
- mhcflurry.cli.downloads_command.releases_subcommand(args)[source]
List catalogue IDs and group identical sources for an optional bundle.
- mhcflurry.cli.downloads_command.info_subcommand(args)[source]
Show resolved configuration or a bundle’s purpose, sources and usage.
mhcflurry.cli.eval_command module
Evaluation command namespace.
mhcflurry eval is the semantic home for model comparison, benchmark score
generation, and paper-style figures. Historical top-level commands remain
available as compatibility entry points:
mhcflurry eval compare-modelsdelegates tomhcflurry compare-models.mhcflurry eval plot-comparisondelegates tomhcflurry plot-model-comparison.mhcflurry eval paper-figures renderdelegates tomhcflurry paper-figures.mhcflurry eval paper-figures score-predictionsderives reusable AUC/PPV score tables from saved benchmark prediction tables.mhcflurry eval paper-figures external-predictorsoptionally shells out to external-predictor runners such asmhctoolsto add columns to a canonical saved benchmark prediction table.mhcflurry eval paper-figures runruns compare-models, paper-figures, and plot-model-comparison as one local evaluation/figure pipeline.mhcflurry eval affinity-candidate-figuresbuilds a common held-out prediction table and paper-figure suite for shortlisted affinity candidates.mhcflurry eval merge-external-predictionsconsolidates aligned per-sample external-predictor benchmark groups for those figures.
Future benchmark-prediction and external-predictor registration commands should be added under this namespace rather than as new top-level commands.
- mhcflurry.cli.eval_command.make_parser(prog='mhcflurry eval')[source]
Return the lightweight namespace parser used by docs/help tooling.
- mhcflurry.cli.eval_command.run_argv(argv, prog='mhcflurry eval')[source]
Dispatch
mhcflurry evalsubcommands.
mhcflurry.cli.experiment_snapshot module
CLI for timestamped, plot-reconstructable experiment snapshots.
mhcflurry.cli.figure_style module
Shared styling helpers for model-comparison and paper figures.
- mhcflurry.cli.figure_style.apply_paper_style()[source]
Apply the shared publication-style matplotlib defaults.
- mhcflurry.cli.figure_style.despine(ax)[source]
Hide nonessential axes decoration for paper-style panels.
mhcflurry.cli.generate_training_hyperparameters module
Generate release-training hyperparameter grids.
- mhcflurry.cli.generate_training_hyperparameters.unique_hyperparameters(items)[source]
Return
itemswith duplicate dicts removed, preserving order.
- mhcflurry.cli.generate_training_hyperparameters.build_affinity_grid(minibatch_size=128, optimizer_implementation='keras', data_dependent_initialization_target='post_activation', init='glorot_uniform')[source]
Return the 35-architecture Class I pan-allele affinity grid.
- mhcflurry.cli.generate_training_hyperparameters.processing_base_grid_iter(minibatch_size=512, optimizer_implementation='keras', init='glorot_uniform')[source]
Yield the base processing architecture grid before flank variants.
- mhcflurry.cli.generate_training_hyperparameters.build_processing_base_grid(minibatch_size=512, optimizer_implementation='keras', init='glorot_uniform')[source]
Return the processing architecture grid before flank variants.
- mhcflurry.cli.generate_training_hyperparameters.transform_processing_hyperparameters(kind, hyperparameters, optimizer_implementation=None, init=None)[source]
Return one processing flank variant for a hyperparameter dict.
optimizer_implementationandinitare optional explicit overrides. When omitted, the values already serialized inhyperparametersare retained.
- mhcflurry.cli.generate_training_hyperparameters.build_processing_variant_grid(production_hyperparameters, kind, optimizer_implementation=None, init=None)[source]
Return a flank-mode variant of a processing hyperparameter grid.
- mhcflurry.cli.generate_training_hyperparameters.build_affinity_ablation_panels()[source]
Return the paired, representative affinity audit panels.
- mhcflurry.cli.generate_training_hyperparameters.build_processing_ablation_panels()[source]
Return the paired, representative processing audit panels.
- mhcflurry.cli.generate_training_hyperparameters.build_processing_batch_sweep_panels(minibatch_sizes=(128, 256, 512, 1024))[source]
Return focused 5-aa and no-flank processing batch-sweep panels.
This is a screening grid, not the production architecture grid. It keeps the two representative architectures used by the initializer/optimizer audit and prunes recipe combinations that were dominated on the primary held-out metrics. The same candidate recipes are emitted at every batch size so minibatch size is the only between-panel change.
- mhcflurry.cli.generate_training_hyperparameters.read_hyperparameters_yaml(path)[source]
Read a YAML hyperparameter list from
path.
- mhcflurry.cli.generate_training_hyperparameters.dump_hyperparameters(hyperparameters, stream=None)[source]
Write hyperparameter dictionaries as safe YAML.
- mhcflurry.cli.generate_training_hyperparameters.add_minibatch_argument(parser, default=128)[source]
Add the common training-minibatch-size argument to
parser.
- mhcflurry.cli.generate_training_hyperparameters.add_optimizer_implementation_argument(parser)[source]
Add the optimizer-equation implementation argument.
- mhcflurry.cli.generate_training_hyperparameters.make_parser(prog=None)[source]
Build the argparse parser for the unified generator command.
- mhcflurry.cli.generate_training_hyperparameters.run_argv(argv=None, prog=None)[source]
Run the unified generator command.
- mhcflurry.cli.generate_training_hyperparameters.run_affinity_argv(argv=None, prog=None)[source]
Compatibility entry point for the old affinity generator script.
mhcflurry.cli.help module
Readable argparse help with optional terminal styling.
- mhcflurry.cli.help.color_enabled(stream)[source]
Return whether the destination supports optional terminal styling.
- mhcflurry.cli.help.style_help(text, stream)[source]
Style headings and flags only when writing to a color-capable terminal.
mhcflurry.cli.main module
Top-level mhcflurry CLI dispatcher.
Every subcommand is registered as (module path, entry attr, one-line
help) and lazy-imported on dispatch, so mhcflurry --help and
mhcflurry <subcommand> --help only pull in the modules they actually
need. This keeps the torch-import cost off the top-level help path.
Two flavors of subcommand:
New under the parent:
train,eval,compare-models,plot-model-comparison,paper-figures. Each module exposesrun_argv(argv)which does its own argparse.Historical mhcflurry-* commands:
predict,predict-scan,downloads,calibrate-percentile-ranks, theclass1-train-*/class1-select-*family,pseudosequences. Each is wrapped by invoking the command module’s existingrun(argv)(ormain(argv)forpseudosequences) on the post-subcommand argv. Allmhcflurry-*console_scripts insetup.pyremain installed as compat shims pointing at the same underlying entry functions.
- mhcflurry.cli.main.build_parser()[source]
Return the top-level parser used by
--helpand tooling.Subparsers are registered by name only — no per-subcommand arguments are added here. The legacy module owns its own
--helpoutput and is invoked lazily bymain(). Keeps this function torch-free.
mhcflurry.cli.materialize_affinity_checkpoint module
Materialize retained affinity checkpoints as an ordinary predictor.
- mhcflurry.cli.materialize_affinity_checkpoint.make_parser(prog='mhcflurry train materialize-affinity-checkpoint')[source]
mhcflurry.cli.merge_external_predictions module
Merge precomputed external-predictor benchmark groups into one table.
- mhcflurry.cli.merge_external_predictions.make_parser(prog='mhcflurry eval merge-external-predictions')[source]
Build the external-prediction merge parser.
- mhcflurry.cli.merge_external_predictions.merge_external_prediction_groups(group_specs, out_path)[source]
Merge aligned per-predictor group members and return provenance.
mhcflurry.cli.model_comparison_constants module
Shared constants for model comparison metrics and plots.
mhcflurry.cli.paper_figures module
Generate paper-style figures from retraining/evaluation outputs.
This command ports the figure families from the 2023 retraining notebooks
into a reproducible CLI. It reads saved evaluation tables: raw saved
prediction tables such as benchmark.multiallelic.csv.bz2, derived score
tables such as accuracy_scores.multiallelic.csv, and the current
compare-models output directory when supplied. Missing inputs are written
to missing_inputs.md and manifest.csv so a training run can distinguish
“not generated because the data is absent” from “plotting silently drifted.”
- class mhcflurry.cli.paper_figures.PredictorConfig(candidate: 'str', external_baselines: 'tuple', preferred_predictors: 'tuple', monoallelic_panel_predictors: 'tuple', presentation_panel_predictors: 'tuple', presentation_panel_baselines: 'tuple')[source]
Bases:
object- candidate: str
- external_baselines: tuple
- preferred_predictors: tuple
- monoallelic_panel_predictors: tuple
- presentation_panel_predictors: tuple
- presentation_panel_baselines: tuple
- class mhcflurry.cli.paper_figures.FigureInputs(scores_dir: 'Path', comparison_dir: 'Optional[Path]', run_dir: 'Optional[Path]', multiallelic_predictions: 'Optional[Path]', monoallelic_predictions: 'Optional[Path]')[source]
Bases:
object- scores_dir: Path
- comparison_dir: Path | None
- run_dir: Path | None
- multiallelic_predictions: Path | None
- monoallelic_predictions: Path | None
- mhcflurry.cli.paper_figures.make_parser()[source]
Return a standalone parser for documentation tooling.
- mhcflurry.cli.paper_figures.run_argv(argv)[source]
Entry point for the lazy
mhcflurry paper-figuresdispatcher.
- class mhcflurry.cli.paper_figures.FigureWriter(out_dir, formats, combined_pdf)[source]
Bases:
objectWrite one plot to SVG/PDF/PNG and track the output manifest.
- mhcflurry.cli.paper_figures.score_saved_prediction_table(path, index_column=None, kind=None, predictor_info=None, external_baselines=(('netmhcpan4.ba', 'ba'), ('netmhcpan4.el', 'el'), ('mixmhcpred', 'mixmhcpred')), predictor_columns=None)[source]
Return notebook-style AUC/PPV rows from a saved prediction table.
The input table must contain
hitand one grouping column (sample_id,allele, orhlaunlessindex_columnis passed). Canonical score columns are selected from the built-in predictor registry. Custom columns must be declared in apredictor_infoDataFrame withpredictorandhigher_is_bettercolumns.predictor_columnscan explicitly restrict selection to canonical or declared columns. Numeric metadata is never inferred to be a predictor score. Whenkind="monoallelic", allele identifiers are preferred oversample_idfor automatic grouping.
mhcflurry.cli.plot_loss_curves module
Plot training loss curves for every network in a run + highlight which ones were picked by the selection step.
The fit_info arrays (loss, val_loss, plus the optional
per-epoch timing breakdowns) are serialized inside each row’s config_json in
the trained models’ manifest.csv. Before selection, that manifest
has one row per candidate; after selection it contains only the models
that entered the final ensemble.
Usage — aggregate curves over the full candidate pool, colored by selection status:
- mhcflurry train plot-loss-curves
–unselected-dir results/new_run/models.unselected.combined –selected-dir results/new_run/models.combined –out results/plots/loss_curves
Usage — just the selected ensemble:
- mhcflurry train plot-loss-curves
–selected-dir results/new_run/models.combined –out results/plots/loss_curves
- Produces (in
--out): loss_curves_by_model.png— per-model train+val curves. Non- selected models in gray (if--unselected-diris provided), selected in color.loss_curves_by_arch.png— curves colored by architecture identity.per_fold_summary.png— one panel per fold showing val_loss convergence of selected vs non-selected.summary.csv— final-epoch losses and explicit checkpoint metadata.
- mhcflurry.cli.plot_loss_curves.make_parser(prog='mhcflurry train plot-loss-curves')[source]
Return the command-line parser for training-loss diagnostics.
- mhcflurry.cli.plot_loss_curves.run(args)[source]
Generate training-loss plots and a per-model summary.
mhcflurry.cli.plot_model_comparison module
Plot the metric outputs from mhcflurry compare-models.
Reads the CSVs + JSON written by compare-models and renders ROC / PR
/ scatter / per-allele delta plots under <input>/plots/ for affinity,
processing, and presentation comparisons. Kept as a separate subcommand so
the metric pipeline doesn’t pay the matplotlib import cost.
- mhcflurry.cli.plot_model_comparison.make_parser()[source]
Return a standalone parser for documentation tooling (autoprogram).
mhcflurry.cli.predict_command module
Predict binding affinity, antigen processing and presentation for peptides.
Provide a CSV or use –alleles and –peptides. Use –affinity-only for binding affinity alone. Results go to stdout unless –out is given.
- Examples:
mhcflurry predict INPUT.csv –out RESULT.csv mhcflurry predict –alleles HLA-A0201 –peptides SIINFEKL DENDREKLLL mhcflurry predict –alleles ‘HLA-A*02:01;HLA-A*03:01’ –peptides SIINFEKL
CSV columns: allele, peptide, n_flank and c_flank. Available N/C flanks are used by default; use –no-flanking to compare without sequence context. Flank columns may be omitted when context is unavailable. Input columns are preserved in the output. Separate –alleles arguments are independent queries, each scored against every peptide. Delimit alleles within a CSV cell or quoted argument with commas, semicolons or spaces to score one MHC allele set; its row reports the strongest binding allele.
mhcflurry.cli.predict_scan_command module
Scan proteins for peptide presentation using a CSV, FASTA or –sequences.
With –alleles, return peptides with affinity percentile ranks at most 2.0 by default. Use –results-all for every peptide or –threshold-* to change the filters. Without –alleles, predict processing only. Results go to stdout unless –out is given.
- Examples:
mhcflurry predict-scan proteins.fasta –alleles HLA-A0201 –out hits.csv mhcflurry predict-scan proteins.csv –alleles ‘HLA-A*02:01;HLA-A*03:01’ mhcflurry predict-scan –sequences SIINFEKLGGGNLVPMVATV –alleles HLA-A0201
CSV columns: sequence_id, sequence. Each –alleles argument is one sample; delimit alleles within a quoted argument with commas or semicolons to give a sample MHC allele set.
- class mhcflurry.cli.predict_scan_command.ChunkResult(chunk_num, predictions, comparison_quantity)
Bases:
tupleCreate new instance of ChunkResult(chunk_num, predictions, comparison_quantity)
- chunk_num
Alias for field number 0
- comparison_quantity
Alias for field number 2
- predictions
Alias for field number 1
mhcflurry.cli.processing_affinity_control module
Evaluate processing scores on affinity-controlled hit/decoy risk sets.
- mhcflurry.cli.processing_affinity_control.make_parser(prog='mhcflurry eval processing-affinity-control')[source]
Return the command-line parser.
- mhcflurry.cli.processing_affinity_control.score_risk_sets(frame, score_columns)[source]
Return overall, sample, and peptide-length metrics.
- mhcflurry.cli.processing_affinity_control.summarize_metrics(metrics, baseline)[source]
Return score summaries and paired deltas from the baseline.
mhcflurry.cli.reassign_mass_spec_training_data module
Reassign affinity values for mass-spec training rows.
- mhcflurry.cli.reassign_mass_spec_training_data.make_parser(prog=None)[source]
Build the command-line parser.
- mhcflurry.cli.reassign_mass_spec_training_data.reassign_mass_spec_training_data(data, ms_only=False, drop_negative_ms=False, set_measurement_value=None, exclude_pmhcs=None, out_csv=None, verbose=False, exclude_source_samples_manifest=None, sample_aliases=None)[source]
Return a curated training dataframe with requested MS-row edits.
mhcflurry.cli.select_allele_specific_models_command module
Model select class1 single allele models.
- mhcflurry.cli.select_allele_specific_models_command.run(argv=['-b', 'html', '-v', '-d', '_build/doctrees', '-W', '.', '_build/html'])[source]
- class mhcflurry.cli.select_allele_specific_models_command.ScrambledPredictor(predictor)[source]
Bases:
object
- class mhcflurry.cli.select_allele_specific_models_command.ScoreFunction(function, summary=None)[source]
Bases:
objectThin wrapper over a score function (Class1AffinityPredictor -> float). Used to keep a summary string associated with the function.
- class mhcflurry.cli.select_allele_specific_models_command.CombinedModelSelector(model_selectors, weights=None, min_contribution_percent=1.0)[source]
Bases:
objectModel selector that computes a weighted average over other model selectors.
- class mhcflurry.cli.select_allele_specific_models_command.ConsensusModelSelector(predictor, num_peptides_per_length=10000, multiply_score_by_value=10.0)[source]
Bases:
objectModel selector that scores sub-ensembles based on their Kendall tau consistency with the full ensemble over a set of random peptides.
- class mhcflurry.cli.select_allele_specific_models_command.MSEModelSelector(df, predictor, min_measurements=1, multiply_score_by_data_size=True)[source]
Bases:
objectModel selector that uses mean-squared error to score models. Inequalities are supported.
mhcflurry.cli.select_pan_allele_models_command module
Model select class1 pan-allele models.
APPROACH: For each training fold, we select at least min and at most max models (where min and max are set by the –{min/max}-models-per-fold argument) using a step-up (forward) selection procedure. The final ensemble is the union of all selected models across all folds.
- mhcflurry.cli.select_pan_allele_models_command.mse(predictions, actual, inequalities=None, affinities_are_already_01_transformed=False)[source]
Mean squared error of predictions vs. actual
- Parameters:
- predictionslist of float
- actuallist of float
- inequalitieslist of string (“>”, “<”, or “=”)
- affinities_are_already_01_transformedboolean
Predictions and actual are taken to be nanomolar affinities if affinities_are_already_01_transformed is False, otherwise 0-1 values.
- Returns:
- float
- mhcflurry.cli.select_pan_allele_models_command.run(argv=['-b', 'html', '-v', '-d', '_build/doctrees', '-W', '.', '_build/html'])[source]
- mhcflurry.cli.select_pan_allele_models_command.do_model_select_task(item, constant_data={})[source]
- mhcflurry.cli.select_pan_allele_models_command.model_select(fold_num, models, min_models, max_models, save_validation_predictions=False, constant_data={})[source]
Model select for a fold.
- Parameters:
- fold_numint
- modelslist of Class1NeuralNetwork
- min_modelsint
- max_modelsint
- constant_datadict
- save_validation_predictionsbool
Include per-model validation predictions in the result.
- Returns:
- dict with keys ‘fold_num’, ‘selected_indices’, ‘summary’
mhcflurry.cli.select_processing_models_command module
Model select antigen processing models.
APPROACH: For each training fold, we select at least min and at most max models (where min and max are set by the –{min/max}-models-per-fold argument) using a step-up (forward) selection procedure. The final ensemble is the union of all selected models across all folds. AUC is used as the metric.
- mhcflurry.cli.select_processing_models_command.run(argv=['-b', 'html', '-v', '-d', '_build/doctrees', '-W', '.', '_build/html'])[source]
- mhcflurry.cli.select_processing_models_command.do_model_select_task(item, constant_data={})[source]
- mhcflurry.cli.select_processing_models_command.model_select(fold_num, models, min_models, max_models, save_validation_predictions=False, constant_data={})[source]
Model select for a fold.
- Parameters:
- fold_numint
- modelslist of Class1ProcessingNeuralNetwork
- min_modelsint
- max_modelsint
- constant_datadict
- save_validation_predictionsbool
Include per-model validation predictions in the result.
- Returns:
- dict with keys ‘fold_num’, ‘selected_indices’, ‘summary’
mhcflurry.cli.train_allele_specific_models_command module
Train Class1 single allele models.
- mhcflurry.cli.train_allele_specific_models_command.run(argv=['-b', 'html', '-v', '-d', '_build/doctrees', '-W', '.', '_build/html'])[source]
mhcflurry.cli.train_command module
Release-training workflow commands.
mhcflurry train is the semantic home for maintainer-level training
workflows that compose lower-level training CLIs. The first command is a
source-checkout wrapper around scripts/release/retrain_evaluate_deploy.sh:
it keeps Brev/runplz orchestration, artifact sync, evaluation, plotting, and
optional deployment in one maintained implementation while giving users a
stable mhcflurry entry point.
mhcflurry.cli.train_pan_allele_models_command module
Train Class1 pan-allele models.
- mhcflurry.cli.train_pan_allele_models_command.assign_folds(df, num_folds, held_out_fraction, held_out_max, seed=None)[source]
Split training data into multiple test/train pairs, which we refer to as folds. Note that a given data point may be assigned to multiple test or train sets; these folds are NOT a non-overlapping partition as used in cross validation.
A fold is defined by a boolean value for each data point, indicating whether it is included in the training data for that fold. If it’s not in the training data, then it’s in the test data.
Folds are balanced in terms of allele content.
- Parameters:
- dfpandas.DataFrame
training data
- num_foldsint
- held_out_fractionfloat
Fraction of data to hold out as test data in each fold
- held_out_max
For a given allele, do not hold out more than held_out_max number of data points in any fold.
- seedint, optional
Master seed. When given, numpy’s global RNG (which the per-allele
.sample()calls below draw from) is seeded up front, so fold membership is reproducible. When None, fold assignment is left in the current NumPy RNG state.
- Returns:
- pandas.DataFrame
index is same as df.index, columns are “fold_0”, … “fold_N” giving whether the data point is in the training data for the fold
- mhcflurry.cli.train_pan_allele_models_command.pretrain_data_iterator(filename, master_allele_encoding, peptides_per_chunk=1024, shard_rank=0, num_shards=1)[source]
Step through a CSV file giving predictions for a large number of peptides (rows) and alleles (columns).
- Parameters:
- filenamestring
- master_allele_encodingAlleleEncoding
- peptides_per_chunkint
- shard_rankint
Zero-based worker shard index.
- num_shardsint
Assign each CSV chunk to one shard by its chunk index.
- Returns:
- Generator of (AlleleEncoding, EncodableSequences, float affinities) tuples
- mhcflurry.cli.train_pan_allele_models_command.pretrain_network_input_iterator(filename, master_allele_encoding, peptide_encoding, peptides_per_chunk=1024, worker_id=0, num_workers=1, compact_peptide_repeats=False, peptide_amino_acid_encoding_torch=True)[source]
Yield pretrain batches as network-input
(x_dict, y)tuples.
- mhcflurry.cli.train_pan_allele_models_command.run(argv=['-b', 'html', '-v', '-d', '_build/doctrees', '-W', '.', '_build/html'])[source]
- mhcflurry.cli.train_pan_allele_models_command.train_model(work_item_name, work_item_num, num_work_items, architecture_num, num_architectures, fold_num, num_folds, replicate_num, num_replicates, hyperparameters, pretrain_data_filename, verbose, progress_print_interval, predictor, save_to, save_all_checkpoints=False, compile_warmup_only=False, constant_data={}, resource_probe_only=False)[source]
mhcflurry.cli.train_presentation_models_command module
Train Class1 presentation models.
- mhcflurry.cli.train_presentation_models_command.run(argv=['-b', 'html', '-v', '-d', '_build/doctrees', '-W', '.', '_build/html'])[source]
- mhcflurry.cli.train_presentation_models_command.estimate_presentation_feature_worker_gb(args, predictor)[source]
Estimate steady per-worker VRAM for presentation feature generation.
Presentation workers are inference workers, not model-fit workers. Their persistent device memory is the optimized affinity predictor plus the two processing predictor ensembles. Transient activation memory is handled by each predictor’s auto batch-size resolver, but the local worker-count resolver still needs a per-worker base footprint so it does not schedule too many independent processes onto one GPU.
- mhcflurry.cli.train_presentation_models_command.iter_presentation_feature_networks(predictor)[source]
Yield torch networks that can become GPU-resident in one worker.
- mhcflurry.cli.train_presentation_models_command.network_parameter_bytes(network)[source]
Return unique parameter + buffer bytes for a torch module.
- mhcflurry.cli.train_presentation_models_command.presentation_network_peak_bytes_per_row(network)[source]
Estimate forward peak memory for a presentation feature network.
- mhcflurry.cli.train_presentation_models_command.predict_features_parallel(args, predictor, df, experiment_to_alleles)[source]
Predict BA/AP features in local worker processes.
Presentation training fits only logistic-regression weights, but feature generation is expensive: it runs the affinity predictor and both processing predictors over tens of millions of rows. Split by sample and row chunk so each worker can use its assigned GPU independently.
mhcflurry.cli.train_processing_models_command module
Train Class1 processing models.
- mhcflurry.cli.train_processing_models_command.assign_folds(df, num_folds, held_out_samples, seed=None)[source]
Split training data into multiple test/train pairs, which we refer to as folds. Note that a given data point may be assigned to multiple test or train sets; these folds are NOT a non-overlapping partition as used in cross validation.
A fold is defined by a boolean value for each data point, indicating whether it is included in the training data for that fold. If it’s not in the training data, then it’s in the test data.
- Parameters:
- dfpandas.DataFrame
training data
- num_foldsint
- held_out_samplesint
- seedint, optional
Master seed. When given, numpy’s global RNG (which the per-fold
.sample()call below draws from) is seeded up front, so held-out-sample membership is reproducible. When None, fold assignment is left in the current NumPy RNG state.
- Returns:
- pandas.DataFrame
index is same as df.index, columns are “fold_0”, … “fold_N” giving whether the data point is in the training data for the fold
- mhcflurry.cli.train_processing_models_command.run(argv=['-b', 'html', '-v', '-d', '_build/doctrees', '-W', '.', '_build/html'])[source]
- mhcflurry.cli.train_processing_models_command.processing_work_item_count(args)[source]
Best-effort number of model fits in this command.
- mhcflurry.cli.train_processing_models_command.estimate_processing_worker_gb(args)[source]
Estimate steady per-worker VRAM for processing training.
Processing is not shaped like affinity training: workers keep encoded flank tensors resident on-device, then run Conv1d models whose peak activation width scales with flank length and convolutional filters. This estimate sizes worker concurrency from the exact encoded dataset, model/gradient/optimizer state, and training activation peak for the largest architecture in the hyperparameter sweep. Validation and post-fit scoring use the remaining per-worker live-memory budget at runtime, so they do not need a second fixed launch-time batch assumption.
- mhcflurry.cli.train_processing_models_command.train_model(work_item_name, work_item_num, num_work_items, architecture_num, num_architectures, fold_num, num_folds, replicate_num, num_replicates, hyperparameters, verbose, progress_print_interval, predictor, save_to, compile_warmup_only=False, constant_data={}, resource_probe_only=False)[source]