Training models
Train a model when you have new measurements, need a controlled experiment, or are preparing a release. Always evaluate a trained ensemble on held-out data before using it for scientific conclusions.
Choose a workflow
Data and goal |
Command |
|---|---|
One or a few well-covered alleles |
|
Measurements spanning many alleles |
|
Full retrain, selection, calibration, and evaluation |
|
The low-level commands fit candidate models. Production ensembles also require model selection and percentile calibration; use the release workflow when you need the complete pipeline.
The release workflow and the mhcflurry train data-preparation/sweep wrappers
require a source checkout containing scripts/. Run them from the checkout root
or use an editable install of that checkout. The low-level mhcflurry class1-train-* commands are included in the installed Python package.
Small allele-specific example
Fetch the curated example data:
mhcflurry downloads fetch data_curated
The training command accepts a YAML list of architectures. This single, four-member architecture is suitable for learning the workflow; it is not a replacement for the released ensemble search:
- activation: tanh
dropout_probability: 0.0
early_stopping: true
layer_sizes: [8]
locally_connected_layers: []
loss: custom:mse_with_inequalities
max_epochs: 500
minibatch_size: 16384
n_models: 4
output_activation: sigmoid
patience: 20
peptide_amino_acid_encoding: BLOSUM62
random_negative_affinity_max: 50000.0
random_negative_affinity_min: 20000.0
random_negative_constant: 25
random_negative_rate: 0.0
validation_split: 0.1
Save it as hyperparameters.yaml, then run:
mhcflurry class1-train-allele-specific-models \
--data "$(mhcflurry downloads path data_curated)/curated_training_data.csv.bz2" \
--hyperparameters hyperparameters.yaml \
--allele 'HLA-A*02:01' \
--min-measurements-per-allele 75 \
--out-models-dir models
The output directory is a complete predictor. Keep its manifest, metadata, and
weights together; pass the directory to prediction commands with --models.
Prepare training data
Affinity training tables use one row per measurement. Allele-specific training
requires allele, peptide, measurement_value, and measurement_type.
Use measurement_type to identify quantitative or qualitative measurements.
The optional measurement_inequality column records =, <, or >; if the
column is omitted, all measurements are treated as equalities.
Alleles must be sequence-resolved MHC class I names. Allele-specific training
does not require a pseudosequence table unless the hyperparameters enable
cross-allele pretraining with pretrain_min_points.
The curated data used in the example and for released models is the
data_curated download bundle. Use mhcflurry downloads path data_curated to
locate it.
Pan-allele training
Pan-allele training additionally needs an allele pseudosequence table and a model-selection step. Each training allele must resolve to a key in that table; rows for alleles without a matching pseudosequence are excluded. Start with the command help and the maintained training recipes rather than copying individual flags from an old run:
mhcflurry class1-train-pan-allele-models --help
mhcflurry train pan-allele-release --help
Release workflow
The release workflow can run locally or on a configured remote backend. It records source and workflow provenance, resumes completed phases, compares the candidate with public models, and copies review artifacts back to the control machine. Deployment is opt-in.
Leave worker counts, GPU packing, DataLoader workers, and prediction batches on
auto initially. See Configuration and performance before pinning resource values and
the maintained training scripts
for the current release recipe.
Evaluate before use
Run mhcflurry eval compare-models against a public baseline and inspect the
summary metrics and plots. Prediction-affecting changes need held-out evidence,
not only a successful unit-test run. See Evaluating trained models for the evaluation
workflow.
Refit calibration after changing an ensemble’s members or weights. The release workflow handles affinity and presentation; standalone processing calibration is explicit. See Percentile calibration for all predictors for reference requirements and preserving previous calibrations for comparisons.
Antigen-processing training data
mhcflurry train processing-data needs a source checkout, like the release
workflow.
New processing training and model selection require affinity/length-matched hit/decoy risk sets. Generate them with the maintained command:
mhcflurry train processing-data \
--hits hits_with_tpm.csv.bz2 \
--affinity-predictor PUBLIC_AFFINITY/models.combined \
--proteome-reference-csv uniprot_proteins.csv.bz2 \
--exclude-samples-file RELEASE_HOLDOUT/processing_samples.csv \
--negative-policy matched --decoys-per-hit 1 \
--max-affinity-distance 0.25 --ppv-multiplier 100 \
--random-seed 42 --out processing_train.csv
mhcflurry train validate-processing-data --data processing_train.csv
The frozen reference is independent of the newly trained affinity candidate.
Each negative has the same sample and peptide length as its hit and differs
by at most 0.25 log10 predicted-affinity units. Same-protein matches are
preferred. If the scored candidate pool is insufficient, generation fails;
increase the pool multiplier in a fresh experiment instead of dropping hits
or using unmatched fallback peptides. Preserve processing_train.csv.matching/
alongside the final table: it contains reference/input fingerprints, the scored
candidate pools, matching diagnostics, and any failure records. The table
itself retains risk-set/source identities and matching metadata.
Training and selection default to --processing-data-policy matched, including
resume validation. Inner early stopping keeps whole samples together so reused
decoys do not leak across its split. For historical exact-data replay only,
pass --processing-data-policy legacy; reproducing the old generator also
requires --negative-policy legacy-top-binders. The exact-public-processing
workflow sets the legacy policy explicitly. Do not call a matched-data retrain
an exact-public-data comparison.
Processing evaluation defaults to the same matching algorithm with ten decoys
per hit. It saves both the full benchmark prediction table for external joins
and the matched risk-set predictions used for metrics. Training uses one decoy
per hit by default to remain balanced; these two prevalences do not give
directly comparable AUPRC values. Historical random-decoy evaluation requires
--processing-negative-policy random-diagnostic, and its plotting requires
--allow-legacy-processing-plots. Affinity regression and end-to-end
presentation retain their separate objectives/cohorts; matched processing
results are not plotted on their absolute-performance axes.