Downloading and selecting model weights

Code and weights have separate versions. A code patch can keep the same default weights: the current package uses the 2.3.0 weight release. New weight releases identify changed model artifacts; they need not accompany every code release. Historical catalogue names remain available for reproducibility.

Browse what is available

mhcflurry downloads releases models_class1_presentation
mhcflurry downloads list --kind models
mhcflurry downloads info models_class1_presentation

releases lists valid catalogue identifiers. With a bundle name, it shows only releases containing that bundle and groups entries with the same archive URLs. For example, catalogues 2.2.0 and 2.0.0 point to the same presentation archive. list groups prediction models, legacy models, experiments, and supporting data. The first table shows the latest weights, distinct older archive versions and installed catalogue directories, with the recommended presentation bundle first, followed by standalone affinity and processing bundles. Affinity and processing rows also show components found inside installed presentation bundles as RELEASE via presentation; standalone bundle installs remain separate entries. Components must have a model directory and manifest.csv to appear in the table. A bundle containing only one processing variant is marked with flanks only or without flanks only. The presentation bundle’s ? or ! marker indicates unknown or different source URLs. Aliases sharing the same archive are grouped in the availability columns; releases DOWNLOAD lists every valid identifier. Historical resources appear below the main predictors. Terminal output uses restrained color; redirected output and NO_COLOR=1 remain plain. Use --kind data for data only, or --release 2.2.0 to inspect an older catalogue. info DOWNLOAD adds descriptions, archive locations, fetch/use commands and exact embedded component paths. These presence checks do not verify weight file integrity or prove that embedded and standalone ensembles are identical. All three commands support --json and read the installed package’s catalogue offline. Update the package to obtain a newer catalogue.

Bundle

Contents and use

models_class1_presentation

Full presentation predictor, including affinity, processing, and their combiner. The normal prediction bundle.

models_class1_pan

Selected pan-allele affinity ensemble, for standalone affinity prediction.

models_class1_processing

Standalone processing ensembles with and without N/C flanks.

models_class1

Legacy allele-specific affinity models. The name does not mean the current general-purpose class I bundle.

*_unselected

Candidate networks before ensemble selection.

*_variants, *_refined, *_experiments1

Historical experimental configurations.

*_with_mass_spec, *_no_mass_spec, *_minimal

Historical training variants or smaller subsets.

data_*, allele_sequences, analysis_predictor_info, random_peptide_predictions, cross_validation_class1

Supporting data and analyses.

The presentation bundle contains its own affinity and processing components. The standalone bundles can be absent while full presentation prediction is ready to use. Class1AffinityPredictor.load() can fall back to the default presentation bundle when standalone affinity weights are absent and no affinity-path override is set. Class1ProcessingPredictor.load() has no such fallback: pass an embedded component path explicitly, or use Class1PresentationPredictor.load().processing_predictor_with_flanks (or processing_predictor_without_flanks). Explicit paths and release selection retain their existing precedence.

In JSON output, downloaded, status and path still describe the named bundle. The additive presentation_components entries describe embedded paths, directory/manifest presence and the presentation bundle’s source status.

Compare new and historical weights

The public 2.1.5, 2.2.0, and 2.2.1 packages used the same 2020 model archives, registered under catalogue 2.2.0. There is no separate 2.1.5 or 2.2.1 weight catalogue. The 2.3.0 models use the 2023 curated affinity snapshot and an updated processing training recipe; this does not mean every component gained new 2023 observations. See Release training recipe.

Fetch both presentation bundles once:

mhcflurry downloads fetch models_class1_presentation --release 2.3.0
mhcflurry downloads fetch models_class1_presentation --release 2.2.0

Use a CSV with peptide, allele, n_flank, and c_flank columns. An allele cell can contain a semicolon-separated MHC allele set. Run both models on the same rows; available N/C flanks are used by default:

mhcflurry predict eval.csv --model-release 2.3.0 --out new.csv
mhcflurry predict eval.csv --model-release 2.2.0 --out old.csv

For the corresponding comparison without flanks:

mhcflurry predict eval.csv --model-release 2.3.0 --no-flanking --out new-no-flanks.csv
mhcflurry predict eval.csv --model-release 2.2.0 --no-flanking --out old-no-flanks.csv

predict-scan also accepts --model-release. The selector uses a presentation bundle, including when predict --affinity-only requests only its affinity component. It selects installed weights; it does not download them automatically. --models DIR instead selects an arbitrary local predictor and cannot be combined with --model-release.

These commands compare weights using the installed code. Reproducing a historical software version also requires that version in its own environment. See Evaluation of the 2.3.0 weights for the completed comparison across public software versions, NetMHCpan BA/EL outputs, and MixMHCpred on identical rows.

Older allele-specific models

MHCflurry still distributes the allele-specific predictors described in the 2018 paper for reproducing earlier results; for new work, use the current presentation bundle. They are a separate, affinity-only bundle selected with --models rather than --model-release:

mhcflurry downloads fetch models_class1
mhcflurry predict \
    --alleles HLA-A0201 HLA-A0301 \
    --peptides SIINFEKL SIINFEKD SIINFEKQ \
    --models "$(mhcflurry downloads path models_class1)/models" \
    --affinity-only \
    --out predictions.csv

Local storage and overrides

mhcflurry downloads info
mhcflurry downloads path models_class1_presentation --release 2.3.0
mhcflurry downloads url models_class1_presentation --release 2.3.0

info starts with model availability, followed by historical resources and resolved configuration. Use mhcflurry downloads --verbose info for configured default predictor paths, components inside the default presentation predictor, and environment variables. The variables are optional overrides: unset means the default is in use, not that the local path is missing. On macOS the default root is ~/Library/Application Support/mhcflurry/4/; each weight release has its own subdirectory. The 4 is the cache-layout version, not a model version. The presentation predictor is inside models_class1_presentation/models.

MHCFLURRY_DATA_DIR changes the parent of the release directories. MHCFLURRY_DOWNLOADS_CURRENT_RELEASE changes the active catalogue. MHCFLURRY_DOWNLOADS_DIR selects an unversioned custom download directory. The MHCFLURRY_DEFAULT_CLASS1_* variables shown by info override individual predictor defaults. Explicit --model-release selects from the requested catalogue instead of those individual model-path overrides.

With an unversioned custom directory, --model-release requires recorded source URLs matching the requested bundle; otherwise it asks you to use a versioned cache or select the local model explicitly with --models. A release selector must not silently load another release’s weights.

“Installed” means a directory exists. “Source matches” means its DOWNLOAD_INFO.csv URLs match the catalogue; it does not verify every file’s integrity. Public model checksums are attached to the corresponding GitHub release. The --release option is supported consistently by fetch, list, info, path, and url.

Why archive tags differ

A catalogue collects many resources. Unchanged older resources keep their original URLs, so a current catalogue can include storage tags such as pre-2.0 alongside 2.3.0. Those tags describe where an archive was uploaded; they do not require prerelease code. Several catalogues can reference the same archive, and a single catalogue can include model and data archives from different dates. Inspect releases DOWNLOAD or info DOWNLOAD to see this mapping explicitly.

Percentile calibration

Released affinity and presentation predictors include percentile calibration; released processing predictors do not. For custom models, or for processing percentiles, Percentile calibration for all predictors documents compact calibration, which uses a small validated knot representation instead of dense histogram tables. Calibration changes percentile mappings, not raw model scores or weights. Use a separate model copy and an appropriate background reference; do not fit calibration to evaluation labels.