Downloading and selecting model weights
Code and weights have separate versions. A code patch can keep the same default weights: the current package uses the 2.3.0 weight release. New weight releases identify changed model artifacts; they need not accompany every code release. Historical catalogue names remain available for reproducibility.
Browse what is available
mhcflurry downloads releases models_class1_presentation
mhcflurry downloads list --kind models
mhcflurry downloads info models_class1_presentation
releases lists valid catalogue identifiers. With a bundle name, it shows only
releases containing that bundle and groups entries with the same archive URLs.
For example, catalogues 2.2.0 and 2.0.0 point to the same presentation archive.
list groups prediction models, legacy models, experiments, and supporting data.
The first table shows the latest weights, distinct older archive versions and
installed catalogue directories, with the recommended presentation bundle first,
followed by standalone affinity and processing bundles.
Affinity and processing rows also show components found inside installed
presentation bundles as RELEASE via presentation; standalone bundle installs
remain separate entries. Components must have a model directory and
manifest.csv to appear in the table. A bundle containing only one processing
variant is marked with flanks only or without flanks only. The presentation
bundle’s ? or ! marker indicates unknown or different source URLs.
Aliases sharing the same archive are grouped in the availability columns;
releases DOWNLOAD lists every valid identifier. Historical resources appear
below the main predictors. Terminal output uses restrained color; redirected
output and NO_COLOR=1 remain plain.
Use --kind data for data only, or --release 2.2.0 to inspect an older catalogue.
info DOWNLOAD adds descriptions, archive locations, fetch/use commands and
exact embedded component paths. These presence checks do not verify weight
file integrity or prove that embedded and standalone ensembles are identical.
All three commands support --json and read the installed package’s catalogue
offline. Update the package to obtain a newer catalogue.
Bundle |
Contents and use |
|---|---|
|
Full presentation predictor, including affinity, processing, and their combiner. The normal prediction bundle. |
|
Selected pan-allele affinity ensemble, for standalone affinity prediction. |
|
Standalone processing ensembles with and without N/C flanks. |
|
Legacy allele-specific affinity models. The name does not mean the current general-purpose class I bundle. |
|
Candidate networks before ensemble selection. |
|
Historical experimental configurations. |
|
Historical training variants or smaller subsets. |
|
Supporting data and analyses. |
The presentation bundle contains its own affinity and processing components.
The standalone bundles can be absent while full presentation prediction is
ready to use. Class1AffinityPredictor.load() can fall back to the default
presentation bundle when standalone affinity weights are absent and no
affinity-path override is set. Class1ProcessingPredictor.load() has no such
fallback: pass an embedded component path explicitly, or use
Class1PresentationPredictor.load().processing_predictor_with_flanks (or
processing_predictor_without_flanks). Explicit paths and release selection
retain their existing precedence.
In JSON output, downloaded, status and path still describe the named
bundle. The additive presentation_components entries describe embedded paths,
directory/manifest presence and the presentation bundle’s source status.
Compare new and historical weights
The public 2.1.5, 2.2.0, and 2.2.1 packages used the same 2020 model archives,
registered under catalogue 2.2.0. There is no separate 2.1.5 or 2.2.1
weight catalogue. The 2.3.0 models use the 2023 curated affinity snapshot and an updated
processing training recipe; this does not mean every component gained new
2023 observations. See Release training recipe.
Fetch both presentation bundles once:
mhcflurry downloads fetch models_class1_presentation --release 2.3.0
mhcflurry downloads fetch models_class1_presentation --release 2.2.0
Use a CSV with peptide, allele, n_flank, and c_flank columns. An allele
cell can contain a semicolon-separated MHC allele set. Run both models on the
same rows; available N/C flanks are used by default:
mhcflurry predict eval.csv --model-release 2.3.0 --out new.csv
mhcflurry predict eval.csv --model-release 2.2.0 --out old.csv
For the corresponding comparison without flanks:
mhcflurry predict eval.csv --model-release 2.3.0 --no-flanking --out new-no-flanks.csv
mhcflurry predict eval.csv --model-release 2.2.0 --no-flanking --out old-no-flanks.csv
predict-scan also accepts --model-release. The selector uses a presentation
bundle, including when predict --affinity-only requests only its affinity
component. It selects installed weights; it does not download them automatically.
--models DIR instead selects an arbitrary local predictor and cannot be
combined with --model-release.
These commands compare weights using the installed code. Reproducing a historical software version also requires that version in its own environment. See Evaluation of the 2.3.0 weights for the completed comparison across public software versions, NetMHCpan BA/EL outputs, and MixMHCpred on identical rows.
Older allele-specific models
MHCflurry still distributes the allele-specific predictors described in the
2018 paper for reproducing earlier results; for new work, use the current
presentation bundle. They are a separate, affinity-only bundle selected with
--models rather than --model-release:
mhcflurry downloads fetch models_class1
mhcflurry predict \
--alleles HLA-A0201 HLA-A0301 \
--peptides SIINFEKL SIINFEKD SIINFEKQ \
--models "$(mhcflurry downloads path models_class1)/models" \
--affinity-only \
--out predictions.csv
Local storage and overrides
mhcflurry downloads info
mhcflurry downloads path models_class1_presentation --release 2.3.0
mhcflurry downloads url models_class1_presentation --release 2.3.0
info starts with model availability, followed by historical resources and
resolved configuration. Use mhcflurry downloads --verbose info for configured
default predictor paths, components inside the default presentation predictor,
and environment variables. The variables are optional
overrides: unset means
the default is in use, not that the local path is missing. On macOS the default
root is ~/Library/Application Support/mhcflurry/4/; each weight release has
its own subdirectory. The 4 is the cache-layout version, not a model version.
The presentation predictor is inside models_class1_presentation/models.
MHCFLURRY_DATA_DIR changes the parent of the release directories.
MHCFLURRY_DOWNLOADS_CURRENT_RELEASE changes the active catalogue.
MHCFLURRY_DOWNLOADS_DIR selects an unversioned custom download directory.
The MHCFLURRY_DEFAULT_CLASS1_* variables shown by info override individual
predictor defaults. Explicit --model-release selects from the requested
catalogue instead of those individual model-path overrides.
With an unversioned custom directory, --model-release requires recorded
source URLs matching the requested bundle; otherwise it asks you to use a
versioned cache or select the local model explicitly with --models. A
release selector must not silently load another release’s weights.
“Installed” means a directory exists. “Source matches” means its
DOWNLOAD_INFO.csv URLs match the catalogue; it does not verify every file’s
integrity. Public model checksums are attached to the corresponding GitHub
release. The --release option is supported consistently by fetch, list,
info, path, and url.
Percentile calibration
Released affinity and presentation predictors include percentile calibration; released processing predictors do not. For custom models, or for processing percentiles, Percentile calibration for all predictors documents compact calibration, which uses a small validated knot representation instead of dense histogram tables. Calibration changes percentile mappings, not raw model scores or weights. Use a separate model copy and an appropriate background reference; do not fit calibration to evaluation labels.