Introduction and installation
MHCflurry predicts which peptides are likely to be displayed by MHC class I molecules. Its pretrained models answer three related questions:
Binding affinity: can this peptide bind a particular MHC allele?
Antigen processing: is cellular processing likely to produce this peptide?
Presentation: is the peptide likely to reach the cell surface, considering both binding and processing?
For most epitope-prioritization work, start with the presentation predictor. Use binding affinity when you specifically need peptide–MHC binding estimates, or processing alone when you do not have an allele or MHC allele set.
The default pan-allele models support most sequenced human MHC I alleles and several other species. GPUs and Apple Silicon (MPS) are optional and are detected automatically.
Install MHCflurry
Install the latest stable MHCflurry 2.3 patch release with:
pip install --upgrade "mhcflurry>=2.3,<2.4"
Download the pretrained presentation models:
mhcflurry downloads fetch models_class1_presentation
This bundle includes the binding-affinity and antigen-processing components needed for presentation prediction. Conda users can follow Using conda instead.
Make a first prediction
Provide each peptide, its MHC alleles, and its N- and C-terminal source-protein flanks in a CSV. This small example uses windows from the sequences in example.fasta:
allele,peptide,n_flank,c_flank
HLA-A*02:01;HLA-A*03:01,TPVCPNGPG,MSSSS,NCQV
HLA-A*02:01;HLA-A*03:01,RLLEGMEMI,MVENK,FGQVI
Save it as peptides.csv, then run:
mhcflurry predict peptides.csv --out predictions.csv
If the source context is unavailable, omit the flank columns; MHCflurry uses its no-flank predictor. See Command-line tutorial for direct peptide arguments and protein scanning.
Understand the results
The output contains one row per peptide and allele or MHC allele set query. These are the main prediction columns:
Column |
Interpretation |
|---|---|
|
Combined binding and processing score from 0–1; higher is stronger. |
|
Predicted binding affinity in nM; lower is stronger. |
|
Allele-specific rank from 0–100; lower is stronger. |
|
Processing score from 0–1; higher is stronger. |
Separate allele arguments request separate predictions. A delimited allele list represents one MHC allele set and reports its strongest-binding allele. See MHC alleles and samples for examples.
Where to go next
Command-line tutorial or Python library tutorial: the full prediction workflow, including protein scanning.
Evaluation of the 2.3.0 weights: how the released models compare with other predictors.
Downloading and selecting model weights: other weight releases, if you need to reproduce an older analysis.
If you have your own measurements, Training models and Evaluating trained models cover fitting and validating custom models.
Using conda
The Bioconda package has a separate release process. Check which versions are available before installing:
conda search --override-channels -c conda-forge -c bioconda mhcflurry
For a code version not yet packaged by Bioconda, create a conda environment and install from PyPI:
conda create -q -n mhcflurry-env python=3.10
conda activate mhcflurry-env
pip install --upgrade "mhcflurry>=2.3,<2.4"
mhcflurry downloads fetch models_class1_presentation
MHCflurry supports Python 3.10+ on Linux and macOS. Windows may work but is not currently part of the supported test matrix.
Getting help and citing MHCflurry
For questions and bug reports, use the GitHub issue tracker.
If you use MHCflurry in research, cite the MHCflurry 2.0 presentation-model paper and the original binding-affinity paper listed in the project README.