Immunogenicity predictors¶
Predictors of whether a peptide elicits a T-cell response. They emit
immunogenicity, read with result.immunogenicity. Calis needs only the
peptide; the others also need an allele.
| Predictor | Class | Needs | Notes |
|---|---|---|---|
| Calis | I | peptides only | Built in; the baseline |
| PRIME | I | peptides + alleles | Calls MixMHCpred |
| DeepImmuno | I | peptides + alleles | 9- and 10-mers only |
| TLimmuno2 | II | peptides + class II alleles | Slow percentile rank |
| BigMHC_IM | I | peptides + alleles | Described under BigMHC |
Read this before trusting a score¶
Every current CD8 immunogenicity predictor, including PRIME, BigMHC IM and DeepImmuno, does well on well-characterized epitopes and poorly on novel neoepitopes. Independent benchmarks put the field at AUC 0.5–0.65 on unseen tumor neoepitopes (ITSNdb ~0.52–0.60, ICERFIRE ~0.56, IMPROVE ~0.60). Use the scores to prioritize candidates, not as ground truth.
In the one neutral head-to-head that scored both (NeoaPred, Bioinformatics 2024), BigMHC IM edged PRIME on cancer neoepitopes, while PRIME tends to do better on viral epitopes. Its training positives are mostly viral and cancer-testis antigens, with only ~129 (v1) / ~596 (v2) true immunogenic neoepitopes, and its higher self-reported numbers are partly explained by train/test overlap (IMPROVE flagged ~70% overlap with its evaluation set).
Calis¶
Calis is the classic sequence-only IEDB class-I immunogenicity model (Calis et al. 2013): a fixed per-amino-acid log-enrichment scale weighted by per-position importance, with the anchor positions (P1/P2/C-terminus) masked out.
It needs no external install and no downloaded weights. Its ~30 published
parameters (from the open-access CC-BY paper) are built in, so it is a fast,
dependency-free, allele-independent baseline. It emits one immunogenicity
prediction per peptide (empty allele); score > 0 leans immunogenic.
from mhctools import Calis
predictor = Calis()
results = predictor.predict(["GILGFVFTL", "NLVPMVATV"])
results[0].immunogenicity.score # 0.30484 (higher = more immunogenic)
PRIME¶
PRIME predicts CD8+ T-cell immunogenicity of class-I peptides by combining
MHC-I binding (via MixMHCpred, which it calls internally) with a
TCR-recognition propensity model. It emits one immunogenicity prediction per
(peptide, allele): score is the PRIME score (higher = more immunogenic) and
percentile_rank is the PRIME %Rank (lower = better).
PRIME is academic / non-commercial licensed, so mhctools shells out to an install you provide rather than vendoring it.
from mhctools import PRIME
predictor = PRIME(
alleles=["HLA-A*02:01", "HLA-B*07:02"],
program_name="PRIME", # or an absolute path
mixmhcpred_path="/path/to/MixMHCpred", # v3.0+, optional if on PATH
timeout=300)
results = predictor.predict(["GILGFVFTL", "NLVPMVATV"])
results[0].immunogenicity.score
mhctools verifies MixMHCpred's reported version before PRIME inference and rejects versions older than 3.0 or an unparseable version. The timeout covers the PRIME process tree, including its nested MixMHCpred call.
DeepImmuno¶
DeepImmuno predicts class-I CD8+ immunogenicity from the peptide and its
HLA-A/B/C allele with a small CNN (Li et al. 2021). It scores 9- and 10-mers
only and supports a fixed set of ~62 alleles, snapping anything else to the
nearest it knows. It emits one immunogenicity prediction per (peptide,
allele); score is in 0–1 (higher = more immunogenic).
DeepImmuno ships its weights in-repo and is MIT-licensed, but its script loads
them with an old Keras 2 / TensorFlow stack, so mhctools shells out to
DeepImmuno's own CLI in a separate checkout. Run mhctools fetch deepimmuno,
or point at a manual clone with DEEPIMMUNO_HOME, and set DEEPIMMUNO_PYTHON
to an interpreter that has TensorFlow with Keras 2, or newer TensorFlow plus
the tf-keras shim, since the wrapper sets TF_USE_LEGACY_KERAS=1 for the
subprocess.
from mhctools import DeepImmuno
DeepImmuno.fetch()
predictor = DeepImmuno(alleles=["HLA-A*02:01"]) # resolves DEEPIMMUNO_HOME / ~/DeepImmuno
results = predictor.predict(["NLVPMVATV", "GILGFVFTL"])
results[0].immunogenicity.score # 0.9568 (higher = more immunogenic)
TLimmuno2¶
TLimmuno2 predicts class-II (CD4+) immunogenicity. Calis, PRIME, BigMHC IM, and DeepImmuno predict class-I immunogenicity.
It scores a peptide against a class-II allele (transfer-learned from class-II
binding) and emits one immunogenicity prediction per (peptide, allele):
score in 0–1 (higher = more immunogenic) and percentile_rank from its %Rank
against a background set, rescaled to 0–100 (lower = more immunogenic).
Native NetMHCIIpan-style keys (DRB1_0803, HLA-DPA10103-DPB10101) pass
through; common DR forms (HLA-DRB1*08:03) are converted; anything TLimmuno2
does not know raises. Its upstream license is ambiguous (an Apache-2.0 README
badge, no LICENSE file), which mhctools treats the same way as NetCleave: it
can fetch a pinned snapshot, but only when you confirm your own use is
authorized, so the first fetch requires --accept-license. TLIMMUNO2_PYTHON
names an interpreter that has TensorFlow (Keras 2, or newer TensorFlow plus
tf-keras).
from mhctools import TLimmuno2
predictor = TLimmuno2(alleles=["DRB1_0803"]) # TLIMMUNO2_HOME, ~/TLimmuno2, then snapshot
results = predictor.predict(["FHTMWHVTRGAVLMY"])
results[0].immunogenicity.score # 0.9874 (higher = more immunogenic)
TLimmuno2's %Rank is computed against ~90,000 background peptides for each distinct allele, so a call costs about a minute per allele no matter how many peptides you pass. Batch peptides by allele. Class-II immunogenicity is noisier than class-I, so use the score to prioritize, not as ground truth.