TCR specificity predictors¶
Predictors of whether a T-cell receptor recognises a peptide-MHC complex. Their
input is a peptide plus a TCR (the CDR loops, optionally V/J genes), not
an allele, and they emit pMHC_TCR_binding, read with result.tcr_binding.
| Predictor | Input | Needs |
|---|---|---|
| NetTCR | (peptide, TCR) pairs |
mhctools fetch nettcr --accept-license + a TFLite runtime (pip install mhctools[nettcr]) |
| Tulip | (peptide, TCR) pairs plus the presenting allele |
mhctools fetch tulip + a TULIP-capable Python (TULIP_HOME, TULIP_PYTHON) |
| MixTCRpred | TCRs, scored against one fixed pMHC target per model | pip install "mhctools[mixtcrpred]" + mhctools fetch mixtcrpred --accept-license |
NetTCR and Tulip take explicit pairs through predict_pairs() or every
peptide-by-TCR combination through predict(peptides, tcrs). They have no
--mhc-predictor name on the command line, because the input does not fit that
interface; MixTCRpred has its own mhctools mixtcrpred subcommand.
NetTCR¶
NetTCR predicts whether a paired αβ T-cell receptor recognises a (class-I)
peptide. Unlike the MHC-ligand predictors, its input is a peptide plus a TCR
(the six CDR loops), not an allele, and it emits the pMHC_TCR_binding kind.
NetTCR ships its pretrained weights in its git repository as small TFLite models. This wrapper runs the pan cross-validation ensemble in-process and does not need NetTCR's conda environment.
from mhctools import NetTCR, TCR
NetTCR.fetch(accept_license=True) # downloads only the ~8 MB pan ensemble
predictor = NetTCR() # resolves NETTCR_DIR / ~/NetTCR-2.2
tcr = TCR(
cdr1a="NSASQS", cdr2a="VYSSG", cdr3a="VVEGDKVI",
cdr1b="MGHRA", cdr2b="YSYEKL", cdr3b="ASSHSGYEQF", name="clone1")
# Score explicit (peptide, TCR) pairs...
results = predictor.predict_pairs([("LLWNGPMAV", tcr)])
results[0].tcr_binding.score # ensemble-mean recognition probability
# ...or every peptide x TCR combination.
results = predictor.predict(["LLWNGPMAV", "GILGFVFTL"], [tcr])
Tulip¶
from mhctools import Tulip, TCR
tcr = TCR(cdr3a="CAGASGNTGKLIF", cdr3b="CASSIRASYEQYF", name="clone1")
Tulip.fetch() # pinned code, tokenizers, and weights
predictor = Tulip() # also needs a TULIP-capable Python
results = predictor.predict(["GILGFVFTL"], [tcr], mhc="HLA-A*02:01")
results[0].preds[0].score # higher = more likely binding
TULIP-TCR is GPLv3 and pinned
to transformers==4.32.1; mhctools is Apache-2.0 and depends on neither torch
nor transformers. The Tulip wrapper therefore vendors none of TULIP. It runs
an upstream checkout out-of-process, in an isolated interpreter, via TULIP's own
predict.py. mhctools fetch tulip obtains the tested code, tokenizers, and
weights; scripts/setup_tulip_env.sh can build the separate runtime.
You may instead provide your own checkout and interpreter:
TULIP_HOME: a clone of TULIP-TCR (providespredict.py,src/, tokenizers, and the releasedmodel_weights/);TULIP_PYTHON: an isolated Python 3.11 interpreter withtorchandtransformers==4.32.1. Python 3.11 specifically, sotokenizersinstalls from a prebuilt wheel and needs no Rust toolchain.
MixTCRpred¶
MixTCRpred has a different, deliberately explicit shape: each checkpoint is trained for one fixed peptide/MHC target. Its catalog currently contains 146 models (43 marked high-confidence by upstream), spanning human/mouse class I and II targets.
The input TCR stores paired CDR3s and optional trav, traj, trbv, and
trbj assignments. Released models use CDR3alpha/beta plus CDR1/2 derived from
the V genes; J assignments are accepted and QC-reported but are not network
inputs.
Each checkpoint is about 31.8 MB. The immutable Zenodo record contains 146 checkpoints totaling 4.64 GB (4.33 GiB); the 43 high-confidence checkpoints total 1.37 GB (1.27 GiB). To avoid turning every mhctools installation into a multi-gigabyte download, the pinned upstream artifact includes its two bundled reference checkpoints, and additional models are fetched individually with bounded retries, atomic installation, and checksum verification.
pip install "mhctools[mixtcrpred]"
mhctools fetch mixtcrpred --accept-license
mhctools ls mixtcrpred --models --high-confidence
mhctools fetch mixtcrpred --model A0201_GILGFVFTL
mhctools mixtcrpred --model A0201_GILGFVFTL \
--input paired-tcrs.csv --out scored-tcrs.csv
from mhctools import MixTCRpred, TCR
models = MixTCRpred.catalog()
target = MixTCRpred.resolve_model("GILGFVFTL", "HLA-A*02:01")
predictor = MixTCRpred(target.name)
tcr = TCR(
cdr3a="CAGASGNTGKLIF", cdr3b="CASSIRASYEQYF",
trav="TRAV27", traj="TRAJ42", trbv="TRBV19", trbj="TRBJ2-6",
)
prediction = predictor.predict_tcrs([tcr])[0]
prediction.score
prediction.percentile_rank
annotate_dataframe() and the CSV command retain the original table and add
the raw score, percentile rank, fixed target metadata, corrected V/J names,
V-derived CDR1/2 loops, and upstream-equivalent QC warning.
MixTCRpred code is academic/non-commercial and fetched directly from its original repository only after explicit acceptance. Optional checkpoints come from the authors' immutable CC-BY-4.0 Zenodo record and are checksum-verified before use.
As with any PyTorch checkpoint, explicit overrides and user-managed model files must come from a trusted source, because loading can execute serialized code.