Results and DataFrames¶
The predict() method returns a list of PeptideResult objects, one per input
peptide. Each result contains predictions across the requested alleles and
prediction kinds.
Read a peptide result¶
An accessor such as result.affinity selects the best prediction of that kind.
It returns None when the predictor does not produce the kind. For example,
an affinity-only model has no stability result.
from mhctools import MHCflurry
predictor = MHCflurry(alleles=["HLA-A*02:01", "HLA-B*07:02"])
results = predictor.predict(["SIINFEKL", "GILGFVFTL"])
r = results[0]
r.peptide # "SIINFEKL"
r.offset # position in source protein (if scanned)
r.kinds # {"pMHC_affinity", "pMHC_presentation", "antigen_processing"}
r.alleles # {"HLA-A*02:01", "HLA-B*07:02"}
# best prediction of each kind, or None when the kind is absent
r.affinity
r.presentation
r.stability
if r.affinity:
r.affinity.value # IC50 in nM
r.affinity.percentile_rank # 0-100, lower = better
r.affinity.score # predictor-specific scale, higher = better
r.affinity.allele # best allele for this kind
r.best_affinity_by_rank # by lowest percentile rank instead of score
r.preds # tuple of every underlying Prediction
r.filter(kind="pMHC_affinity")
r.filter(allele="HLA-A*02:01")
Read a prediction¶
Each peptide result contains a tuple of Prediction objects, one per allele
and kind. These immutable objects also carry the predictor and source-sequence
information:
from mhctools import Prediction
pred = Prediction(
kind="pMHC_affinity",
score=0.85, # predictor-specific scale, higher = better
peptide="SIINFEKL",
allele="HLA-A*02:01",
value=120.5, # IC50 in nM
percentile_rank=0.8,
source_sequence_name="TP53",
offset=42,
predictor_name="netMHCpan",
predictor_version="4.1",
)
The main numerical fields are:
score: always present and higher-is-better, on a predictor-specific scale.value: a physical quantity on a linear scale, in that kind's canonical unit (nM for affinity, hours for stability). Empty when the kind has no unit, or when the wrapper cannot convert to it.percentile_rank: 0-100, lower is stronger, present when the predictor scores against a background distribution.
A predictor can emit more than one kind. NetMHCpan 4.1, for example, produces
both pMHC_affinity and pMHC_presentation for every peptide-allele pair.
See prediction kinds for units and measurement context, and the predictor matrix for each model's output kinds.
Which accessor for which kind¶
Each kind has an accessor on PeptideResult that returns the best prediction of
that kind, or None:
| Kind | Accessor |
|---|---|
pMHC_affinity |
result.affinity |
pMHC_presentation |
result.presentation |
pMHC_stability |
result.stability |
antigen_processing |
result.processing |
proteasome_cleavage |
result.cleavage |
endolysosomal_cleavage |
result.endolysosomal_cleavage |
tap_transport |
result.tap_transport |
erap_trimming |
result.erap_trimming |
immunogenicity |
result.immunogenicity |
pMHC_TCR_binding |
result.tcr_binding |
peptide_half_life |
result.peptide_half_life (also serum_half_life, plasma_half_life, blood_half_life by matrix) |
The other kinds (systemic_clearance, cellular_uptake, ...) are reached with
result.filter(kind=...); they have no best-of ordering, see peptide PK,
uptake, and tissue exposure.
result.to_dict() and result.to_dataframe() serialise every underlying
prediction.
DataFrames¶
Every level has a _dataframe variant that flattens to a pandas DataFrame:
df = predictor.predict_dataframe(["SIINFEKL"], sample_name="pat001")
df = predictor.predict_proteins_dataframe({"TP53": "MEEPQ..."}, sample_name="pat001")
The columns are the same for every predictor, and are defined once as
mhctools.pred.COLUMNS:
| Column | |
|---|---|
sample_name |
who this was predicted for |
peptide, n_flank, c_flank |
the peptide and its flanking residues |
source_sequence_name, offset |
where it came from, if scanned |
predictor_name, predictor_version |
what produced it |
allele, tcr |
the MHC allele and/or TCR, empty when not applicable |
kind, score, value, percentile_rank |
the prediction itself |
measurement_context, peptide_input, cache_key |
assay context and exact-input identity |
Ordering and repeated peptides¶
Results preserve input order and repeated peptides. results[i] corresponds
to peptides[i], so you can use zip(peptides, results). A repeated peptide
gets its own PeptideResult for each occurrence. The NetMHC family and the
legacy binding-prediction fallback score each distinct peptide once and expand
its predictions back to every occurrence. If those backends omit a requested
peptide/allele pair, predict() raises ValueError instead of returning a
shorter, misaligned list.