API reference¶
Use this reference to look up constructors, fields, and methods. For a complete first prediction, start with getting started.
| Task | Reference | Explanation and examples |
|---|---|---|
| Read predictions and scores | Results | Results and DataFrames |
| Describe a peptide or TCR | Inputs | Input shapes |
| Add predictions to a table | Table annotation | Annotation recipe |
| Download or inspect models | Models and artifacts | Getting models |
| Assess individual peptide bonds | Peptidase activity | Activity guide |
Predictor-specific setup and examples are in the family guides. The predictor matrix lists Python classes and command-line names for every supported model.
Results¶
predict() returns one PeptideResult per peptide.
Its predictions are Prediction objects, each describing
one endpoint and MHC or TCR context. Kind names the endpoint;
MeasurementContext records its measurement
semantics. MultiSample runs a predictor for several
named MHC allele sets.
Read results and DataFrames for examples and prediction kinds for score meanings and units. Source: prediction types and sample helper.
PeptideResult
dataclass
¶
All predictions for one peptide (a tuple of Prediction objects).
endolysosomal_cleavage
property
¶
endolysosomal_cleavage
Best endolysosomal (MHC-II) cleavage prediction, or None.
peptide_half_life
property
¶
peptide_half_life
Peptide half-life when all available results share one context.
Comparing different matrices or systemic scopes is intentionally rejected; callers can select those contexts explicitly instead.
best_by ¶
best_by(kind, field)
Return the prediction of kind with the best field, or None.
"Best" direction comes from :func:best_direction —
score is max-better, percentile_rank is min-better, and
value is kind-dependent (IC50 lower-better, half-life higher-better).
Predictions with None for field are skipped. Predictions
with an allele are preferred; if none of the matching kind have an
allele, falls back to allele-less predictions (e.g. processing
predictors that emit allele-independent scores).
best_by_value ¶
best_by_value(kind)
Best prediction of kind by value. Direction is kind-specific
(see :data:VALUE_BEST_DIRECTIONS). Raises ValueError for kinds
without a registered value direction.
Prediction
dataclass
¶
Kind ¶
String constants for prediction kinds.
You can use Kind.pMHC_affinity or just "pMHC_affinity" —
they're the same string. These constants name what is measured, but
predictor instances define the MHC context required for their supported
kinds through kind_support().
MeasurementContext
dataclass
¶
Versioned semantics shared by every prediction.
unit and transform describe :attr:Prediction.value; the stored
physical value remains linear, so transform must currently be
"linear" when a value is present. Predictor-native or confidence
outputs belong in score and are identified by score_semantics.
Ordinary model outputs receive a small cached default; assay-specific
wrappers fill only the fields they actually know. Unknown descriptive
fields are None, never a plausible biological default.
MultiSample ¶
Run a predictor across multiple samples, each with its own alleles.
Parameters:
-
samples(dict) –Mapping of sample_name -> list of allele strings. Example: {"pat001": ["HLA-A02:01", "HLA-B07:02"], "pat002": ["HLA-A01:01", "HLA-B08:01"]}
-
predictor_class(class) –A predictor class (e.g. NetMHCpan41) that accepts an
alleleskeyword argument. -
**predictor_kwargs–Additional keyword arguments forwarded to the predictor constructor.
predict_dataframe ¶
predict_dataframe(peptides)
predict() flattened to a DataFrame with sample_name column.
predict_proteins ¶
predict_proteins(sequence_dict, peptide_lengths=None)
Returns:
-
dict mapping sample_name -> {sequence_name: list of PeptideResult}–
predict_proteins_dataframe ¶
predict_proteins_dataframe(sequence_dict, peptide_lengths=None)
predict_proteins() flattened to a DataFrame with sample_name column.
Inputs¶
TCR carries receptor sequences and gene identifiers.
PeptideInput records a peptide's sequence, terminal
chemistry, attachments, and source occurrence.
PeptideContext adds study, formulation, and assay
context when those details are known.
See input shapes, the TCR guide, and peptide exposure inputs. Source: TCR types and peptide types.
TCR
dataclass
¶
A paired αβ T-cell receptor, described by its six CDR loops.
Parameters:
-
cdr1a(str, default:'') –CDR1/2/3 amino-acid sequences of the α chain (NetTCR
A1/A2/A3). -
cdr2a(str, default:'') –CDR1/2/3 amino-acid sequences of the α chain (NetTCR
A1/A2/A3). -
cdr3a(str, default:'') –CDR1/2/3 amino-acid sequences of the α chain (NetTCR
A1/A2/A3). -
cdr1b(str, default:'') –CDR1/2/3 amino-acid sequences of the β chain (NetTCR
B1/B2/B3). -
cdr2b(str, default:'') –CDR1/2/3 amino-acid sequences of the β chain (NetTCR
B1/B2/B3). -
cdr3b(str, default:'') –CDR1/2/3 amino-acid sequences of the β chain (NetTCR
B1/B2/B3). -
name(str, default:'') –Human-readable identifier (e.g. a clonotype name). If omitted, :attr:
identifierfalls back to the CDR3α/CDR3β pair. -
trav(str, default:'') –TCR alpha/beta V and J gene assignments. These fields intentionally preserve the caller's spelling: predictor-specific correction and allele handling belong in the predictor adapter.
-
traj(str, default:'') –TCR alpha/beta V and J gene assignments. These fields intentionally preserve the caller's spelling: predictor-specific correction and allele handling belong in the predictor adapter.
-
trbv(str, default:'') –TCR alpha/beta V and J gene assignments. These fields intentionally preserve the caller's spelling: predictor-specific correction and allele handling belong in the predictor adapter.
-
trbj(str, default:'') –TCR alpha/beta V and J gene assignments. These fields intentionally preserve the caller's spelling: predictor-specific correction and allele handling belong in the predictor adapter.
Notes
CDR3β (cdr3b) carries most of the antigen-contact specificity, but
NetTCR-2.2's pan model expects all six loops; leaving loops empty will
encode as fully-padded (uninformative) features.
identifier
property
¶
identifier
Stable string identifier for this receptor.
Uses :attr:name if set, otherwise the CDR3α/CDR3β pair, which
together define the clonotype for most practical purposes.
from_dict
classmethod
¶
from_dict(d)
Deserialize from a dict.
Accepts canonical cdr1a..cdr3b and trav..trbj keys
case-insensitively, plus NetTCR/IMMREP-style short aliases
a1..b3 / A1..B3. Canonical field names win if both a
canonical key and an alias are present.
from_series
classmethod
¶
from_series(row)
Deserialize from a pandas Series or other mapping-like row.
PeptideInput
dataclass
¶
Canonical L-peptide chemical form plus provenance and context.
attachments contains (site, identity) pairs. Sites and identities
are deliberately descriptive strings: adapters may support a documented
subset, and must reject every other form rather than dropping it. Source
occurrence provenance is kept separate from the chemical/context identity
used for inference caching.
chemical_identity_sha256
property
¶
chemical_identity_sha256
Identity of sequence, termini, and attachments only.
inference_identity_sha256
property
¶
inference_identity_sha256
Chemical and model-input context identity, excluding occurrence.
record_identity_sha256
property
¶
record_identity_sha256
Full record identity including occurrence and source provenance.
from_dict
classmethod
¶
from_dict(value)
Restore an input while ignoring unknown same-version fields.
PeptideContext
dataclass
¶
Versioned administration, study, assay, and cellular context.
Every field is descriptive. None means unknown; this class does not
infer patient-specific properties, physiological defaults, or a preferred
formulation, route, or modification.
Annotating tables¶
annotate_table() adds predictions to a DataFrame.
An AnnotationSpec selects the model, endpoint, and
filter to apply. Start with the annotation recipe for a complete
example.
Source: table annotation.
annotate_table ¶
annotate_table(table, specs, peptide_column='peptide', allele_column=None, allele_sep=_DEFAULT_ALLELE_SEP, overwrite=False)
Append predictor score columns to a table, best-allele per row.
Parameters:
-
table(DataFrame) –Input table; returned unmodified with new columns appended to a copy.
-
specs(sequence of AnnotationSpec) –One entry per output column.
-
peptide_column(str, default:'peptide') –Column holding the peptide sequence for each row.
-
allele_column(str, default:None) –Column holding the row's allele(s). May contain several alleles per cell (see
allele_sep). If omitted, predictors run allele-free and no best-allele provenance column is written. -
allele_sep(compiled regex, default:_DEFAULT_ALLELE_SEP) –Splits multiple alleles packed into one cell. Defaults to splitting on whitespace, commas, and semicolons.
-
overwrite(bool, default:False) –If False (default), raise when an output column already exists.
Returns:
-
DataFrame–A copy of
tablewith one numeric column per spec appended (plus a<output_column>_best_allelecolumn when an allele column is given). Rows with no usable prediction getNaN(andNonebest allele).
Notes
Every allele in the table must be supported by each predictor;
commandline predictors raise UnsupportedAllele at construction
otherwise, matching mhctools' behavior elsewhere.
AnnotationSpec
dataclass
¶
One predictor mapped to one output column.
Parameters:
-
predictor(predictor instance, or callable ``alleles -> predictor``) –A built predictor (used as-is), or a factory that builds one given the union of alleles in the table. Commandline predictors validate their alleles at construction, so passing a factory lets the predictor be built with exactly the alleles the table needs.
-
output_column(str) –Name of the appended score column.
-
field(str, default:'affinity') –Which prediction field to write, one of :func:
output_field_tokens("affinity","score","percentile_rank", ...). Determines both the value written and the best-allele direction. -
add_best_allele(bool, default:True) –If True (default) and the table has an allele column, also append a provenance column naming the allele that won each row.
-
best_allele_column(str, default:None) –Name of the provenance column. Defaults to
f"{output_column}_best_allele".
build_predictor ¶
build_predictor(alleles)
Return a predictor instance for the given union of alleles.
A callable that is not itself a predictor (no predict method) is
treated as a factory and called with alleles; anything else is
assumed to already be a built predictor and returned unchanged.
matches_context ¶
matches_context(prediction)
Whether a prediction satisfies this field's context selector.
Models and artifacts¶
fetch() installs managed model artifacts.
list_artifacts() reports their installation status;
integration_status() checks whether a predictor
can run. See getting models for the installation workflow and
licensing for upstream terms.
Source: artifact management and integration checks.
fetch ¶
fetch(name, version=None, data_dir=None, accept_license=False, models=None, all_models=False, high_confidence=False)
Fetch a predictor's external artifacts and return their status.
Native download managers retain ownership of their files. Tested upstream
snapshots are installed under :func:data_path. Predictors whose
parameters are already packaged are successful no-ops.
list_artifacts ¶
list_artifacts(names=None, data_dir=None)
List known packaged, native, managed, and manual artifacts.
integration_status ¶
integration_status(name, check='runnable', data_dir=None, timeout=10)
Report located, runnable, and reproduced state for one integration.
Parameters:
-
name(str) –Artifact/integration name accepted by :func:
artifact_status. -
check((located, runnable, reproduced), default:"located") –Highest capability to attempt. Reproduction probes are intentionally sparse and reference-backed; unsupported checks remain
None. -
data_dir(path - like, default:None) –Override the mhctools-managed artifact directory.
-
timeout(float, default:10) –Bound for command launch and import probes.
Peptidase activity¶
predict_cleavage() assesses individual bonds for
selected peptidases. predict_cleavage_batch()
assesses a set of inputs and processing scenarios.
CleavageInput records sequence and terminal chemistry.
The Python names retain “cleavage”; the biological workflow is explained in the
peptidase activity guide.
See choosing processing models, reading the evidence, and batch assessments. Source: model dispatch, input and result types, and batch assessment.
predict_cleavage ¶
predict_cleavage(peptide, models=None, compartment=None, *, enzyme_states=None)
Return distinct model results without aggregation or enzyme ranking.
compartment filters annotated enzyme locations, not assay validation.
Explicit selections incompatible with that filter are rejected.
predict_cleavage_batch ¶
predict_cleavage_batch(inputs, scenarios, *, predictors=None, reference_panels=(), raise_on_error=False)
Assess named occurrences under explicitly selected biological scenarios.
Parameters:
-
inputs(iterable of dict) –Named sequences/peptide occurrences accepted by normalize_cleavage_input.
-
scenarios(iterable of dict) –Named tumor/APC/extracellular questions with explicit model lists. Compartments filter enzyme locations, not experimental validation.
-
predictors(mapping, default:None) –Already constructed canonical predictors keyed by exact model name. Useful for user-managed assets and source-observation adapters.
-
reference_panels(iterable of dict, default:()) –Experimental source catalogs with
modelandcasesfields, using the PeptidaseSubstrateReference contract. One named panel per assay/condition; exact sequence/chemistry lookup never extrapolates. Panels are retained in the output so imported observations round-trip. -
raise_on_error(bool, default:False) –Raise backend errors instead of retaining failed assessment records.
Returns:
-
dict–Versioned JSON manifest retaining inputs, scenarios, model results, conditional fragments, epitope overlays and declared coverage gaps.
CleavageInput
dataclass
¶
Linear, canonical L-peptide; defaults explicitly assume free termini.
Other chemistry (D-residues, cyclization, conjugation, disulfides, etc.)
is outside this input type. No sequence normalization is performed.
source_start is a zero-based offset in an optional parent sequence.
fragment ¶
fragment(start, end, *, n_term, c_term)
Describe a conditional fragment, explicitly supplying its chemistry.
This does not predict whether or when the fragment is produced. Retained parent termini must retain their original chemistry.