Skip to content

API reference

Use this reference to look up constructors, fields, and methods. For a complete first prediction, start with getting started.

Task Reference Explanation and examples
Read predictions and scores Results Results and DataFrames
Describe a peptide or TCR Inputs Input shapes
Add predictions to a table Table annotation Annotation recipe
Download or inspect models Models and artifacts Getting models
Assess individual peptide bonds Peptidase activity Activity guide

Predictor-specific setup and examples are in the family guides. The predictor matrix lists Python classes and command-line names for every supported model.

Results

predict() returns one PeptideResult per peptide. Its predictions are Prediction objects, each describing one endpoint and MHC or TCR context. Kind names the endpoint; MeasurementContext records its measurement semantics. MultiSample runs a predictor for several named MHC allele sets.

Read results and DataFrames for examples and prediction kinds for score meanings and units. Source: prediction types and sample helper.

PeptideResult dataclass

All predictions for one peptide (a tuple of Prediction objects).

kinds property

kinds

Set of Kind values present in this result.

alleles property

alleles

Set of allele strings present in this result.

tcrs property

tcrs

Set of TCR identifiers present in this result.

affinity property

affinity

Best affinity prediction, or None.

presentation property

presentation

Best presentation prediction, or None.

stability property

stability

Best stability prediction, or None.

immunogenicity property

immunogenicity

Best immunogenicity prediction, or None.

processing property

processing

Best antigen-processing prediction, or None.

cleavage property

cleavage

Best proteasomal cleavage prediction, or None.

endolysosomal_cleavage property

endolysosomal_cleavage

Best endolysosomal (MHC-II) cleavage prediction, or None.

tap_transport property

tap_transport

Best TAP transport prediction, or None.

erap_trimming property

erap_trimming

Best ERAP1 trimming prediction, or None.

peptide_half_life property

peptide_half_life

Peptide half-life when all available results share one context.

Comparing different matrices or systemic scopes is intentionally rejected; callers can select those contexts explicitly instead.

serum_half_life property

serum_half_life

Best peptide half-life measured in serum, or None.

plasma_half_life property

plasma_half_life

Best peptide half-life measured in plasma, or None.

blood_half_life property

blood_half_life

Best peptide half-life measured in whole blood, or None.

tcr_binding property

tcr_binding

Best pMHC:TCR binding prediction, or None.

filter

filter(kind=None, allele=None)

Filter preds. None means don't filter on that field.

to_dict

to_dict()

Serialize to a JSON-friendly dict.

from_dict classmethod

from_dict(d)

Deserialize from a dict (as produced by :meth:to_dict).

best_by

best_by(kind, field)

Return the prediction of kind with the best field, or None.

"Best" direction comes from :func:best_direction — score is max-better, percentile_rank is min-better, and value is kind-dependent (IC50 lower-better, half-life higher-better).

Predictions with None for field are skipped. Predictions with an allele are preferred; if none of the matching kind have an allele, falls back to allele-less predictions (e.g. processing predictors that emit allele-independent scores).

best_by_score

best_by_score(kind)

Best prediction of kind by score (max-better).

best_by_rank

best_by_rank(kind)

Best prediction of kind by percentile_rank (min-better).

best_by_value

best_by_value(kind)

Best prediction of kind by value. Direction is kind-specific (see :data:VALUE_BEST_DIRECTIONS). Raises ValueError for kinds without a registered value direction.

Prediction dataclass

Single prediction from one model on one peptide.

to_dict

to_dict()

Serialize to a JSON-friendly dict.

from_dict classmethod

from_dict(d)

Deserialize from a dict (as produced by :meth:to_dict).

Kind

String constants for prediction kinds.

You can use Kind.pMHC_affinity or just "pMHC_affinity" — they're the same string. These constants name what is measured, but predictor instances define the MHC context required for their supported kinds through kind_support().

MeasurementContext dataclass

Versioned semantics shared by every prediction.

unit and transform describe :attr:Prediction.value; the stored physical value remains linear, so transform must currently be "linear" when a value is present. Predictor-native or confidence outputs belong in score and are identified by score_semantics. Ordinary model outputs receive a small cached default; assay-specific wrappers fill only the fields they actually know. Unknown descriptive fields are None, never a plausible biological default.

to_dict

to_dict()

Serialize to a JSON-friendly dictionary.

from_dict classmethod

from_dict(value)

Deserialize a context while ignoring forward-compatible fields.

MultiSample

Run a predictor across multiple samples, each with its own alleles.

Parameters:

  • samples (dict) –

    Mapping of sample_name -> list of allele strings. Example: {"pat001": ["HLA-A02:01", "HLA-B07:02"], "pat002": ["HLA-A01:01", "HLA-B08:01"]}

  • predictor_class (class) –

    A predictor class (e.g. NetMHCpan41) that accepts an alleles keyword argument.

  • **predictor_kwargs –

    Additional keyword arguments forwarded to the predictor constructor.

predict

predict(peptides)

Returns:

  • dict mapping sample_name -> list of PeptideResult –

predict_dataframe

predict_dataframe(peptides)

predict() flattened to a DataFrame with sample_name column.

predict_proteins

predict_proteins(sequence_dict, peptide_lengths=None)

Returns:

  • dict mapping sample_name -> {sequence_name: list of PeptideResult} –

predict_proteins_dataframe

predict_proteins_dataframe(sequence_dict, peptide_lengths=None)

predict_proteins() flattened to a DataFrame with sample_name column.

Inputs

TCR carries receptor sequences and gene identifiers. PeptideInput records a peptide's sequence, terminal chemistry, attachments, and source occurrence. PeptideContext adds study, formulation, and assay context when those details are known.

See input shapes, the TCR guide, and peptide exposure inputs. Source: TCR types and peptide types.

TCR dataclass

A paired αβ T-cell receptor, described by its six CDR loops.

Parameters:

  • cdr1a (str, default: '' ) –

    CDR1/2/3 amino-acid sequences of the α chain (NetTCR A1/A2/A3).

  • cdr2a (str, default: '' ) –

    CDR1/2/3 amino-acid sequences of the α chain (NetTCR A1/A2/A3).

  • cdr3a (str, default: '' ) –

    CDR1/2/3 amino-acid sequences of the α chain (NetTCR A1/A2/A3).

  • cdr1b (str, default: '' ) –

    CDR1/2/3 amino-acid sequences of the β chain (NetTCR B1/B2/B3).

  • cdr2b (str, default: '' ) –

    CDR1/2/3 amino-acid sequences of the β chain (NetTCR B1/B2/B3).

  • cdr3b (str, default: '' ) –

    CDR1/2/3 amino-acid sequences of the β chain (NetTCR B1/B2/B3).

  • name (str, default: '' ) –

    Human-readable identifier (e.g. a clonotype name). If omitted, :attr:identifier falls back to the CDR3α/CDR3β pair.

  • trav (str, default: '' ) –

    TCR alpha/beta V and J gene assignments. These fields intentionally preserve the caller's spelling: predictor-specific correction and allele handling belong in the predictor adapter.

  • traj (str, default: '' ) –

    TCR alpha/beta V and J gene assignments. These fields intentionally preserve the caller's spelling: predictor-specific correction and allele handling belong in the predictor adapter.

  • trbv (str, default: '' ) –

    TCR alpha/beta V and J gene assignments. These fields intentionally preserve the caller's spelling: predictor-specific correction and allele handling belong in the predictor adapter.

  • trbj (str, default: '' ) –

    TCR alpha/beta V and J gene assignments. These fields intentionally preserve the caller's spelling: predictor-specific correction and allele handling belong in the predictor adapter.

Notes

CDR3β (cdr3b) carries most of the antigen-contact specificity, but NetTCR-2.2's pan model expects all six loops; leaving loops empty will encode as fully-padded (uninformative) features.

identifier property

identifier

Stable string identifier for this receptor.

Uses :attr:name if set, otherwise the CDR3α/CDR3β pair, which together define the clonotype for most practical purposes.

cdr_dict

cdr_dict()

Return the six CDRs keyed by NetTCR feature name (a1..b3).

gene_dict

gene_dict()

Return V/J assignments using standard AIRR-style gene names.

to_dict

to_dict()

Serialize to a JSON-friendly dict.

from_dict classmethod

from_dict(d)

Deserialize from a dict.

Accepts canonical cdr1a..cdr3b and trav..trbj keys case-insensitively, plus NetTCR/IMMREP-style short aliases a1..b3 / A1..B3. Canonical field names win if both a canonical key and an alias are present.

from_series classmethod

from_series(row)

Deserialize from a pandas Series or other mapping-like row.

PeptideInput dataclass

Canonical L-peptide chemical form plus provenance and context.

attachments contains (site, identity) pairs. Sites and identities are deliberately descriptive strings: adapters may support a documented subset, and must reject every other form rather than dropping it. Source occurrence provenance is kept separate from the chemical/context identity used for inference caching.

chemical_identity_sha256 property

chemical_identity_sha256

Identity of sequence, termini, and attachments only.

inference_identity_sha256 property

inference_identity_sha256

Chemical and model-input context identity, excluding occurrence.

record_identity_sha256 property

record_identity_sha256

Full record identity including occurrence and source provenance.

to_dict

to_dict()

Return a JSON-friendly representation.

from_dict classmethod

from_dict(value)

Restore an input while ignoring unknown same-version fields.

PeptideContext dataclass

Versioned administration, study, assay, and cellular context.

Every field is descriptive. None means unknown; this class does not infer patient-specific properties, physiological defaults, or a preferred formulation, route, or modification.

to_dict

to_dict()

Return a JSON-friendly representation.

from_dict classmethod

from_dict(value)

Restore a context while ignoring unknown same-version fields.

Annotating tables

annotate_table() adds predictions to a DataFrame. An AnnotationSpec selects the model, endpoint, and filter to apply. Start with the annotation recipe for a complete example.

Source: table annotation.

annotate_table

annotate_table(table, specs, peptide_column='peptide', allele_column=None, allele_sep=_DEFAULT_ALLELE_SEP, overwrite=False)

Append predictor score columns to a table, best-allele per row.

Parameters:

  • table (DataFrame) –

    Input table; returned unmodified with new columns appended to a copy.

  • specs (sequence of AnnotationSpec) –

    One entry per output column.

  • peptide_column (str, default: 'peptide' ) –

    Column holding the peptide sequence for each row.

  • allele_column (str, default: None ) –

    Column holding the row's allele(s). May contain several alleles per cell (see allele_sep). If omitted, predictors run allele-free and no best-allele provenance column is written.

  • allele_sep (compiled regex, default: _DEFAULT_ALLELE_SEP ) –

    Splits multiple alleles packed into one cell. Defaults to splitting on whitespace, commas, and semicolons.

  • overwrite (bool, default: False ) –

    If False (default), raise when an output column already exists.

Returns:

  • DataFrame –

    A copy of table with one numeric column per spec appended (plus a <output_column>_best_allele column when an allele column is given). Rows with no usable prediction get NaN (and None best allele).

Notes

Every allele in the table must be supported by each predictor; commandline predictors raise UnsupportedAllele at construction otherwise, matching mhctools' behavior elsewhere.

AnnotationSpec dataclass

One predictor mapped to one output column.

Parameters:

  • predictor (predictor instance, or callable ``alleles -> predictor``) –

    A built predictor (used as-is), or a factory that builds one given the union of alleles in the table. Commandline predictors validate their alleles at construction, so passing a factory lets the predictor be built with exactly the alleles the table needs.

  • output_column (str) –

    Name of the appended score column.

  • field (str, default: 'affinity' ) –

    Which prediction field to write, one of :func:output_field_tokens ("affinity", "score", "percentile_rank", ...). Determines both the value written and the best-allele direction.

  • add_best_allele (bool, default: True ) –

    If True (default) and the table has an allele column, also append a provenance column naming the allele that won each row.

  • best_allele_column (str, default: None ) –

    Name of the provenance column. Defaults to f"{output_column}_best_allele".

build_predictor

build_predictor(alleles)

Return a predictor instance for the given union of alleles.

A callable that is not itself a predictor (no predict method) is treated as a factory and called with alleles; anything else is assumed to already be a built predictor and returned unchanged.

direction_op

direction_op()

max or min callable for reducing candidate predictions.

matches_context

matches_context(prediction)

Whether a prediction satisfies this field's context selector.

Models and artifacts

fetch() installs managed model artifacts. list_artifacts() reports their installation status; integration_status() checks whether a predictor can run. See getting models for the installation workflow and licensing for upstream terms.

Source: artifact management and integration checks.

fetch

fetch(name, version=None, data_dir=None, accept_license=False, models=None, all_models=False, high_confidence=False)

Fetch a predictor's external artifacts and return their status.

Native download managers retain ownership of their files. Tested upstream snapshots are installed under :func:data_path. Predictors whose parameters are already packaged are successful no-ops.

list_artifacts

list_artifacts(names=None, data_dir=None)

List known packaged, native, managed, and manual artifacts.

integration_status

integration_status(name, check='runnable', data_dir=None, timeout=10)

Report located, runnable, and reproduced state for one integration.

Parameters:

  • name (str) –

    Artifact/integration name accepted by :func:artifact_status.

  • check ((located, runnable, reproduced), default: "located" ) –

    Highest capability to attempt. Reproduction probes are intentionally sparse and reference-backed; unsupported checks remain None.

  • data_dir (path - like, default: None ) –

    Override the mhctools-managed artifact directory.

  • timeout (float, default: 10 ) –

    Bound for command launch and import probes.

Peptidase activity

predict_cleavage() assesses individual bonds for selected peptidases. predict_cleavage_batch() assesses a set of inputs and processing scenarios. CleavageInput records sequence and terminal chemistry. The Python names retain “cleavage”; the biological workflow is explained in the peptidase activity guide.

See choosing processing models, reading the evidence, and batch assessments. Source: model dispatch, input and result types, and batch assessment.

predict_cleavage

predict_cleavage(peptide, models=None, compartment=None, *, enzyme_states=None)

Return distinct model results without aggregation or enzyme ranking.

compartment filters annotated enzyme locations, not assay validation. Explicit selections incompatible with that filter are rejected.

predict_cleavage_batch

predict_cleavage_batch(inputs, scenarios, *, predictors=None, reference_panels=(), raise_on_error=False)

Assess named occurrences under explicitly selected biological scenarios.

Parameters:

  • inputs (iterable of dict) –

    Named sequences/peptide occurrences accepted by normalize_cleavage_input.

  • scenarios (iterable of dict) –

    Named tumor/APC/extracellular questions with explicit model lists. Compartments filter enzyme locations, not experimental validation.

  • predictors (mapping, default: None ) –

    Already constructed canonical predictors keyed by exact model name. Useful for user-managed assets and source-observation adapters.

  • reference_panels (iterable of dict, default: () ) –

    Experimental source catalogs with model and cases fields, using the PeptidaseSubstrateReference contract. One named panel per assay/condition; exact sequence/chemistry lookup never extrapolates. Panels are retained in the output so imported observations round-trip.

  • raise_on_error (bool, default: False ) –

    Raise backend errors instead of retaining failed assessment records.

Returns:

  • dict –

    Versioned JSON manifest retaining inputs, scenarios, model results, conditional fragments, epitope overlays and declared coverage gaps.

CleavageInput dataclass

Linear, canonical L-peptide; defaults explicitly assume free termini.

Other chemistry (D-residues, cyclization, conjugation, disulfides, etc.) is outside this input type. No sequence normalization is performed. source_start is a zero-based offset in an optional parent sequence.

fragment

fragment(start, end, *, n_term, c_term)

Describe a conditional fragment, explicitly supplying its chemistry.

This does not predict whether or when the fragment is produced. Retained parent termini must retain their original chemistry.