Skip to content

Recipes

Short answers to common tasks. Each assumes you have a predictor built as in the quickstart.

Scan proteins instead of peptides

predict_proteins() takes a dictionary of sequences and returns {sequence_name: list[PeptideResult]}, with each result's offset set:

proteins = predictor.predict_proteins(
    {"TP53": "MEEPQSDPSVEPPLSQETFS...", "KRAS": "MTEYKLVVVGAGGVGKS..."},
    peptide_lengths=[9, 10],
)

for r in proteins["TP53"]:
    if r.affinity and r.affinity.value < 500:
        print(f"  offset={r.offset} {r.peptide} IC50={r.affinity.value:.0f}")

Run many samples with different genotypes

from mhctools import MultiSample, MHCflurry

ms = MultiSample(
    samples={
        "pat001": ["HLA-A*02:01", "HLA-B*07:02"],
        "pat002": ["HLA-A*01:01", "HLA-B*08:01"],
    },
    predictor_class=MHCflurry,
)

results = ms.predict(["SIINFEKL", "GILGFVFTL"])       # {sample: [PeptideResult]}
protein_results = ms.predict_proteins({"TP53": "MEEPQ..."})  # {sample: {seq: [...]}}

df = ms.predict_dataframe(["SIINFEKL"])               # flat, with sample_name
df = ms.predict_proteins_dataframe({"TP53": "MEEPQ..."})

Add predictor scores to an existing table

A benchmark table usually has columns like sample_id, hit, peptide and a per-row genotype, and needs scores appended. annotate_table is I/O-free and works on any DataFrame:

from mhctools import annotate_table, AnnotationSpec, NetMHCpan42_BA

annotated = annotate_table(
    df,
    [AnnotationSpec(
        predictor=lambda alleles: NetMHCpan42_BA(alleles=alleles),
        output_column="netmhcpan4.2.ba",
        field="affinity")],
    peptide_column="peptide",
    allele_column="hla")

There's a CLI equivalent that reads and writes CSV.