Skip to content

Antigen processing

Antigen processing includes the generation, trimming, destruction, and transport of peptides before MHC loading. Peptidases are the enzymes that hydrolyze peptide bonds within these pathways. The relevant compartment depends on the pathway:

Location Role Models and evidence
Cytosol Proteasomal peptide generation and further trimming or degradation Pepsickle, NetChop, cytosolic peptidases
ER N-terminal trimming of class I precursors after TAP transport ERAMER, ERAP2 motif evidence
Endosomes and lysosomes Class II antigen processing and some class I cross-presentation routes NetCleave, IRAP source observations, coverage limits
Cell surface and extracellular fluids Peptide turnover before or outside cellular uptake Extracellular peptidase models

Experimental work establishes roles for proteasomes in class I peptide generation, ERAP1 in ER trimming, and cathepsin S in class II processing. IRAP-mediated cross-presentation also illustrates why an endosomal location does not imply class II alone. Extracellular turnover can affect which peptides reach a cell; its prediction does not establish uptake or presentation.

Choose the output you need

Use the peptide-level predictors below for a processing score, TAP transport, or precursor trimming. Use the peptidase activity API when you need individual bonds, enzyme identity, and experimental or motif evidence. Both interfaces include Pepsickle and ERAMER; the distinction is the result format and endpoint.

The model-selection guide recommends models for specific biological questions. For peptide stability and delivery endpoints, see peptide half-life and uptake and tissue exposure.

Peptide-level results

The following predictors are allele-independent: Prediction.allele is empty. Each emits one kind, available through a dedicated result accessor:

Kind Accessor Predictors
proteasome_cleavage result.cleavage Pepsickle, NetChop, NetCleave (class I)
endolysosomal_cleavage result.endolysosomal_cleavage NetCleave (class II)
tap_transport result.tap_transport DeepTAP
erap_trimming result.erap_trimming ERAMER

The antigen_processing kind is specifically MHCflurry's combined processing score (result.processing); it is not the name of a compartment or a complete simulation of antigen processing. See MHCflurry.

Proteasome predictors summarize cleavage evidence into a peptide-level score. C-terminal scores depend on the residues after the peptide. Pass c_flanks= and n_flanks= where supported, or scan a protein with predict_proteins().

Pepsickle

Pepsickle is the in-vivo epitope proteasome model of Weeder et al. (2021). Install the package with pip install pepsickle; there is nothing to fetch.

from mhctools import Pepsickle

predictor = Pepsickle()
results = predictor.predict(["SIINFEKL"], c_flanks=["GGG"])
results[0].cleavage.score        # C-terminal cleavage score

by_protein = predictor.predict_proteins({"TP53": "MEEPQSDPSVEPPLSQETFS"},
                                        peptide_lengths=[9])

Without c_flanks the C-terminal position has no downstream context and scores 0.0, so a flank-free call is not a prediction that the peptide is not cleaved. Use human_only=True for the human-trained model, and isolate_subprocess=True to run inference in a subprocess (this avoids macOS OpenMP crashes). The same models, with explicit constitutive/immunoproteasome selection, are exposed per bond through the cleavage API.

NetChop

NetChop 3.1 ships as 32-bit x86 Linux binaries. On macOS and ARM Linux, NetChop automatically runs a user-supplied licensed installation in a digest-pinned compatibility container. Set NETCHOP_HOME to the directory containing bin/netChop (or set NETMHC_BUNDLE_HOME to its parent bundle) and preload the runtime image once:

docker pull --platform linux/386 \
  i386/debian@sha256:75efd55b326373cf69989912388c0d50c5390638af7378d2fedc3aeb9d100e46

Inference runs with Docker network access disabled and image pulling forbidden. The licensed NetChop files and input directory are mounted read-only. Use NetChop(execution="native") or NetChop(execution="container", netchop_dir="/path/to/netchop-3.1") to select a backend explicitly.

from mhctools import NetChop

predictor = NetChop()                       # NETCHOP_HOME / NETMHC_BUNDLE_HOME / PATH
results = predictor.predict(["SIINFEKL"], c_flanks=["GGG"])
results[0].cleavage.score

NetCleave

NetCleave works differently from Pepsickle and NetChop: it emits a single C-terminal cleavage score per peptide, and it covers both the MHC-I proteasomal (NetCleave_I → proteasome_cleavage) and MHC-II endolysosomal (NetCleave_II → endolysosomal_cleavage) pathways.

It needs the residues downstream of the peptide to build the cleavage site, so pass c_flanks or scan proteins. Its weights ship in the git repo; the R dependency mentioned in NetCleave's README is only for its training pipeline, not for prediction.

from mhctools import NetCleave_II

predictor = NetCleave_II()   # NETCLEAVE_DIR, ~/NetCleave, ~/code/NetCleave, then snapshot
# score peptides with their C-terminal flanking residues (>= 3)
results = predictor.predict(["SIINFEKL"], c_flanks=["DGH"])
results[0].endolysosomal_cleavage.score

# or scan a protein so each peptide is scored in real context
by_protein = predictor.predict_proteins({"TP53": "MEEPQ..."}, peptide_lengths=[15])

mhctools fetch netcleave --accept-license installs a pinned snapshot of about 10 MB: the entry script, predictor/, and data/models/. It skips the ~118 MB of IEDB and UniParc databases that only upstream's --generate/--train paths use. Upstream publishes no license file, so here the flag acknowledges that you have confirmed your own use is authorized rather than accepting stated terms; the recorded manifest says "license": "none published". A checkout you manage yourself still takes precedence.

NetCleave's own paper reports that class-II C-terminal cleavage is a much weaker signal than class I (AUC ~0.66 vs ~0.91), so weigh endolysosomal_cleavage scores accordingly.

DeepTAP

TAP (transporter associated with antigen processing) shuttles cytosolic peptides into the ER for MHC-I loading. It is a distinct step from proteasomal cleavage.

DeepTAP is a BiGRU that scores each peptide once, independent of allele, like the cleavage predictors. It emits one tap_transport prediction per peptide with an empty allele. score is in 0-1 (higher = stronger TAP binding); in task_type="reg" mode the predicted affinity in nM is also surfaced as value (lower = stronger).

DeepTAP ships its weights in-repo and is Apache-2.0, but pins an old pytorch-lightning, so mhctools shells out to DeepTAP's own CLI in a separate interpreter. (The checkpoints load fine under modern Lightning too.) Run mhctools fetch deeptap; if the current interpreter lacks torch, set DEEPTAP_PYTHON to one that has it. DEEPTAP_HOME selects a manual checkout.

from mhctools import DeepTAP

DeepTAP.fetch()
predictor = DeepTAP(task_type="cla")       # resolves DEEPTAP_HOME / ~/DeepTAP
results = predictor.predict(["SIINFEKL", "AEASAAAAY"])
results[1].tap_transport.score             # 0-1, higher = stronger TAP binding

DeepTAP's evaluation is self-reported, and no independent TAP benchmark exists for any tool. Treat the score as a pathway signal for prioritizing, not a validated one.

ERAMER

ERAP1 trims the N-termini of 9–16mer precursor peptides in the ER down to the 8–10mers MHC-I presents, the step between TAP transport and MHC loading.

ERAMER scores a precursor by averaging a per-length position-weight-matrix specificity over each residue trimmed off as it is cut toward a target epitope length. It is allele-independent, emitting one erap_trimming prediction per peptide, with score roughly −1…1 (higher = more likely trimmed).

ERAMER is GPLv3 and its PWM ships in a GPL-licensed PWM.xlsx, so mhctools vendors neither. This is a clean-room Python-3 reimplementation of the (Python-2.7) tool's trimming-cascade average, loading the PWM from an upstream ERAMER checkout at runtime. Run mhctools fetch eramer, or point at a manual clone with ERAMER_HOME.

from mhctools import ERAMER

ERAMER.fetch()
predictor = ERAMER(epitope_length=8)       # resolves ERAMER_HOME / ~/ERAMER
results = predictor.predict(["GGGGGVVVVVVAAAEE"])   # a 9-16mer precursor
results[0].erap_trimming.score

ERAMER's evaluation is self-reported and ERAP1 trimming is inherently noisy. Treat the score as a pathway prior, not a validated one.