Skip to content

Choosing processing and peptidase models

Choose a biological endpoint and a plausible compartment before choosing an enzyme. The recommendations below reflect the model's endpoint and available coverage. They do not rank predictors by independently established accuracy. See antigen processing for the pathway overview.

Class I peptide generation

Start with Pepsickle for proteasomal site prediction, with NetChop as another model to compare. Retain the parent sequence and flanks when asking whether an epitope boundary can be generated. The Pepsickle paper describes separate epitope-trained and digestion-trained families.

Use the epitope-trained family for an epitope-processing question. Use the constitutive and immunoproteasome digestion variants when comparing those enzyme profiles. Select the profile from your biological scenario; a tumor or APC label alone does not establish proteasome composition. For full proteins and vaccine constructs, use batch assessments. A cleavage-site score alone does not establish MHC presentation.

ER trimming

Use ERAMER to assess a 9–16-residue N-terminally extended precursor toward a target epitope. Choose its cascade API for the aggregate trimming score, or ERAMER's single-step model for the first exposed bond. These endpoints differ.

ERAP2 motif evidence, available as erap2-basic, can flag its Arg/Lys N-terminal preference. It is a limited rule, not a full-context ERAP2 predictor. ERAP1 experiments show that trimming can generate or destroy epitopes, so more trimming is not universally favorable. Pair this analysis with TAP transport and MHC binding as separate endpoints.

Cytosolic trimming and degradation

For an exposed precursor terminus, select a cytosolic model matching the question: aminopeptidase P for an N-terminal X-Pro bond, DPP8/9 for dipeptide removal, or TPP2 for tripeptide removal. Their motif rules supply recognition evidence. TPP2 and NPEPPS have permissive rules with weak selectivity information.

THOP1 and neurolysin provide exact-substrate experimental observations. Use these to inspect documented substrates; they abstain on novel sequences. None of these models predicts whole-cell peptide survival or competition with TAP and the proteasome.

Class II endolysosomal processing

Use NetCleave class II for a peptide's C-terminal processing proxy, with at least three downstream residues. Its paper reports weaker performance for class II than class I. It does not identify which cathepsin cuts the bond.

For enzyme-specific APC questions, cathepsins and AEP/legumain are relevant, but mhctools currently has no transferable model for their activity on novel sequences. Cathepsin S experiments establish a role in class II presentation; they do not validate a generic sequence-only predictor. Use the curated reference panels for their measured substrates and conditions. Keep these coverage gaps explicit when combining results with class II binding predictions.

Endosomal cross-presentation

IRAP/LNPEP is an endosomal aminopeptidase implicated in class I cross-presentation. The lnpep-observed model is an exact-substrate lookup, not a general cross-presentation predictor. Its source assay used purified enzyme at pH 8.0; do not interpret it as calibration in an acidified endosome.

Use batch scenarios to record an explicit routing hypothesis and any conditional fragments. Compartment labels filter models by enzyme location; they do not predict where an antigen travels.

Extracellular peptide degradation

For a human DPP4 substrate, use DPP4 qPISA to assess removal of the first two residues from a free N-terminus. The qPISA study models substrate depletion by purified enzyme under its assay conditions, not serum half-life.

For broader candidate screening, select extracellular models by the bond: CPN or activated CPB2 for a C-terminal Lys/Arg; ACE for ordinary C-terminal dipeptide removal; aminopeptidase P for N-terminal X-Pro; or FAP for its proline-associated activities. MME, ENPEP, and ANPEP provide narrower preference flags. Preserve terminal chemistry and, for CPB2, the explicit active-enzyme assumption.

Use --compartment extracellular for the broader annotated panel, or select serum or plasma deliberately. These filters establish neither enzyme concentration nor calibration in that fluid. For a duration endpoint, consult peptide half-life models; for delivery, consult uptake and exposure results. Do not combine the peptidase outputs into a serum-stability probability.