Peptidase models¶
The model panel behind predict_cleavage() and mhctools cleavage. Which
enzyme, where it acts, what pattern it assesses and how strongly a match should
be read is below; what the result states mean is in reading the
evidence.
For model selection by biological endpoint and compartment, see choosing processing and peptidase models.
Running the panel¶
mhctools cleavage --list-models
mhctools cleavage --list-models --json
mhctools cleavage --sequence RPPGFSPFR --model app2-xp --model cpn-basic
mhctools cleavage --sequence VPYGSFKHV --compartment cytosol --out cleavage.json
mhctools cleavage --sequence HAEGTFTSD --model dpp4-qpisa --n-term acetylated
mhctools cleavage --sequence SIINFEKL --model pepsickle-in-vivo-human-only
from mhctools import predict_cleavage
results = predict_cleavage("TSGPNQ", models=["fap-endo-gp", "prep-pro"])
The default panel evaluates all 20 built-in models and returns separate
results. --model and --sequence can be repeated. --list-models prints a
compact discovery table; add --json for its full machine-readable catalog.
The eight optional Pepsickle models cover epitope and C/I digestion families;
see proteasome models.
Missing assets and uninspected external runtimes are listed as unresolved.
Predictions carry SHA-256 identities for the actual runtime's weights,
inference code, feature code and dependency metadata.
Prediction JSON uses schema version 2: its top-level models object stores each
full provenance record once, keyed by model name, and each item in results
references that name in its model field. The output retains unmatched and
unsupported results, source coordinates and chemistry. Numerical scores are
written with six significant digits. --source-start is a zero-based offset
shared by the supplied inputs; use separate calls when fragments have different
offsets.
Compartment filtering uses exact, conservative enzyme-location annotations.
serum, plasma, extracellular, cytosol, endosome and er are
distinct. The extracellular filter is useful for broader candidate
screening; serum is not an exhaustive inventory of everything potentially
present in a serum sample. Presence, concentration, activation, inhibitors and
exposure are not inferred. No combination is converted into overall stability.
The models¶
Strictness grades are explained in reading the evidence.
| Model | Assessed recognition pattern | Strictness | Main scope and primary evidence |
|---|---|---|---|
dpp4-qpisa |
N-terminal P2-P1|P1′ score | scored | Human DPP4; quantitative model described above |
ace-dipeptidyl |
C-terminal |non-Pro–non-Asp/Glu | required | Ordinary human ACE dipeptide activity; angiotensin assays |
mme-hydrophobic |
Selected P1′ residues Phe/Ile/Leu/Tyr | preferred | Human neprilysin; kidney peptide assays; incomplete whole-sequence specificity |
cpb2-basic |
C-terminal |Lys/Arg, explicitly active enzyme | required | Human TAFIa; chemerin cleavage; unknown/zymogen/inactive states abstain |
cpn-basic |
C-terminal |Lys/Arg | required | Human plasma CPN; Oshima et al. 1975 |
app1-xp |
N-terminal X|Pro | required | Cytosolic human XPNPEP1; Cottrell et al. 2000; manganese dependent, distinct gene product from XPNPEP2 |
app2-xp |
N-terminal X|Pro | required | Human XPNPEP2; Molinaro et al. |
fap-dipeptidyl |
N-terminal X-Pro|non-Pro | required | Human FAP; Edosada et al. |
fap-endo-gp |
Gly-Pro|non-Pro | required | FAP endopeptidase; substrate profiling, prime-side constraint |
enpep-acidic |
N-terminal Asp/Glu|X | preferred | Aminopeptidase A; human specificity study; calcium and sequence affect activity |
anpep-ala |
N-terminal Ala|X preference | permissive | Aminopeptidase N; human structure/biochemistry; many other substrates omitted |
dpp8-xp-xa |
N-terminal X-(Pro/Ala)|non-Pro | required | Cytosolic DPP8; characterization, degradomics |
dpp9-xp-xa |
N-terminal X-(Pro/Ala)|non-Pro | required | Cytosolic DPP9; antigen processing, same degradomics study |
tpp2-tripeptidyl |
N-terminal tripeptide, non-Pro P1 and P1′ | permissive | Cytosolic TPP2; RU1 precursor processing; topology, not selectivity |
npepps-n-terminal |
First bond lacking Gly/Pro context | permissive | Puromycin-sensitive aminopeptidase; same RU1 study; broad enzyme, weak flag |
prep-pro |
Internal X-Pro|X | required | PREP/POP, conservative 4–30-residue domain; human profiling; flanking preferences omitted |
erap2-basic |
N-terminal Arg/Lys|X preference | preferred | ERAP2 in ER; biochemistry, peptide structures; not a full-context predictor |
thop1-observed |
Exact-sequence source lookup | source observations | Cytosolic THOP1; Knight et al. 1995 |
nln-observed |
Exact-sequence source lookup | source observations | Cytosolic NLN; human neurolysin structures and LC-MS |
lnpep-observed |
Exact-sequence source lookup | source observations | Endosomal LNPEP/IRAP; Georgiadou et al. 2010 |
eramer-step (optional) |
Initial N-terminal bond; length-specific PWM score | scored | ERAP1 in ER; ERAMER, 9–16 residues |
The motif models return decisions without numerical scores. For example,
APP removes the first residue of RPPGFSPFR at R|PPGFSPFR; DPP-like
activity would remove two residues and is a separate assessment. FAP's
endopeptidase rule can also assess an N-acetylated input, supported by its
blocked-substrate activity. ACE also accepts N-acetylation, and MME accepts
C-amidation. Other rules conservatively accept free termini only;
an unsupported modified input does not establish that
its bonds resist enzymatic cleavage.
Proteasome models¶
Eight optional Pepsickle models cover the epitope and digestion families, with explicit constitutive (C) or immunoproteasome (I) selection where the family has it. They are excluded from the default panel and selected by exact name:
| Model | Family | Context | Proteasome |
|---|---|---|---|
pepsickle-in-vivo-human-only |
epitope-trained neural ensemble | 8 residues before and after P1 | agnostic |
pepsickle-in-vivo-all-mammal |
epitope-trained neural ensemble | 8 residues before and after P1 | agnostic |
pepsickle-in-vitro-2-human-only-constitutive |
digestion-trained neural ensemble | 3 residues before and after P1 | C |
pepsickle-in-vitro-2-human-only-immunoproteasome |
digestion-trained neural ensemble | 3 residues before and after P1 | I |
pepsickle-in-vitro-2-all-mammal-constitutive |
digestion-trained neural ensemble | 3 residues before and after P1 | C |
pepsickle-in-vitro-2-all-mammal-immunoproteasome |
digestion-trained neural ensemble | 3 residues before and after P1 | I |
pepsickle-in-vitro-all-mammal-constitutive |
gradient-boosted digestion model | needs the isolated scikit-learn 0.23.2 runtime (see below) | C |
pepsickle-in-vitro-all-mammal-immunoproteasome |
gradient-boosted digestion model | needs the isolated scikit-learn 0.23.2 runtime | I |
The Python Pepsickle wrapper takes model_type and proteasome_type (C
or I) for digestion models. It rejects a proteasome type for the epitope
model and human_only=True for gradient boosting, which upstream ignores.
Model metadata identifies actual weights, code, population and model family.
The upstream endpoint sentinel is excluded from canonical bond results.
The legacy gradient-boosted runtime¶
Provision it with Docker:
python scripts/setup_test_backends.py pepsickle
source env/test-backends/activate.sh
This pins Python 3.8.20, scikit-learn 0.23.2 and its companion packages in a
separate container. Its generated PEPSICKLE_GB_PYTHON launcher selects that
runtime only for gradient boosting and runs inference with networking
disabled. Neural models keep their existing runtime. A separately managed
interpreter can be selected with Pepsickle(python_executable=...) or
PEPSICKLE_PYTHON for every model family.
Prediction provenance comes from the selected interpreter, including package versions and actual code/weight hashes. A changed identity between inspection and inference causes failure. Catalog listing does not start external runtimes; their identities remain unresolved until the predictor is selected. No model is silently substituted when a runtime is missing or incompatible.
Adapter agreement with upstream inference establishes implementation conformance, not accuracy for tumor/APC processing or vaccine-peptide survival.
Human DPP4 qPISA¶
The Gudipati et al. 2024 paper
reports a model fitted to substrate depletion by purified human DPP4. The
source assay used tryptic HeLa peptides in HEPES pH 7.4, at 21 C for 4 hours.
The implementation independently evaluates the three rearranged terms in
Dataset EV2: P1 + P2:P1 + P1:P1-prime, where the first three peptide
residues are P2, P1 and P1-prime. Only bond 2 is assessed.
Higher scores indicate greater predicted log2 depletion relative to buffer control in that assay. Negative values are retained. Scores are not cleavage probabilities, serum half-lives or stability ranks across enzymes. Compartment metadata records where the enzyme can act, not where the model has been calibrated. Physiological exposure, structure, concentration and competing enzymes are not modeled.
All triplets with complete coefficients can be evaluated, including those without Pro or Ala at P1. Of 8,000 canonical triplets, 6,420 have complete coefficients; the other 1,580 return an explicit missing-coefficient reason. Evaluability does not establish that an individual triplet occurred in the training data. The related C. elegans DPF-3 model is not used for human DPP8/9.
Parameter provenance¶
mhctools/data/dpp4_qpisa.json contains the numeric cells from
44320_2024_71_MOESM3_ESM.xlsx, sheet dpp4_modelParams. Published NA
cells become JSON null; no numerical imputation or refitting is performed.
The source workbook SHA-256 is
ee449da13b5ec66fd6fb08c16203da8c44f2e9c10572a355abdcb985694c6ddc.
It was retrieved via the Europe PMC supplementary archive.
The article assigns associated data to CC0,
unless otherwise credited; Dataset EV2 has no separate credit restriction.
Attribution: Rajani Kanth Gudipati and colleagues, 2024, DOI above. No figures
or upstream R source code are redistributed.
Serum and extracellular candidates¶
mhctools cleavage --sequence DRVYIHPFHL --model ace-dipeptidyl --model mme-hydrophobic
mhctools cleavage --sequence YFPGQFAFSK --model cpb2-basic --enzyme-state CPB2=active
mhctools benchmark --reference-cleavage serum --out serum-reference.json
In Python, supply enzyme_states={"CPB2": "active"} to predict_cleavage,
or get_cleavage_model("cpb2-basic", enzyme_state="active"). The result
records that assumption in conditions. CPB2 requires proteolytic activation;
its presence as a zymogen does not establish activity. unknown, zymogen
and inactive return no assessed sites. Activation experiments
also show that the surrounding coagulation environment matters.
ACE's ordinary rule assesses only the bond before the last two residues.
It recognizes angiotensin I cleavage at bond 8 and the N-acetyl-SDKP
bond 2. It abstains on amidated substance P even though human ACE is known
to cleave that peptide at bonds 8 and 9 through exceptional processing.
MME flags selected hydrophobic residues after a bond, including substance P's
reported bonds 6, 7 and 9; it accepts that peptide's C-terminal amide.
These human enzyme observations
do not establish a transferable model of all substrates. The MME rule's
2–30-residue domain is a conservative implementation scope; longer substrates
can exist. Use the extracellular filter to include MME: location annotations
do not assert that purified-kidney specificity was calibrated in serum.
The packaged serum-candidate reference contains 17 source observations, including two explicitly reported non-cleavages. Fifteen are assessable by the corresponding rules; the two exceptional ACE/substance-P observations abstain. Cross-enzyme combinations are unassessed. The source IDs, chemical forms and known or unspecified experimental conditions are retained. These small reproduction controls do not establish specificity, serum half-life performance, or validity on long peptides. The report explicitly shows missing serum-half-life and long-peptide evidence.
Cytosolic and endosomal candidates¶
mhctools cleavage --sequence VPYGSFKHV --compartment cytosol
mhctools cleavage --sequence KSLYNTVATL --compartment endosome
mhctools cleavage --sequence RPPGFSPFR --model thop1-observed --model app1-xp
mhctools benchmark --reference-cleavage intracellular --out intracellular-reference.json
Antigen-processing peptidases are the reason this panel exists, but three of
them publish their specificity as whole-substrate outcomes rather than as a
transferable pattern. Inventing a motif from those papers would misrepresent
them, so thop1-observed, nln-observed and lnpep-observed are source
references instead: an exact sequence and chemical form returns what the
experiment reported, and anything else abstains. There is no nearest-neighbour
matching and no extrapolation. Results carry a substrate_observation of
cleavage_reported or no_cleavage_detected alongside any bonds the source's
product identities actually pin down. Where a source saw degradation but no
intermediate, cleavage is recorded with no bond rather than guessing one.
A reported non-cleavage never becomes a per-bond label. no_cleavage_detected
means that assay saw no loss of that peptide under its own conditions and
detection limits, which is not the same as a bond that cannot be cleaved.
The remaining cytosolic enzymes are ordinary motif rules, and their grades
matter. app1-xp is required: aminopeptidase P is defined by hydrolysing
the X-Pro bond, so a non-match is real evidence. tpp2-tripeptidyl and
npepps-n-terminal are permissive. Removing three residues describes TPP2's
topology, not which peptides it turns over, and puromycin-sensitive
aminopeptidase is broad enough that its Gly/Pro exclusion is only a hint.
Do not read a TPP2 match as a prediction that the peptide is consumed.
THOP1 and NLN are closely related and are deliberately kept apart. They cleave neurotensin at different bonds and their specificities can be swapped by mutating two active-site residues, so neither model's observations transfer to the other. Only human-enzyme observations are curated: the widely cited bradykinin, enkephalin and neurotensin results for neurolysin come from rat or species-unspecified preparations and are excluded rather than relabelled human.
lnpep-observed covers IRAP as an endosomal cross-presentation candidate, not
a second ER enzyme. Its source digested peptides at pH 8.0 with purified
enzyme, so the records describe that experiment, not an acidified endosome.
The packaged intracellular reference contains 42 source observations across
four models, including six reported non-cleavages. Forty-one reproduce; the
amidated substance P record abstains because C-amidation is outside the
aminopeptidase P rule's documented input domain. One resistant IRAP precursor
from Georgiadou 2010 is excluded entirely: the paper prints DIRSSVQNKL in
its results and Table I but DIRSSQVNKL in the Figure 3F caption, and an
exact-sequence catalog cannot silently pick one. Both strings are retained in
the dataset notice, and the discrepancy is tracked in
known gaps.
Thimet oligopeptidase contributes no non-cleavage record at all. Every resistant peptide in its source is a hydroxyproline analogue or carries an N-terminal pyroglutamate, and both are outside the canonical-peptide input type. Its absence from the negatives is a curation limit, not a finding.
The report shows no evidence for antigen presentation in primary dendritic cells and none for cleavage measured in cytosol rather than purified enzyme. Those gaps are the point: nothing here calibrates how long a peptide survives in the cytosol of an antigen-presenting cell, where the proteasome, competing aminopeptidases and TAP transport all act at once.
ERAP1 and existing processing models¶
mhctools fetch eramer
# Install openpyxl in the environment if it is not already available.
mhctools cleavage --sequence LAAAFGAAA --model eramer-step
Alternatively, construct ERAMERCleavage(pwm_path="/path/to/PWM.xlsx") or set
ERAMER_HOME. This optional model is listed without loading assets and is
excluded from the default panel. Explicit selection fails clearly if the
external asset or its runtime is missing.
eramer-step computes the existing ERAMER intermediate PWM specificity for
one exposed 9–16-residue precursor, assigning it to bond 1. It does not report
the average of later trimming intermediates. Its version contains the SHA-256
of the actual workbook snapshot used for inference. The GPL-licensed workbook
is loaded at runtime and is not included in the mhctools distribution.
The existing ERAMER cascade API, NetChop, Pepsickle and other proteasome
predictors remain available through their existing interfaces. Pepsickle also
has a canonical PepsickleCleavage.predict() facade and the two CLI model names
shown above. Array index i maps to internal bond i + 1; the upstream final
zero is an endpoint sentinel and is never emitted as a bond. Canonical results
default to subprocess isolation because the upstream package deserializes a
pickle, while still requiring users to trust the installed model artifact.
Processing-model output scales must not be mixed with qPISA scores or motif
decisions.