Skip to content

Cleavage coverage and validation status

This page accompanies the batch API. Software conformance, source-observation reproduction and independent prediction validation are different claims. No new held-out biological performance number is claimed.

Proteasome variants

The Pepsickle author implementation provides epitope, gradient-boosted digestion and neural digestion families. The paper describes its independent digestion evaluation. The batch adapter's live checks compare human/all-mammal neural C/I profiles with the upstream functions on the same sequence and assert that C/I can change the result. These are conformance checks, not independent accuracy measurements.

The paper repository's MASTER.sh expects held-out inputs under data/validation_data/digestion_data/raw/ and processed, overlap-filtered validation FASTAs. Those paths are absent from the inspected repository tree at c448c4db81925afad78477e74a7d25e0209d3bce. Its available data/raw/digestion_map_files instead feed the training-data extraction path. Reusing those maps and calling them held-out would be incorrect. The Wada case study below recovers source products from one of the paper's held-out studies and audits the raw training maps. Reconstructing the full processed validation partition and establishing family-level independence are open work (known gaps); the case study does not claim to reproduce the paper's overall metrics.

The gradient-boosted artifact records scikit-learn 0.23.2. It fails under current scikit-learn (sklearn.ensemble._gb_losses is missing). The isolated Python 3.8.20/0.23.2 runtime for the gradient-boosted models executes both C/I routes and matches direct upstream inference with networking disabled. It records the actual subprocess's package, code and weight identity. The MAGE-A3 sequence in this conformance test occurs in upstream training data; it is explicitly not held-out validation. Neural and gradient-boosted predictions are not silently substituted.

APC endolysosomal enzymes

The current built-in panel has no transferable cathepsin S/L/B or AEP model. IRAP is an exact-substrate source catalog, and NetCleave-II is a class-II C-terminal processing proxy, not an enzyme-specific cathepsin predictor. The absence is explicit in batch coverage reports and tracked in known gaps.

Primary sources for the next validation block:

Source Useful evidence Applicability constraints
CatS/AEP processing of MBP Product mapping, epitope generation/destruction and protection by HLA-DR Purified/experimental processing conditions; not a general peptide survival model
DIPPS specificity profiling Cathepsin profiles and pH-dependent legumain specificity Denatured in-gel substrates, condition-specific preferences; cannot turn sequence logos into calibrated probabilities
Legumain activation Activation and substrate-dependent pH response Both activation state and P1 residue affect the result; unconditional cut-after-Asn rules omit this context
Cathepsin MSP-MS comparison Human B/L/S and other cathepsins on 228 14-mers at pH 4.6/7.2, 15/60 minutes, 25 C Distinct CatB dipeptidyl-carboxypeptidase and endopeptidase activity; source-specific product detection thresholds

The MSP-MS publication deposits mass-spectrometry data as MSV000090043 / PXD035641. Its workflow uses quadruplicates and significance/fold-change criteria for detected products. Missing products are not automatically non-cleaved bonds. The available source supports assay-aware curation; it does not supply an already validated predictor of long-vaccine processing. The CatL example now includes seven experimentally identified products from two protected synthetic peptides in Tusar et al. 2023, Supplementary Data 4 page 1. The five observed internal substrate/bond pairs agree with Supplementary Table 17. Each original product, terminal modification, source measurement ID and assay condition survives batch save/reload. The assay is purified CatL at pH 5.5, 37 C for 2 hours; changing sequence or chemical form causes abstention. This is a selected source panel, not a transferable model or validation of whole-cell processing.

The batch reference_panels input can preserve curated experiments from these studies now, with each condition in its own named panel. Actual novel-sequence inference needs a separately verified adapter/data model. The model audit below records why novel-sequence coverage remains open.

Open-model availability audit (as of 2026-09-30)

Only openly licensed models are eligible for this work. Publicly downloadable files without a project license do not satisfy that constraint.

Candidate Verified artifact availability Remaining requirement
Tusar et al. cathepsin S/L/B SVMs CC BY 4.0 Supplementary Data 3 supplies six SVM-light files; B/L/S each declare 192 input features Reproduce the original sequence/structure features and independent reference scores before adapting them to vaccine inputs
PCSS backend LGPL-2.1 code at ea4c3ec81ef30b7a30f3c03508ee2bf1dd78ce34 Author README explicitly reports that hard-coded databases/programs prevent operation outside the Sali lab; weights alone do not supply this pipeline
ProsperousPlus Directories C01.060 (CatB), C01.032 (CatL), C01.034 (CatS), C13.004 (animal legumain) are present GitHub license metadata is null and the root has no project license; excluded under the open-model requirement (known gaps)
panCleave Author describes a pooled, protease-agnostic random forest Its output cannot supply enzyme-specific CatS/L/B/AEP coverage
DIPPS legumain study Experimental pH-dependent specificity evidence It is not a published fitted AEP predictor; an unconditional cut-after-Asn rule would discard the reported context

The CatB/L/S files were extracted from the Europe PMC open-access supplement archive for PMC10124925 and inspected. Their SHA-256 digests are:

  • SuppData3_CatB_SVMmodel.txt: 3acfa1cf903657694a6f5d68ba5151e6ea1ca191b2d7c212d0723d9fa1b5caac
  • SuppData3_CatL_SVMmodel.txt: dc013a06685c727384ec64cf9e0dc6166cf514fb5e849f66e0a07872cfba5951
  • SuppData3_CatS_SVMmodel.txt: 6c2c857c37f56d916c8706f43d008cefe6b47ee2feea43a35d66225fca6df60b

The article describes protein secondary-structure and solvent-exposure inputs. No unverified feature values, replacement model, or independent accuracy claim are supplied here. Completing this open-model runtime and obtaining an openly licensed, verified AEP predictor remain concrete blockers (known gaps). Requests for novel-sequence cathepsin/AEP predictions must continue to report unsupported coverage. Experimental source imports remain available now.

Source-backed long-peptide case study

Wada et al. 2018 is explicitly assigned to validation in Pepsickle's Table 1. The fixture curates all 47 detected products from Figure 2A and 2C: two 31-residue vaccine constructs containing the same epitopes in different orders, joined by RR linkers. The dataset retains first-detection times, figure row IDs, parent endpoints and the purified murine-i20S assay conditions.

The real gradient-boosted immunoproteasome model scores the complete constructs. Its native scores are paired with 24 observed internal construct/bond pairs; parent sequence ends never become cleavage labels. The pinned raw training-map audit covers 79 files and 58 distinct source sequences. It finds no exact construct or observed seven-residue cleavage-context overlap. Wada's DOI is absent from the reported training DOI fields; unresolved source identifiers, homology-family overlap and the original fitted partition remain explicit limitations. This is a small study-held-out case study, not a new calibrated performance estimate or a reconstruction of all 225 author validation windows.

The fixture README gives the primary source, license, curation scope and reproducible command. The report preserves observed products through batch save/reload, with endpoint/boundary/internal overlays. No missing product is relabeled as a negative, and no product detection time is converted into a predicted cleavage rate or presentation outcome.

Additional extracellular coverage is tracked in known gaps: CleaveNet's MMP substrate scores need their own native endpoint and assay validation rather than conversion to per-bond probabilities.

Required validation before biological ranking

Use exact source substrates/products, enzymes, species, chemical forms, activation, pH, temperature, time and detection limits. Separate positive product observations from verified non-cleavage measurements. Audit training overlap by study and sequence/family before claiming held-out performance. Preserve unsupported inputs and failures in the denominator.

Evaluate purified-enzyme cleavage, whole-matrix degradation and antigen presentation as distinct endpoints. In particular, human-DC long-peptide cross-presentation supports a possible proteasome/TAP route; it does not calibrate every construct or imply that cleavage scores alone predict presentation. IFN-gamma tumor-organoid experiments likewise show why a single sequence-only profile cannot encode the tumor's cellular state.