Auditing training-sample overlap

Zero peptide–MHC overlap does not establish biological-sample separation: other peptides from the same specimen may have contributed to training or model selection. The historical affinity bundles lack enough sample provenance to certify this separation. Upgrading the code cannot reconstruct those identities. External predictors’ training overlap also remains uncertified without their complete lineage. Existing benchmark results retain these limitations.

Preserve lineage during new training

New affinity curation adds source_provenance, a JSON array of source records. Each record preserves available study, sample and assay IDs, the raw input file’s SHA-256, and its zero-based row index after parsing the CSV header. Deduplication unions every contributing record while retaining the previous measurement, row order and weighting. Reassignment, training metadata and selection metadata retain this column. Archive the original sources as well: hashes and row indices identify evidence; they do not replace it.

Missing IDs remain unknown. An assay ID is not a biological sample ID. complete on a source record means contributor retention, not known specimen identity. Historical deduplicated tables must not be relabeled as complete sample tables merely because one study or sample can be recovered.

New release-holdout policies also write affinity_source_samples.csv, containing the union of monoallelic and selected presentation evaluation samples. The affinity preparation script applies it in addition to the pMHC exclusions. Direct use is:

mhcflurry class1-reassign-mass-spec-training-data curated.csv \
    --exclude-pmhcs holdout/affinity_pmhcs.csv \
    --exclude-source-samples holdout/affinity_source_samples.csv \
    --sample-aliases specimen_aliases.csv --out-csv training.csv

Any overlapping contributor removes the entire measurement. If training has a known study but no sample identity, the whole matching study is excluded. An exclusion lacking a study conservatively matches its sample ID in every study. Unresolvable training rows remain explicitly unresolved; retaining them cannot support a verified-disjoint claim. The ordinary release-holdout validate command checks the recorded exclusions and labels whole-sample proof unresolved. Old policy files remain readable with their original, narrower scope.

Audit every compared model together

Freeze one cohort and inventory every compared MHCflurry model, including affinity, processing and presentation components. Inventory all biological sources used for pretraining, training, stopping/development and ensemble selection, including the lineage of any teacher used to generate synthetic pretraining targets. An empty stage list explicitly means that stage used no data; it does not mean its data are unavailable.

Run the audit before claiming sample-disjoint evaluation:

mhcflurry train release-holdout audit-samples \
    --inventory lineage.json --cohort frozen_cohort.csv.gz \
    --sample-metadata cohort_studies.csv --aliases specimen_aliases.csv \
    --out-dir results/sample_audit

The output directory must be new. The command writes sample_audit.csv, sample_disjointness.json and cohort.csv.gz, with hashes, sample/row/label counts, and reasons for overlap or uncertainty. It fails if any input sample is overlapping or unresolved against any inventoried MHCflurry model. Use --report-only to inspect incomplete historical evidence without that exit failure. External evidence has a separate status and does not silently become verified when MHCflurry passes.

The JSON also records the generating function and arguments, package/source hashes and Python/pandas versions. The audit is deterministic and uses no random seed. Freeze the inventory and source files while it runs.

The exported cohort contains only samples verified against all inventoried MHCflurry models. Use that one cohort for every comparator, retaining all its positive and negative rows. Disclose changed counts and hashes; never filter each predictor independently. Recheck peptide–MHC overlap separately. A benchmark already used for development remains a regression benchmark, even after overlap filtering; use an untouched confirmation cohort for new generalization claims.

Inventory and identity formats

This example describes one affinity-only model. A full presentation model uses "kind": "presentation" and requires affinity, processing and presentation component entries. Add every compared release as a separate model. External models use "kind": "external"; unavailable components/artifacts yield an explicit unresolved result.

{
  "schema_version": 1,
  "identity_review": "Describe the reviewed study namespaces, specimen aliases and shared specimens here.",
  "models": [{
    "name": "candidate-affinity",
    "kind": "affinity",
    "artifacts": ["affinity-model-bundle.tar.bz2"],
    "components": {
      "affinity": {
        "complete": true,
        "evidence": "Describe how these files cover this model's complete lineage.",
        "pretraining": [],
        "training": [{"path": "train_data.csv.bz2"}],
        "development": [],
        "selection": [{"path": "model_selection_data.csv.bz2"}]
      }
    }
  }]
}

Paths are relative to the inventory file. Include an immutable model archive or all model files in artifacts; their hashes bind the report to those weights. The tool verifies the supplied files, not the truth of a completeness declaration. Supply review evidence and leave complete: false when any lineage is missing.

Data entries default to the source_provenance column. For reviewed raw sample tables, an entry may instead specify "identity_mode": "sample_table", with "study_column": "pmid" and "sample_column": "sample_id" (defaults: study_id and sample_id). This mode is unsuitable for historical affinity tables that already discarded contributors.

cohort_studies.csv has study_id,sample_id columns; it is optional if the cohort already has both. Numeric study IDs mean PubMed IDs and normalize to pmid:ID. Other study IDs need explicit namespaces and reviewed correspondence. Sample IDs retain leading zeroes. Alias CSV columns, in order, are study_id,sample_id,canonical_study_id,canonical_sample_id. Blank sample IDs on both sides map a study alias; filled IDs map a shared specimen, including across studies. Conflicting mappings and cycles are errors. Different spelling or a different publication is not sufficient evidence of a different specimen.