Processing data preparation

mhcflurry train processing-data prepares affinity-matched negative peptides with bounded CPU/scoring overlap and resumable, checksummed artifacts. The matching contract is described in Processing negative matching.

Execution and saved artifacts

Each sample is a resumable workflow. CPU stages sample candidate windows and save scored rounds while a single scoring owner per process runs the affinity predictor. The default pipeline depth is three; use --preparation-pipeline-depth 1 for serial stages. Sample seeds derive from the master seed and sample identity, so scheduling does not change candidate draws.

ProcessingProteome encodes proteins once per worker. ProteinWindows keeps sampled positions and peptide identities numeric; NumericCandidatePool deduplicates scoring inputs before peptide/flank strings are materialized for saved tables. NumericSequences.aligned_tensor() and Class1AffinityPredictor.predict_numeric(..., allele=...) gather and align inputs on CPU, MPS or CUDA. This strict, single-allele API uses the existing networks and log-affinity ensemble aggregation. Sampling remains on the CPU.

Each scored round is saved before matching or expansion. An insufficient pool triggers additional draws for unresolved peptide lengths, within the configured round limit. Completion is recorded only after matched outputs and their hashes are saved. Errors retain committed rounds; they do not publish a completed sample.

Use --resume for the same output or --resume-matching-dir PRIOR.matching when creating a new output. Resume checks input/reference hashes, matching policy, seed, observed rows and saved scores. It never silently reuses old assignments from a different matching policy. The explicit legacy-top-binders recipe uses the historical serial preparation path.

Benchmarking

mhcflurry train benchmark-processing-preparation \
  --out benchmark.json --repeats 5 \
  --scored-pool sample.round-00.csv.bz2 \
  --scored-pool sample.round-01.csv.bz2

Supply all rounds needed for a feasible matching pool. Add --affinity-predictor /path/to/models.combined to compare actual numeric/string predictions and timings. Outputs record input and source hashes, runtime versions, all timing repetitions and parity checks. Sampling inputs are synthetic; supplied scored pools can be real. Inference timings exclude loading and initial staging.

--baseline-matching-source /path/to/processing_matching.py compares two implementations of the same matching policy and requires exact assignment and diagnostic equality. A nearest-neighbor matcher that reuses negatives is not an appropriate parity baseline for the current without-replacement policy.

Component timings do not establish an end-to-end preparation speedup or a gain in trained-model accuracy. The tests separately cover pipeline backpressure and errors, resume integrity, matching invariants, and numeric/string prediction parity.