Processing hyperparameter campaign
The maintained sweep compares processing-specific recipes; it does not assume that affinity’s winning optimizer transfers. All new fits use a frozen affinity/length-matched table and preserve fold and reference identities. Development results and the repeatedly inspected release benchmark are distinct.
Maintained commands
Use mhcflurry train processing-hyperparameter-sweep --design kernel-width
for the six-width, two-family, four-fold comparison (48 networks). The existing
processing-kernel-sweep command remains an alias for the width workflow.
Use mhcflurry train processing-hyperparameter-sweep --design training-recipe
for the optimizer/initialization/batch factorial (64 networks):
Keras-compatible Adam versus native PyTorch RMSprop.
Glorot, orthogonal only, pre-LSUV and post-LSUV.
Minibatches 512/1024, two architecture families and two paired folds.
The recipe screen holds width 11, ReLU, 512 filters, dropout 0.5, learning rate 0.001 and L1/L2 zero fixed. The width screen retains its original L1=1e-6. Thus its width-11 control is not silently reused as a zero-L1 recipe fit. The families are an external-flank CNN and a whole-peptide CNN plus 5x5 boundary branches, not a fixed mixture of networks.
Required arguments include --out, --train-data, --public-root,
--release-holdout-dir and --source-commit. Use --evaluation none for
development screening. --prepare-only writes the exact design without training.
--folds-from inherits the first two folds for the recipe screen; pass the same
frozen table as --train-data to preserve all numerical metadata exactly.
Checkpoints and initialization
save_all_checkpoints=True retains independent best-loss and terminal processing
weights, plus best-AP when ranking monitoring is enabled.
restore_best_weights=True chooses the best state for checkpoint_metric
(val_loss by default; opt-in val_macro_ap). Checkpoints use NPZ sidecars; they do not enter manifest
JSON as weight arrays. Loading a referenced missing checkpoint fails. Refitting
without retention clears old references. Older predictors still load without
checkpoint sidecars, and missing historical terminal states are not synthesized.
monitor_validation_ranking=True records equal-sample AP and PPV@N after every
epoch. It requires explicit sample-disjoint training/stopping-validation masks
and sample IDs. Each validation sample must contain both labels. AP uses grouped
score ties; PPV breaks ties by original input-row order. Exact AP ties retain
the earliest epoch. Inner-best AP is saved as best_ap, separately from best
(loss) and terminal. checkpoint_metric=val_macro_ap enables this monitoring
and chooses that state when restoration is enabled. Patience remains based on
validation loss, so retained-state comparisons share one training trajectory.
Historical defaults and predictions remain unchanged. Fit metadata explicitly
records the restored policy, best loss/ranking epochs, and inner sample identity.
Focused width / optimizer / checkpoint confirmation
Use mhcflurry train processing-hyperparameter-sweep --design ranking-confirmation
with --folds-from FROZEN_TABLE --train-data FROZEN_TABLE --evaluation none.
This creates 24 fits: legacy 5-aa CNN, widths 11/13/15, Adam/Keras versus
RMSprop/PyTorch and four paired folds. Glorot, batch 512, ReLU, 512 filters,
dropout 0.5, no normalization, fixed BLOSUM62 and zero L1/L2 are held constant.
All three checkpoint states have joinable outer-fold predictions and separate
per-sample and pooled-per-fold metric tables. Primary inference uses inner-best
AP; outer evaluation cannot select epochs. The source generator exposes
processing_candidate_hyperparameters for this explicit shortlist; it is not
a global default change or a final release recipe.
Set PROCESSING_RANKING_CONFIRMATION=1 in the runplz Modal launcher, together
with the frozen table, absolute deadline and timeout at most four hours.
This mode skips both the width and full recipe factorials. It is incompatible
with PROCESSING_RECIPE_AFTER_WIDTHS=1. Training tables and independent
checkpoint files survive completed conditions; the budget is never extended
by retries. New experiments require compact selection and a separate full-presentation
evaluation before release, regardless of development improvements.
Analyze collected condition-level checkpoint tables with
mhcflurry eval processing-confirmation-analysis --experiment RUN_DIR --out NEW_DIR.
The command checks identical fold/sample/count identities across every state
and condition, writes paired sample-bootstrap intervals and figures, and only
exports candidate_hyperparameters.yaml after all six conditions are complete
and the declared development gate passes. The gate compares inner-best-AP
candidates against Adam/width-11 best-loss control: positive AP and PPV changes,
95% paired AP lower bound above zero, PPV lower bound above -0.002, and every
fold’s pooled AUROC/AP/PPV changes at least -0.002. These are exploratory
intervals, not corrected for repeated screening or study dependence. A recipe
export is not a model release or a substitute for presentation validation.
initialization_method has explicit none, orthogonal, lsuv_pre and
lsuv_post values. none preserves the configured kernel initialization. Non-default methods
apply to fresh fits; continued fitting does not reinitialize a trained network.
Calibration uses at most initialization_batch_size training rows (default
512), never stopping-validation rows. Their fit-input row indices are saved.
Eligible layers are boundary hidden layers, the main convolution, and hidden pointwise convolutional heads, in forward order. Scalar heads and output/gating weights are protected. Pre/post refer to before/after the configured activation, before normalization or dropout. Variance excludes positions beyond actual configured input extent; missing-context X rows within that extent remain. Dropout is disabled during initialization. Non-convergence fails explicitly and restores the pre-initialization parameters. Diagnostics include each layer’s variance and iteration count. See the LSUV paper.
Recorded development result
The complete panel (processing-ranking-confirmation-20260909; runplz run
a8889cc8dfbb4f4390d4ceb153ad81e1; source 9778c625; 24 fits in 65 minutes
on one A100-40GB) promoted
legacy_5aa__rmsprop_pytorch__k13 with inner-best-AP weights.
confirmed_processing_candidate_hyperparameters() in
scripts/training/generate_processing_recipe.py names that recipe together
with the run, source and decision hashes; it equals the exported
candidate_hyperparameters.yaml. Macro metrics average four paired folds
within each of 37 held-out samples. Differences are against Adam/Keras width
11 best-loss weights (macro AUROC 0.7717, AUPRC 0.7555, PPV@N 0.7001), using
10000 paired sample-bootstrap draws with seed 42.
Inner-best-AP condition |
AUROC |
AUPRC |
PPV@N |
AUPRC difference |
PPV@N difference |
Gate |
|---|---|---|---|---|---|---|
Adam/Keras, width 11 |
0.7810 |
0.7641 |
0.7128 |
+0.0086 [0.0043, 0.0126] |
+0.0127 [0.0085, 0.0169] |
pass |
Adam/Keras, width 13 |
0.7857 |
0.7671 |
0.7162 |
+0.0116 [0.0048, 0.0182] |
+0.0162 [0.0105, 0.0219] |
pass |
Adam/Keras, width 15 |
0.7838 |
0.7666 |
0.7153 |
+0.0111 [0.0043, 0.0176] |
+0.0153 [0.0085, 0.0216] |
pass |
RMSprop/PyTorch, width 11 |
0.7798 |
0.7629 |
0.7090 |
+0.0074 [0.0021, 0.0128] |
+0.0089 [0.0035, 0.0144] |
fail: pooled fold AUROC -0.0039, AUPRC -0.0074 |
RMSprop/PyTorch, width 13 |
0.7870 |
0.7701 |
0.7155 |
+0.0146 [0.0085, 0.0207] |
+0.0154 [0.0101, 0.0206] |
pass, promoted |
RMSprop/PyTorch, width 15 |
0.7869 |
0.7688 |
0.7146 |
+0.0133 [0.0076, 0.0191] |
+0.0145 [0.0088, 0.0203] |
pass |
The retained-state choice is the largest single effect. In all six conditions inner-best-AP and terminal weights outscore best-loss weights on AUPRC and PPV@N; best-loss epochs ranged from 4 to 34 and inner-best-AP epochs from 14 to 44. Within the control, switching states alone gains +0.0086 AUPRC. Terminal weights scored similarly but were never eligible, by design. Widths 13 and 15 lead width 11 with either optimizer, but the promoted recipe’s AUPRC margin over the other passing conditions (0.001 to 0.006) lies inside every paired interval; it is the recorded point-estimate tie-break, not a resolved difference.
These are single-fit, four-fold development scores on one training trajectory per fold, without multiple-screening or study-cluster correction. They do not change public downloads, default hyperparameters or predictions, and they say nothing about ensembles or presentation. Compact selection on a common held-out panel and the separate presentation gate remain required before any release use; family and width diversity is preserved for that step.
Recovery and outputs
--resume-from PRIOR_WIDTH_RUN copies complete conditions and frozen folds to
a separate run after checking input/design identity. Incomplete conditions stay
in the old run and are reinitialized in the new one. Every copied file is hashed.
Existing recovery artifacts must match those hashes on resume. A lossless CSV
parser preserves saved reference scores; strict fold checks are not relaxed.
Every new condition saves epoch losses, optimizer steps, epoch timing, stop reason, initialization diagnostics, best/terminal weights, and context-joinable per-member predictions on each model’s held-out fold. Root summaries give each sample equal weight after averaging repeated folds within sample. These one-decoy validation AP values are not comparable to ten-decoy processing or full-presentation AP levels. Four/eight-member selection still requires a common held-out panel unseen by every member; this screen does not perform it.
The runplz Modal launcher accepts PROCESSING_KERNEL_TRAIN_DATA and
PROCESSING_KERNEL_PRIOR_SWEEP to reuse completed preparation and fits.
PROCESSING_KERNEL_EVALUATION=none is the default.
PROCESSING_RECIPE_AFTER_WIDTHS=1 queues the factorial behind width completion
on the same single GPU. It requires an absolute
MHCFLURRY_EXPERIMENT_DEADLINE_EPOCH; command descendants are terminated at
that deadline, without marking the stage complete. Set an absolute deadline
within the allocated compute budget, leaving margin for teardown. The launcher
caps each function timeout at 15.5 hours; retries must respect the same deadline.
The launcher writes width figures first, then a combined
campaign-all-figures.pdf after the recipe screen. Raw inputs and prediction
caches remain independently collectible even if a later stage fails. A generated
PDF is not a substitute for inspected figures, paired uncertainty, four-fold
confirmation, compact selection and an end-to-end presentation gate.
Analyze an incomplete or completed recipe screen
Run the maintained read-only analysis against a collected snapshot from one run:
mhcflurry eval processing-recipe-analysis \
--experiment collected/recipe_sweep/experiment.json \
--metrics collected/recipe_sweep/validation_per_sample.csv \
--out experiments/recipe-analysis-snapshot
The output directory must be new or empty. The command preserves unreported conditions as pending and rejects incomplete folds, mismatched sample cohorts or counts, nonfinite metrics, and uncontrolled hyperparameter differences. It averages folds within each sample, then bootstraps samples jointly for each one-factor contrast. It does not average across unmatched factorial cells to claim an overall optimizer or initialization effect. Optimizer comparisons include their implementation (for example, Adam/Keras versus RMSprop/PyTorch).
Outputs include a PDF and page PNGs, condition means, design completion status, paired contrasts, per-sample differences, and a provenance manifest with input and analysis-source hashes. Gray plot cells mean pending, not poor performance. Pareto flags describe AUPRC/PPV point estimates only; they are not significance tests, ensemble selections, or release acceptance. Intervals are exploratory, conditional on trained fits, and uncorrected for screening or study clustering.
The shared mhcflurry train plot-loss-curves command recognizes processing CNN
architecture identities and convolutional L1/L2 regularization. Its curves and
final_loss/final_val_loss columns describe final training epochs, even when
inference restores an earlier checkpoint. Separate best-epoch, best-validation-
loss, and checkpoint-policy columns record that distinction. Historical fits
without explicit restoration metadata report an unknown checkpoint policy.
Cached within-fold ensemble diagnostics
mhcflurry eval processing-fold-ensembles --ensembles ensembles.json --out NEW_DIR
accepts a JSON object mapping ensemble names to ordered checkpoint prediction
cache paths (relative to that JSON). It verifies each cache checksum/sidecar,
training-table fingerprint, sample/row/context/label identity, and one distinct
member per condition per fold. Every mean uses only the models from the row’s
held-out fold. Best and terminal are reported separately by default; use
--checkpoint-policy best to request only the primary policy.
Outputs preserve averaged predictions, fold-to-member identities, per-sample
and pooled-per-fold metrics, sample means, provenance and source. Pass the
sample means to mhcflurry eval paired-sample-metrics for paired plots. This
is a no-training, no-inference development diagnostic of multiple fold-specific
ensembles. It does not measure one fixed ensemble spanning different folds,
select members, or establish release acceptance. A final fixed ensemble still
requires a common panel unseen by every one of its members.