Processing hyperparameter campaign

The maintained sweep compares processing-specific recipes; it does not assume that affinity’s winning optimizer transfers. All new fits use a frozen affinity/length-matched table and preserve fold and reference identities. Development results and the repeatedly inspected release benchmark are distinct.

Maintained commands

Use mhcflurry train processing-hyperparameter-sweep --design kernel-width for the six-width, two-family, four-fold comparison (48 networks). The existing processing-kernel-sweep command remains an alias for the width workflow.

Use mhcflurry train processing-hyperparameter-sweep --design training-recipe for the optimizer/initialization/batch factorial (64 networks):

  • Keras-compatible Adam versus native PyTorch RMSprop.

  • Glorot, orthogonal only, pre-LSUV and post-LSUV.

  • Minibatches 512/1024, two architecture families and two paired folds.

The recipe screen holds width 11, ReLU, 512 filters, dropout 0.5, learning rate 0.001 and L1/L2 zero fixed. The width screen retains its original L1=1e-6. Thus its width-11 control is not silently reused as a zero-L1 recipe fit. The families are an external-flank CNN and a whole-peptide CNN plus 5x5 boundary branches, not a fixed mixture of networks.

Required arguments include --out, --train-data, --public-root, --release-holdout-dir and --source-commit. Use --evaluation none for development screening. --prepare-only writes the exact design without training. --folds-from inherits the first two folds for the recipe screen; pass the same frozen table as --train-data to preserve all numerical metadata exactly.

Checkpoints and initialization

save_all_checkpoints=True retains independent best-loss and terminal processing weights, plus best-AP when ranking monitoring is enabled. restore_best_weights=True chooses the best state for checkpoint_metric (val_loss by default; opt-in val_macro_ap). Checkpoints use NPZ sidecars; they do not enter manifest JSON as weight arrays. Loading a referenced missing checkpoint fails. Refitting without retention clears old references. Older predictors still load without checkpoint sidecars, and missing historical terminal states are not synthesized.

monitor_validation_ranking=True records equal-sample AP and PPV@N after every epoch. It requires explicit sample-disjoint training/stopping-validation masks and sample IDs. Each validation sample must contain both labels. AP uses grouped score ties; PPV breaks ties by original input-row order. Exact AP ties retain the earliest epoch. Inner-best AP is saved as best_ap, separately from best (loss) and terminal. checkpoint_metric=val_macro_ap enables this monitoring and chooses that state when restoration is enabled. Patience remains based on validation loss, so retained-state comparisons share one training trajectory. Historical defaults and predictions remain unchanged. Fit metadata explicitly records the restored policy, best loss/ranking epochs, and inner sample identity.

Focused width / optimizer / checkpoint confirmation

Use mhcflurry train processing-hyperparameter-sweep --design ranking-confirmation with --folds-from FROZEN_TABLE --train-data FROZEN_TABLE --evaluation none. This creates 24 fits: legacy 5-aa CNN, widths 11/13/15, Adam/Keras versus RMSprop/PyTorch and four paired folds. Glorot, batch 512, ReLU, 512 filters, dropout 0.5, no normalization, fixed BLOSUM62 and zero L1/L2 are held constant. All three checkpoint states have joinable outer-fold predictions and separate per-sample and pooled-per-fold metric tables. Primary inference uses inner-best AP; outer evaluation cannot select epochs. The source generator exposes processing_candidate_hyperparameters for this explicit shortlist; it is not a global default change or a final release recipe.

Set PROCESSING_RANKING_CONFIRMATION=1 in the runplz Modal launcher, together with the frozen table, absolute deadline and timeout at most four hours. This mode skips both the width and full recipe factorials. It is incompatible with PROCESSING_RECIPE_AFTER_WIDTHS=1. Training tables and independent checkpoint files survive completed conditions; the budget is never extended by retries. New experiments require compact selection and a separate full-presentation evaluation before release, regardless of development improvements.

Analyze collected condition-level checkpoint tables with mhcflurry eval processing-confirmation-analysis --experiment RUN_DIR --out NEW_DIR. The command checks identical fold/sample/count identities across every state and condition, writes paired sample-bootstrap intervals and figures, and only exports candidate_hyperparameters.yaml after all six conditions are complete and the declared development gate passes. The gate compares inner-best-AP candidates against Adam/width-11 best-loss control: positive AP and PPV changes, 95% paired AP lower bound above zero, PPV lower bound above -0.002, and every fold’s pooled AUROC/AP/PPV changes at least -0.002. These are exploratory intervals, not corrected for repeated screening or study dependence. A recipe export is not a model release or a substitute for presentation validation.

initialization_method has explicit none, orthogonal, lsuv_pre and lsuv_post values. none preserves the configured kernel initialization. Non-default methods apply to fresh fits; continued fitting does not reinitialize a trained network. Calibration uses at most initialization_batch_size training rows (default 512), never stopping-validation rows. Their fit-input row indices are saved.

Eligible layers are boundary hidden layers, the main convolution, and hidden pointwise convolutional heads, in forward order. Scalar heads and output/gating weights are protected. Pre/post refer to before/after the configured activation, before normalization or dropout. Variance excludes positions beyond actual configured input extent; missing-context X rows within that extent remain. Dropout is disabled during initialization. Non-convergence fails explicitly and restores the pre-initialization parameters. Diagnostics include each layer’s variance and iteration count. See the LSUV paper.

Recorded development result

The complete panel (processing-ranking-confirmation-20260909; runplz run a8889cc8dfbb4f4390d4ceb153ad81e1; source 9778c625; 24 fits in 65 minutes on one A100-40GB) promoted legacy_5aa__rmsprop_pytorch__k13 with inner-best-AP weights. confirmed_processing_candidate_hyperparameters() in scripts/training/generate_processing_recipe.py names that recipe together with the run, source and decision hashes; it equals the exported candidate_hyperparameters.yaml. Macro metrics average four paired folds within each of 37 held-out samples. Differences are against Adam/Keras width 11 best-loss weights (macro AUROC 0.7717, AUPRC 0.7555, PPV@N 0.7001), using 10000 paired sample-bootstrap draws with seed 42.

Inner-best-AP condition

AUROC

AUPRC

PPV@N

AUPRC difference

PPV@N difference

Gate

Adam/Keras, width 11

0.7810

0.7641

0.7128

+0.0086 [0.0043, 0.0126]

+0.0127 [0.0085, 0.0169]

pass

Adam/Keras, width 13

0.7857

0.7671

0.7162

+0.0116 [0.0048, 0.0182]

+0.0162 [0.0105, 0.0219]

pass

Adam/Keras, width 15

0.7838

0.7666

0.7153

+0.0111 [0.0043, 0.0176]

+0.0153 [0.0085, 0.0216]

pass

RMSprop/PyTorch, width 11

0.7798

0.7629

0.7090

+0.0074 [0.0021, 0.0128]

+0.0089 [0.0035, 0.0144]

fail: pooled fold AUROC -0.0039, AUPRC -0.0074

RMSprop/PyTorch, width 13

0.7870

0.7701

0.7155

+0.0146 [0.0085, 0.0207]

+0.0154 [0.0101, 0.0206]

pass, promoted

RMSprop/PyTorch, width 15

0.7869

0.7688

0.7146

+0.0133 [0.0076, 0.0191]

+0.0145 [0.0088, 0.0203]

pass

The retained-state choice is the largest single effect. In all six conditions inner-best-AP and terminal weights outscore best-loss weights on AUPRC and PPV@N; best-loss epochs ranged from 4 to 34 and inner-best-AP epochs from 14 to 44. Within the control, switching states alone gains +0.0086 AUPRC. Terminal weights scored similarly but were never eligible, by design. Widths 13 and 15 lead width 11 with either optimizer, but the promoted recipe’s AUPRC margin over the other passing conditions (0.001 to 0.006) lies inside every paired interval; it is the recorded point-estimate tie-break, not a resolved difference.

These are single-fit, four-fold development scores on one training trajectory per fold, without multiple-screening or study-cluster correction. They do not change public downloads, default hyperparameters or predictions, and they say nothing about ensembles or presentation. Compact selection on a common held-out panel and the separate presentation gate remain required before any release use; family and width diversity is preserved for that step.

Recovery and outputs

--resume-from PRIOR_WIDTH_RUN copies complete conditions and frozen folds to a separate run after checking input/design identity. Incomplete conditions stay in the old run and are reinitialized in the new one. Every copied file is hashed. Existing recovery artifacts must match those hashes on resume. A lossless CSV parser preserves saved reference scores; strict fold checks are not relaxed.

Every new condition saves epoch losses, optimizer steps, epoch timing, stop reason, initialization diagnostics, best/terminal weights, and context-joinable per-member predictions on each model’s held-out fold. Root summaries give each sample equal weight after averaging repeated folds within sample. These one-decoy validation AP values are not comparable to ten-decoy processing or full-presentation AP levels. Four/eight-member selection still requires a common held-out panel unseen by every member; this screen does not perform it.

The runplz Modal launcher accepts PROCESSING_KERNEL_TRAIN_DATA and PROCESSING_KERNEL_PRIOR_SWEEP to reuse completed preparation and fits. PROCESSING_KERNEL_EVALUATION=none is the default. PROCESSING_RECIPE_AFTER_WIDTHS=1 queues the factorial behind width completion on the same single GPU. It requires an absolute MHCFLURRY_EXPERIMENT_DEADLINE_EPOCH; command descendants are terminated at that deadline, without marking the stage complete. Set an absolute deadline within the allocated compute budget, leaving margin for teardown. The launcher caps each function timeout at 15.5 hours; retries must respect the same deadline.

The launcher writes width figures first, then a combined campaign-all-figures.pdf after the recipe screen. Raw inputs and prediction caches remain independently collectible even if a later stage fails. A generated PDF is not a substitute for inspected figures, paired uncertainty, four-fold confirmation, compact selection and an end-to-end presentation gate.

Analyze an incomplete or completed recipe screen

Run the maintained read-only analysis against a collected snapshot from one run:

mhcflurry eval processing-recipe-analysis \
  --experiment collected/recipe_sweep/experiment.json \
  --metrics collected/recipe_sweep/validation_per_sample.csv \
  --out experiments/recipe-analysis-snapshot

The output directory must be new or empty. The command preserves unreported conditions as pending and rejects incomplete folds, mismatched sample cohorts or counts, nonfinite metrics, and uncontrolled hyperparameter differences. It averages folds within each sample, then bootstraps samples jointly for each one-factor contrast. It does not average across unmatched factorial cells to claim an overall optimizer or initialization effect. Optimizer comparisons include their implementation (for example, Adam/Keras versus RMSprop/PyTorch).

Outputs include a PDF and page PNGs, condition means, design completion status, paired contrasts, per-sample differences, and a provenance manifest with input and analysis-source hashes. Gray plot cells mean pending, not poor performance. Pareto flags describe AUPRC/PPV point estimates only; they are not significance tests, ensemble selections, or release acceptance. Intervals are exploratory, conditional on trained fits, and uncorrected for screening or study clustering.

The shared mhcflurry train plot-loss-curves command recognizes processing CNN architecture identities and convolutional L1/L2 regularization. Its curves and final_loss/final_val_loss columns describe final training epochs, even when inference restores an earlier checkpoint. Separate best-epoch, best-validation- loss, and checkpoint-policy columns record that distinction. Historical fits without explicit restoration metadata report an unknown checkpoint policy.

Cached within-fold ensemble diagnostics

mhcflurry eval processing-fold-ensembles --ensembles ensembles.json --out NEW_DIR accepts a JSON object mapping ensemble names to ordered checkpoint prediction cache paths (relative to that JSON). It verifies each cache checksum/sidecar, training-table fingerprint, sample/row/context/label identity, and one distinct member per condition per fold. Every mean uses only the models from the row’s held-out fold. Best and terminal are reported separately by default; use --checkpoint-policy best to request only the primary policy.

Outputs preserve averaged predictions, fold-to-member identities, per-sample and pooled-per-fold metrics, sample means, provenance and source. Pass the sample means to mhcflurry eval paired-sample-metrics for paired plots. This is a no-training, no-inference development diagnostic of multiple fold-specific ensembles. It does not measure one fixed ensemble spanning different folds, select members, or establish release acceptance. A final fixed ensemble still requires a common panel unseen by every one of its members.