Neural training semantics
MHCflurry 2.3.0 uses PyTorch while retaining support for historical model weights. Matching a hyperparameter name does not always match the training equation across frameworks. This reference describes those distinctions; Release training recipe is the authoritative recipe for the released weights. See Evaluation of the 2.3.0 weights for the completed comparison of the released ensembles on identical evaluation rows.
Framework-semantic discrepancies found
Discrepancy |
Release relevance |
Resolution / justification |
|---|---|---|
Processing layers silently used PyTorch Kaiming/fan-in initialization and nonzero random biases |
Direct: every processing candidate |
Made explicit and equation-tested. Initialization is explicit in every generated recipe; see the final recipe for each processing family. |
Affinity LSUV observed raw Linear output instead of the activated Dense output |
Direct: every affinity candidate uses LSUV |
Both targets are explicit. Historical parity uses post-activation; the 2.3.0 affinity recipe selects pre-activation. |
Native PyTorch RMSprop places epsilon outside the square root; Keras places it inside |
Direct: every affinity optimizer step |
Fixed with a public, tested Keras equation and an explicit implementation switch. PyTorch documents this framework difference. |
Native PyTorch and Keras Adam place epsilon differently relative to bias correction |
Direct: every processing optimizer step |
Both equations are public and tested, and the selected implementation is serialized. The final no-flank and boundary families use Keras-compatible Adam; the short-flank family uses native RMSprop. |
Both PyTorch trainers rounded validation rows as |
Direct but usually one row per network |
Fixed centrally and regression-tested. This is deterministic parity, not an optimization question. |
PyTorch BatchNorm updates running variance with an unbiased estimate; Keras uses population variance |
Inactive in the release affinity grid; processing has no BN |
Fixed in |
Generic PyTorch Xavier fan calculation was wrong for the transposed 3-D |
Inactive: release affinity grid has no local layers |
Fixed using Keras’ |
Native SGD fallback used LR 0.001 and unknown optimizer names silently became Adam |
Inactive: release uses RMSprop/Adam |
Fixed: Keras SGD default LR is 0.01 and unknown names now fail. |
Keras Glorot/He normal initializers use variance-corrected truncated normals; PyTorch’s normal initializers are untruncated |
Inactive: release uses Glorot uniform |
Follow-up only if those non-release initializer values are to remain supported for new training. Loaded historical weights are unaffected. |
TensorFlow and PyTorch dropout RNGs, shuffles, reduction order, and GPU kernels differ |
Direct stochastic trajectory, but not a hyperparameter drift |
Irreducible framework difference. Compare distributions and held-out metrics, not byte-identical weights. |
Fixed master/per-fit seeds replace entropy-derived seeds |
Direct identities/trajectory |
Intentional reproducibility improvement. It does not change the sampled distributions. |
Device-side encoding, compact cartesian batches, validation batching, lazy proteome sampling, and prediction chunking |
Execution only |
Algebra/prediction parity is covered by tests. Release training fails if autosizing would shrink a configured minibatch. |
|
Could alter trajectory |
Off for release; eager + |
Native RMSprop adds epsilon outside the square root; Keras-compatible RMSprop adds it inside. Adam implementations also differ in epsilon placement relative to bias correction. These are explicit optimizer choices, not transparent performance switches. See the PyTorch RMSprop documentation, Keras RMSprop documentation, and Adam paper.
The public class defaults, compatibility generators, and frozen release recipe serve different purposes. The compatibility affinity generator defaults to minibatch 128; the 2.3.0 release recipe explicitly selects 1024 with native RMSprop and pre-activation LSUV. Always retain the generated hyperparameter YAML with a training run. Framework RNGs and reduction order preclude byte-identical retraining; validate prediction quality on a common held-out cohort.