Configuration and performance

Automatic hardware planning

Prediction, training, calibration, selection, and evaluation commands share a workload-aware planner. Depending on the command, it considers available GPU memory, model and data size, the configured training batch, host RAM, CPU count, and remaining work items. Container memory limits are detected through cgroups, so a job does not size itself from RAM that belongs to the physical host but is unavailable inside the container. On spawn-based platforms, the planner also measures the loaded worker context before starting the pool; this prevents a compressed input file from hiding the much larger in-memory copy created in each worker. If the remaining safe RAM cannot hold even one additional spawn copy, an automatic plan runs serially in the already-loaded parent process.

The common automatic options are:

Option

Automatic behavior

--gpus

With the auto or gpu backend, discovers every visible CUDA device unless a count is supplied.

--num-jobs auto

Chooses the total worker count from CUDA and host capacity; runs serially on CPU/MPS.

--max-workers-per-gpu auto

Packs complete estimated worker working sets into each GPU.

--dataloader-num-workers auto

Sizes pretraining DataLoader children from CPU and host memory.

--torch-compile auto

Reads MHCFLURRY_TORCH_COMPILE; compilation is off when unset. Enabling it affects CUDA only.

Automatic GPU packing has no default fixed worker cap. Explicit CLI values and explicit per-workload memory estimates remain authoritative; the planner logs the facts and limits used for each decision.

Affinity and processing trainers validate an automatic GPU plan with a bounded single-worker resource probe before starting the production pool. The probe uses the real resident fold and validation path and measures whole-process CUDA usage; it is a resource-safety step even when torch.compile is disabled. Runtime auto batches use a fixed per-worker share of launch-time free device capacity, so their size does not depend on which co-resident worker initialized first. Automatic CUDA processing batches are additionally capped by a successful real-model forward probe; the planner does not extrapolate across unobserved convolution workspace shapes.

If a workload still runs out of memory, first keep the batch and worker settings on auto. If you need to intervene, reduce --max-workers-per-gpu or use a smaller explicit batch. Pinning values can improve repeatability on a known machine, but it also bypasses some automatic safeguards.

See Orchestrator architecture for implementation details and expert environment overrides.

Environment overrides

These variables are intended for custom model locations, debugging, and controlled benchmarks.

MHCFLURRY_DEFAULT_CLASS1_MODELS

Path to the default affinity predictor. Without this override, Class1AffinityPredictor.load() uses the active release’s standalone affinity bundle when installed, then falls back to the affinity component of its presentation bundle. Explicit model paths take precedence.

MHCFLURRY_DEFAULT_CLASS1_PRESENTATION_MODELS_DIR

Path to the default presentation predictor, including its affinity and processing components. Otherwise, uses the active release’s presentation bundle.

MHCFLURRY_DEFAULT_CLASS1_PROCESSING_MODELS_DIR

Path to the default standalone processing predictor. Otherwise, uses the active release’s standalone processing bundle.

See Downloading and selecting model weights for release selection and download-directory overrides.

MHCFLURRY_OPTIMIZATION_LEVEL

Controls pan-allele ensemble merging. The default, 1, enables the faster merged-network path. Set it to 0 when debugging or comparing the unmerged implementation.

MHCFLURRY_DEFAULT_PREDICT_BATCH_SIZE

Overrides the default prediction batch size. The normal value is auto, which sizes batches from available device memory and can retry with a smaller batch after an allocator out-of-memory error. A fixed integer disables that resizing behavior and should be used only when you know the workload fits.

Reproducible training

Training and calibration commands accept --random-seed. The command-line default is 42, covering data splits and shuffles, initial weights, random negatives, allele sampling, and calibration peptides. Ensemble members derive distinct sub-seeds from that master seed.

Compact percentile knot selection has a separate, fixed seed of 403 for its grouped 80/20 background split. --random-seed controls background peptide and MHC allele set generation, not that selection seed. The selected knot budget, reference hash, and available validation diagnostics are saved with the curve; see Percentile calibration for all predictors.

The direct Python training APIs use seed=None by default, preserving their historical stochastic behavior unless the caller opts in.

A fixed seed does not guarantee identical weights across devices, dependency versions or numerical settings. Preserve those settings and the effective training batch alongside the seed. See the PyTorch reproducibility guide for backend-specific determinism controls.

Unified and historical command names

MHCflurry 2.3 groups commands under one mhcflurry entry point. For example:

mhcflurry-predict                       = mhcflurry predict
mhcflurry-downloads fetch              = mhcflurry downloads fetch
mhcflurry-class1-train-pan-allele-models = mhcflurry class1-train-pan-allele-models

Both forms invoke the same implementation. New automation can use the unified form; existing scripts do not need to change.