Evaluating trained models
Use the evaluation commands after training a model to compare it with a public release or another run. The normal workflow separates reusable metrics from plot rendering, so you can change figures without rerunning predictions.
Quick evaluation
Compare a candidate run with the installed public models:
mhcflurry eval compare-models \
--a results/new_run \
--b public \
--out results/new_run/eval_comparison
Render the diagnostic plots and combined PDF:
mhcflurry eval plot-comparison \
--input results/new_run/eval_comparison \
--summary-pdf results/new_run/eval_comparison/plots/model_comparison_figures.pdf
Use public:<release_name> instead of public when the baseline version must
be fixed. Components are compared only when both sides contain the corresponding
model artifacts; explicitly requested processing modes must exist on both sides.
Outputs
Stage |
Command |
Main outputs |
|---|---|---|
Metrics |
|
Component CSV/JSON files, |
Diagnostics |
|
ROC, precision-recall, scatter, and delta plots plus an optional combined PDF |
Paper figures |
|
SVG/PDF/PNG panels, |
The metrics directory is the reusable contract between evaluation and plotting. Keep it when iterating on figure style or assembling a review packet.
Interpreting comparisons
compare-models defines PPV@N with N equal to the group’s positive count.
If a score tie crosses the cutoff, its contribution is the expected number of
positives under uniform selection within that tie. This makes the metric
independent of input row order. metric_policy.json records this rule; older
comparison directories may use a different tie policy. Existing published
reports are not rewritten when the software is upgraded.
Report raw-score and percentile metrics separately, recording the calibration method and background. Preserve public baselines unchanged; see Percentile calibration for all predictors for controlled recalibration comparisons.
Before describing a cohort as sample-disjoint, audit the union of all compared models’ training, development and selection lineage. Peptide–MHC exclusion alone does not prove this. See Auditing training-sample overlap for the audit command, shared-cohort export, and limitations of historical and external evidence.
Paper-style figures
For an already-trained local model, compose comparison, diagnostics, and any available paper panels with one command:
mhcflurry eval paper-figures run \
--a results/new_run \
--b public \
--out results/new_run/eval_comparison
Paper panels can also use cached benchmark predictions or score tables. A saved prediction table has one row per evaluated peptide–MHC example and contains:
hit;sample_idfor multiallelic data, orallele/hlafor monoallelic data;optional peptide and flank metadata; and
canonical predictor score columns, or custom score columns declared in
predictor_info.csv.
Derive reusable AUC and PPV score tables with:
mhcflurry eval paper-figures score-predictions \
--kind multiallelic \
--input benchmark.multiallelic.csv.bz2 \
--out accuracy_scores.multiallelic.csv
Score direction must be explicit. Common MHCflurry, NetMHCpan, and MixMHCpred
column names have built-in orientation. Describe custom columns in
predictor_info.csv using predictor and higher_is_better; those rows also
select the custom columns. --predictor-columns can restrict scoring to an
explicit subset, but custom names still need their direction in
predictor_info.csv. Other numeric columns are treated as metadata rather than
guessed to be predictor scores.
Optional licensed predictors run outside the MHCflurry core package. The
paper-figures external-predictors adapter can invoke a locally installed
runner such as mhctools and join its output into the canonical benchmark
table. Missing optional inputs are listed in missing_inputs.md; they are not
silently replaced with synthetic panels.
Keep paper figures in their own directory (the default is
<out>/plots/paper_figures). The combined diagnostic --summary-pdf may be a
top-level file under <out>/plots or live outside the plot tree, but it cannot
be placed inside the paper-figure, affinity, processing, presentation, or
diagnostic-paper subdirectories. Commands reject overlapping output paths
before clearing or rendering anything.
External predictor comparison
mhcflurry eval presentation-external-predictors compares saved compare-models
scores with the NetMHCpan 4.0 BA, NetMHCpan 4.0 EL and MixMHCpred columns
distributed in data_evaluation. It runs no predictor:
mhcflurry eval presentation-external-predictors \
--comparison-dir results/new_run/eval_comparison \
--data-dir "$(mhcflurry downloads path data_evaluation)" \
--cohort multiallelic --coverage common \
--a-label "MHCflurry 2.3.0" --b-label "MHCflurry 2.2" \
--out results/new_run/external_comparison
Rows join by benchmark source file and row identity, with MHC allele sets
canonicalized as compare-models saves them; any unmatched row fails the command.
data_evaluation ships NetMHCpan 4.0 BA/EL and MixMHCpred columns. Pass
--external-dir once per directory of locally generated scores, such as
NetMHCpan 4.1/4.2, whose files may be plain CSV. External training overlap
is not certified: a tool version or release date alone does not prove that
these evaluation samples were excluded from its training.
multiallelic uses the saved presentation scores with and without flanks;
monoallelic uses saved affinity predictions (pass --skip-joined-table for
that large cohort). Metrics use the compare-models definitions. All requested
NetMHCpan versions, both BA and EL, receive paired comparisons.
Use --coverage common for reported comparisons. Every table, paired interval
and figure then uses identical rows across all requested models, and every
sample must remain represented with both classes. The default diagnostic
available mode instead preserves per-model coverage: each predictor uses its
covered rows and each pair uses their intersection. coverage.csv records
unscored rows, original missing-score counts and common exclusions; provenance
records original and scored hit/row counts.
Paired intervals resample whole samples (10000 draws, seed 42 by default) and
are exploratory. Two reference baselines are included by default: seeded random
scores, and a logistic regression on the one-hot first and last four residues
with no MHC or flank input, fitted leave-one-sample-out so no sample’s labels
reach its own scores. Pass --baselines none to skip them. Outputs include
per-sample, macro and pooled metrics, paired differences, a joined score table,
external_comparison.pdf with PNG pages, summary.md and a provenance
manifest.
Specialized workflows
Reusing saved affinity predictions
Affinity comparison summaries include a benchmark_identity hash calculated
after holdout selection, allele intersection, peptide-length filtering, and
training-overlap exclusion. A saved prediction column can be reused without
rerunning a baseline only when that identity matches exactly:
mhcflurry eval compare-models \
--a results/candidate/models.combined \
--b public \
--b-affinity-predictions results/public-comparison/affinity/predictions.csv.bz2 \
--b-affinity-prediction-column b_pred \
--out results/candidate/comparison-vs-public
Matched processing cohorts
For processing pools that cannot supply ten distinct negatives per hit, expand and freeze a single cohort before comparing models:
mhcflurry eval prepare-processing-cohort \
--data-dir DATA_EVALUATION \
--release-holdout-dir results/new_run/release_holdout \
--affinity-predictor PUBLIC_MODELS_CLASS1_PAN/models.combined \
--proteome-reference-csv HUMAN_PROTEOME.csv.bz2 \
--out results/processing_cohort
mhcflurry eval compare-models \
--a results/new_run --b public:2.2.0 \
--include processing --data-dir DATA_EVALUATION \
--release-holdout-dir results/new_run/release_holdout \
--processing-matched-cohort results/processing_cohort \
--out results/new_run/processing_comparison
The preparation command retains every held-out hit and the original cached
public affinities, samples additional candidates only for unresolved lengths,
and scores them with the frozen reference. It first checks that this reference
reproduces 256 spread-out cached predictions per sample within 0.0001 log10
units. Input, model and output checksums, the checked rows, seeds, scored rounds
and unique assignments are saved. --resume reuses verified rounds; insufficient
capacity remains an error. This processing-only expansion does not alter the
presentation benchmark. The remote launcher accepts the resulting directory
through PROCESSING_EVALUATION_COHORT.
Count-matched processing ensembles
To compare a four-network candidate with an eight-network public ensemble, evaluate all four-of-eight public subsets rather than choosing a subset using release-holdout performance:
mhcflurry eval processing-ensemble-subsets \
--input results/matched/matched_predictions.csv.bz2 \
--models-dir /path/to/public/models.selected.short_flanks \
--subset-size 4 --reference-score public_5aa \
--comparison-score new_legacy_cnn --comparison-score new_boundary_5x5 \
--out results/public_four_network_subsets
This source-checkout command requires strict length/affinity-matched risk sets, caches each network’s predictions, verifies reconstruction against the named full-ensemble score, and preserves all subset scores and memberships. It reports median, range and quartiles across the complete subset set. This spread measures ensemble composition sensitivity; it is not a confidence interval or a validation-based model selection procedure. Matching network count does not match the historical training or architecture-search budget.
Use a fresh output directory. --member-cache-dir can reuse an earlier run’s
predictions after checking input, model and execution identities. Cross-device
score tolerances must be explicitly justified and recorded; do not increase
--verification-atol to hide changed weights or a misaligned prediction table.
Affinity candidate figures
For affinity-factorial finalists, mhcflurry eval affinity-candidate-figures
combines the direct, row-identical candidate/public prediction tables into one
monoallelic score table and figure suite. Its paginated AUROC, AUPRC, and PPV
grids show every requested candidate against every available baseline, and its
overview ranks all predictors by macro allele-level metrics. Canonical
NetMHCpan BA/EL and MixMHCpred columns are retained when present.
Use --external-predictions to add a benchmark-aligned table containing
NetMHCpan BA/EL or MixMHCpred scores. The table can be built from official
per-sample data_evaluation groups with mhcflurry eval merge-external-predictions; both commands validate stable row identity and
record input hashes and coverage in figure provenance.
Release and remote runs
mhcflurry train pan-allele-release runs comparison and diagnostic plotting on
the training machine before copying results home. Paper inputs that depend on
locally licensed tools can remain on the control machine and be rendered after
the remote artifacts arrive.
Release maintainers should use the
release workflow guide
for lifecycle, synchronization, and deployment options. The complete argument
reference is in Command-line reference and mhcflurry eval --help.