Presentation percentile calibration
New calibration now uses the shared CompactPercentRankTransform, also used
for affinity and processing. See shared APIs and compatibility.
The automatic budget starts at 64 and permits 128 only after a meaningful
label-free background-validation improvement. Historical calibrations still
load as HistogramPercentRankTransform, without conversion or refitting.
Experimental compact-curve comparison
mhcflurry eval presentation-percentiles compares histogram mappings with
continuous curves, without loading or retraining networks. The original
experiment included 64/128/256 knots; the current production policy permits
only 64/128. Specify those two budgets explicitly for a new comparison:
mhcflurry eval presentation-percentiles fit \
--reference-dir /path/to/cached-reference --out /path/to/new-comparison \
--knots 64 128
mhcflurry eval presentation-percentiles evaluate \
--calibration-dir /path/to/new-comparison \
--predictions /path/to/predictions_with_flanks.csv.bz2 --mode with-flanks
mhcflurry eval presentation-percentiles evaluate \
--calibration-dir /path/to/new-comparison \
--predictions /path/to/predictions_without_flanks.csv.bz2 --mode without-flanks
mhcflurry eval presentation-percentiles plot \
--calibration-dir /path/to/new-comparison
The reference directory contains reference_peptides.csv.bz2 and one aligned
reference_scores_*.npy vector per allele. Prediction tables contain
sample_id, hla, hit, and a_/b_ presentation score/percentile columns;
side a is the candidate and side b is the contextual public baseline.
This experimental command splits the reference 60/20/20 by unique peptide,
keeping all allele queries for a peptide together. Knot count is selected using
background validation accuracy, not presentation labels. Its separate test
split checks calibration. This is distinct from the production helper’s
80/20 selection/refit policy and minimum-count safeguards; see
Percentile calibration for all predictors for that rule. The experiment command’s
historical default still includes 256, so use --knots 64 128 as shown above.
Full-reference refits are then compared on the frozen presentation rows.
Outputs retain the splits, counts at percentile cutoffs, versioned curves,
micro/per-sample metrics, tie-neutral PPV, ranking audits, transformation timings,
and joinable prediction arrays, plus a report and PNG/SVG figures.
The experimental curves interpolate in logit-score/logit-survival coordinates, using weighted midranks at distinct reference scores. End slopes estimated from up to 32 distinct reference scores provide explicitly modeled extrapolation. Extreme extrapolated percentiles are not empirically established probabilities. Existing raw-score ties and floating-point saturation cannot be recovered by this mapping. The original experiment was isolated; its generic compact implementation is now used by new predictor calibration and versioned loading.
Historical tail-adaptive histogram method
With no explicit bins, method="histogram" calibration retains up to 10,000
uniform-quantile bins and adds up to 10,000 bins in the top 1% of calibration
scores. The base grid is bounded by the finite reference count. Additional quantiles
are spaced logarithmically in upper-tail probability, down to 1 / N for N
finite calibration scores. Their count is also capped by the number of
calibration observations in that tail. Duplicate score edges are removed.
Uniform quantiles can place thousands of informative high-scoring evaluation peptides in one percentile bin even though their raw scores are distinct. Preserving the original broad grid also retains the resolution needed when the presentation combiner’s score range is compressed near zero.
This remains available for reproducing the earlier binning experiment. Neither method changes network predictions or any existing saved calibration merely by loading it. No calibration rule is fitted to evaluation labels.
Calibrating a predictor
The existing mhcflurry-calibrate-percentile-ranks command uses the new default
automatically for --predictor-kind class1_presentation (compact by default;
--percentile-method histogram reproduces the earlier tail-bin method). Re-run the original
calibration recipe on a separate copy of the predictor, retaining
the same calibration reference policy, random seed, and source provenance.
For full candidates, that recipe is in
scripts/training/pan_allele_release_full.sh and the saved GENERATE.sh.
The calibration command writes into --models-dir; do not point it at the
public baseline or the only preserved copy of the old candidate.
An old percent_ranks.csv does not contain the raw calibration scores inside
each bin. Merely subdividing that table would invent within-bin CDF detail.
Regenerate calibration predictions when those scores were not retained, then
repeat both with-flanks and without-flanks held-out comparisons. More bins
cannot resolve missing calibration-tail support, true raw-score ties, or
clipping beyond the calibration maximum.
Historical histogram experiment verification
The earlier histogram regression experiment used an independent million-score reference and 4,000 evaluation scores in a rare upper-tail interval. Its AUPRC was 0.93023 for raw scores, 0.50025 with the old uniform-quantile grid, and 0.92849 with the tail-adaptive grid. These are historical synthetic results, not a recalibrated release-candidate performance claim. The current default calibration regression tests exercise compact interpolation instead.
Tests also cover preservation of all base-grid edges, bounded extra-bin count, small reference sets, repeated/nonfinite values, monotonicity and endpoint behavior, compressed scores, explicit bin overrides, and saved-table loading.
The archived real-candidate comparison found that 64 compact knots recovered
the raw presentation ranking on the frozen with-flanks cohort: macro AUPRC
0.337292 and macro PPV@N 0.387948 (using tie-neutral PPV), with the same ranking
metrics at 128 and 256 knots. This used the experiment’s independent cached
background, not a fresh calibration on the complete release-policy reference.
The artifact directory is
experiments/20260909T023206Z-compact-presentation-percentiles-comparison-9778c6256b7e.
Neither compact nor histogram calibration can turn those raw rankings into both metrics above 0.4. PPV@N’s row-order tie-breaking is a separate evaluation issue; it is not changed by this fix.