Cached predictions¶
Running an MHC predictor twice on the same peptide is wasted work.
CachedPredictor serves predictions from a pre-computed table — an
external predictor's output, a previous topiary run, or shards from
parallel jobs — and plugs into TopiaryPredictor(models=…) the same
way a live predictor does.
Use it when you want to:
- Iterate on filters / ranking / the DSL without re-running a slow predictor.
- Pin scores for reproducibility (papers, benchmarks).
- Predict per-allele / per-sample in parallel jobs, persist each shard, merge later.
- Ingest predictions from a tool topiary doesn't natively run.
Basic usage¶
Load a cache, pass it as the model to TopiaryPredictor, go:
from topiary import CachedPredictor, TopiaryPredictor
cache = CachedPredictor.from_topiary_output("run.parquet")
predictor = TopiaryPredictor(models=cache)
df = predictor.predict_from_variants(variants)
The cache answers predict_proteins_dataframe and
predict_peptides_dataframe calls from the table, so every
predict_from_* method on TopiaryPredictor works unchanged.
For whole-peptide half-life tables, use predict_from_named_peptides to replay
the complete peptides originally scored. Known allele-independent kinds retain
mhc_dependence="none" without a live fallback; caching a measurement does not
make it allele-specific. A fallback's explicit kind metadata takes precedence.
The cache does not yet identify different chemistry, assay conditions or model
settings beyond the stored method/version: do not mix those experiments in one
cache (#288).
Incomplete cache coverage¶
By default, a coverage gap raises CachedPredictorCoverageError and stops the
run. This includes missing peptides and incompatible flank, kind or genotype
contexts. raise_on_error=False only affects variant annotation errors.
To retain successful model/input pairs, explicitly supply a report handler:
misses = []
predictor = TopiaryPredictor(models=cache, cache_miss_handler=misses.append)
df = predictor.predict_from_named_sequences(proteins)
# Persist misses alongside df. An empty list means no cache misses were skipped.
A failed sequence loses all its rows for the affected model, including any
covered windows. Other sequences and other models continue. Full sequences are
used when isolating misses, so flank-dependent predictions keep their original
context. Reports identify source_sequence_name, model_key, method/version,
stage (protein, peptide, wildtype, or self_nearest), error type and
message. WT or self-nearest failures leave their comparator scores missing;
they do not remove successful primary predictions. Successful rows remain
ordinary DataFrame rows, so retain the separate report even after filtering.
PartialPredictionWarning also signals incomplete results. Unrelated exceptions
and failures in the handler still raise.
The CLI uses a mandatory JSON sidecar for this opt-in mode:
topiary --fasta proteins.fasta --mhc-cache-file predictions.csv \
--cache-miss-report missing-predictions.json --output-csv results.csv
The sidecar records complete, prediction_rows (after filters), and failures,
including when all inputs fail or all succeed. Exit status 3 means a partial
prediction run; 0 means no reported cache misses. A malformed command line
exits 2 with the usage text; any other fatal error, including a cache that
cannot answer without this option, exits 1 with one line. CSV stdout remains
parseable. The report is written before the
prediction output, and its path must differ from inputs and outputs.
complete concerns cache coverage only, not biological evidence, filtering,
or variant annotation skipped under --skip-variant-errors.
Lower-level callers can use the public predict_with_cache_miss_report policy.
See the implementation and compatibility notes.
Selecting lengths and retrieving different values¶
A cache replays stored measurements. Different peptide, allele, prediction-kind,
or stored context entries can have different scores. Selecting a different set
of entries does not recompute those scores. For example, a table containing
SIINFEKL → 11, IINFEKLA → 22, and SIINFEKLA → 33 returns the first two
values when scanning 8-mers and the last value when scanning 9-mers. These
numbers are illustrative test data, not biological predictions.
cache = CachedPredictor.from_topiary_output("predictions.csv")
print(cache.available_peptide_lengths) # lengths in the table or its fallback
cache.default_peptide_lengths = [9]
predictor = TopiaryPredictor(models=cache)
df = predictor.predict_from_named_sequences({"protein": "SIINFEKLA"})
cache.default_peptide_lengths = None # restore all available lengths
Length selection limits the protein windows generated before cache lookup. An unavailable requested length raises an error; available lengths do not guarantee that every peptide/allele/context is covered. An individual miss still raises a coverage error or uses the configured fallback.
The selection leaves the cache table intact, including measurements at other
lengths. Explicit peptide inputs (--peptide-csv, --peptide-fasta, or
predict_from_named_peptides) query the supplied peptides as-is. Saving and
reloading a cache preserves its measurements; scan-length selection is a
per-run setting and must be applied again after loading.
From the CLI:
topiary --fasta proteins.fasta --mhc-cache-file predictions.csv \
--mhc-peptide-lengths 9 --mhc-alleles 'HLA-A*02:01' \
--subset-output-columns peptide allele kind value score percentile_rank \
--output-csv -
Omitting the length flag uses all available lengths. The legacy
--mhc-epitope-lengths flag also works; --mhc-peptide-lengths takes precedence
if both are supplied. Requested length coverage is checked after selecting
the requested alleles.
To inspect other metrics, select columns such as value, score, and
percentile_rank when present; kind identifies which prediction each row
describes. Filters and ranking select or reorder results from the stored
measurements. To obtain new predictions for an existing entry, run the intended
live predictor and save a new table, then load that table as the cache.
--mhc-cache-predictor-version is a provenance label, not a request to run a
different model. CLI cache flags and --mhc-predictor are mutually exclusive.
Loaders¶
From topiary's own prediction output¶
Round-trip a prior TopiaryPredictor run through Parquet (preferred)
or TSV:
# Save
first_run_df.to_parquet("run.parquet", index=False)
# Reload on subsequent iterations
cache = CachedPredictor.from_topiary_output("run.parquet")
Live MHCflurry output now records its package and official model release when
the weights are loaded (mhctools 3.44.26+, required by Topiary). That identity
survives replay without consulting the MHCflurry installation on the machine
reading the cache. Custom or injected weights need an explicit
predictor_version when constructing the MHCflurry wrapper; they are not labeled
as the default release.
For older output missing provenance, supply the identity of the original run, not whichever model is installed now:
topiary --peptide-csv peptides.csv --mhc-cache-file old-run.csv \
--mhc-cache-format topiary_output \
--mhc-cache-predictor-version '2.2.1+release-2.2.0' \
--output-csv replay.csv
The version above is an example: use it only if it describes the generating
run. --mhc-cache-predictor-name likewise fills missing method names. These
arguments fill absent columns and unstated cells; conflicting recorded values
raise an error, never silently relabel the predictions. The library loaders
accept the same prediction_method_name and predictor_version arguments,
including from_directory. Numeric-looking versions stay text: 2.10 is not
changed into 2.1.
Schema: every column _predict_raw* produces is preserved. Topiary
overlay columns (fragment_id, wt_peptide, etc.) are kept in the
file but ignored on lookup — the cache is the predictor output, not
the full pipeline output.
From NetMHC-family stdout captures¶
The DTU NetMHC suite (NetMHCpan, NetMHC, NetMHCcons, NetMHCIIpan, NetMHCstabpan) prints binding predictions to stdout when run from the command line. If you've captured that output to a file, load it directly — topiary reuses the parsers shipped by mhctools so every NetMHC version they support is covered here too.
cache = CachedPredictor.from_netmhcpan_stdout("netmhcpan_run.out")
cache = CachedPredictor.from_netmhc_stdout("netmhc_run.out", version="4")
cache = CachedPredictor.from_netmhcpan_cons_stdout("cons_run.out")
cache = CachedPredictor.from_netmhciipan_stdout("ii_run.out", version="4.3")
cache = CachedPredictor.from_netmhcstabpan_stdout("stab_run.out")
Each loader parses the version out of the stdout preamble (e.g.
NetMHCpan version 4.1b) and stamps it on predictor_version.
Pass predictor_version="..." if you want a different label, or
if your capture stripped the preamble.
Multi-kind output: NetMHCpan run with -BA (which produces
both eluted-ligand and binding-affinity scores per peptide) is
loaded as both kinds in the cache — pMHC_affinity and
pMHC_presentation rows for each (peptide, allele). The new
6-tuple index key distinguishes them, and downstream topiary DSL
scopes (Affinity.* vs Presentation.*) read the right rows.
The loaders parse the stdout text format. NetMHC's -xlsfile
tab-delimited output is a different format and isn't supported
today — open an issue if you need it.
From mhcflurry output¶
cache = CachedPredictor.from_mhcflurry("mhcflurry_predictions.csv")
The loader maps mhcflurry's mhcflurry_affinity /
mhcflurry_affinity_percentile / mhcflurry_presentation_score
columns onto topiary's canonical affinity / percentile_rank /
score. predictor_version is auto-composed from the installed
mhcflurry — see mhcflurry version composition
below. Automatic composition is appropriate only when that installation and
bundle generated the file; otherwise pass its original predictor_version.
Generic TSV / CSV with column mapping¶
For any tab- or comma-delimited file that doesn't match a format topiary ships a dedicated loader for:
cache = CachedPredictor.from_tsv(
"third_party.tsv",
columns={
"peptide": "Peptide",
"allele": "HLA",
"affinity": "IC50_nM",
"percentile_rank": "Rank%",
},
prediction_method_name="netchop",
predictor_version="3.1",
)
columns maps canonical cache columns to the column names in your
file. prediction_method_name and predictor_version are required
when the file doesn't embed that identity. They also fill unstated cells and
reject conflicts with recorded values.
Every row must carry kind, such as pMHC_affinity, either under that name
or mapped from another column (columns={"kind": "Assay"}). The loader does
not guess the kind from a score column or predictor label.
Pass sep="," for CSV files.
Sharding: merge multiple caches¶
Predict per-allele or per-sample in parallel, persist each shard separately, then merge them:
cache = CachedPredictor.concat([shard_a, shard_b, shard_c])
# Or: load every matching file from a directory
cache = CachedPredictor.from_directory(
"caches/",
pattern="*.parquet",
)
Every shard must share the same
(prediction_method_name, predictor_version) — the core invariant
applies across shards the same way it applies inside one.
Malformed shards report their filename and the required topiary_output
format. Missing or unreadable files report the underlying file-access error.
Overlap resolution (on_overlap=):
"raise"(default) — fail if any(peptide, allele, peptide_length)appears in more than one shard. A sample of conflicting keys is included in the error. Use this if shards should be disjoint."last"— later shard in the input list wins. Useful when the sort order represents "newer overwrites older.""first"— earlier shard wins.callable(row_a, row_b) -> row— custom resolver. Called pairwise per duplicate group. Pattern for "keep stronger binder":
def keep_lower_affinity(a, b):
return a if a["affinity"] <= b["affinity"] else b
cache = CachedPredictor.concat(shards, on_overlap=keep_lower_affinity)
from_directory passes on_overlap through to concat; file order
is sorted lexicographically, so shard_a.tsv is always earlier than
shard_b.tsv.
From an in-memory DataFrame¶
cache = CachedPredictor.from_dataframe(
df,
prediction_method_name="custom",
predictor_version="v7",
)
For programmatic construction — tests, scripts that built predictions some other way, or intermediates in a larger pipeline.
Version invariant¶
A single CachedPredictor holds predictions from exactly one
(prediction_method_name, predictor_version) pair. Scores from
different model versions aren't interchangeable — even percentile
ranks aren't directly comparable across versions. Mixing them would
produce an output that passes every downstream filter invisibly.
Enforcement:
- On construction: a DataFrame with multiple
(name, version)pairs raises.None/NaN/ empty-string values are also rejected — silent "I don't know" would mask the invariant. - On fallback attachment (below): the fallback's
(name, version)must equal the cache's, verified on the first fallback call. - On concat /
from_directory: every shard must agree.
Explicit opt-in equivalence¶
Sometimes two similar versions do produce identical predictions — release candidate → final, or a timestamp-only model-data reflash. Opt in explicitly:
cache = CachedPredictor.from_mhcflurry(
"path.csv",
predictor_version="2.2.0rc2",
also_accept_versions={"2.2.0rc1", "2.2.0"},
)
A fallback or shard passes if its version equals the cache's or is
in also_accept_versions. Names are still strict — mixing mhcflurry
and NetMHCpan is a real type mismatch, not a version wiggle.
mhcflurry version composition¶
Unlike NetMHCpan (which bakes models into the binary), mhcflurry
fetches model weights separately via mhcflurry-downloads fetch.
Two systems on the same mhcflurry package version can produce
different predictions if they have different model bundles installed.
predictor_version must capture both.
topiary.mhcflurry_composite_version() introspects the installed
mhcflurry and returns a composite string like
"2.2.1+release-2.2.0". CachedPredictor.from_mhcflurry(path)
calls it automatically when you omit predictor_version — you
never have to enumerate model bundles manually:
# Auto-composed from the local install
cache = CachedPredictor.from_mhcflurry("predictions.csv")
# Explicit override if you want a custom label
cache = CachedPredictor.from_mhcflurry(
"predictions.csv",
predictor_version="my-project-20260414",
)
The helper raises with a clear message if mhcflurry isn't installed or no model release is configured.
Other tools (NetMHCpan, NetMHCIIpan, etc.) don't need this — their binary version is their full identity.
Fallback: delegate misses to a live predictor¶
If you want a cache that falls back to a live predictor for peptides not yet in the table:
from mhctools import MHCflurry
live = MHCflurry(alleles=[...])
cache = CachedPredictor.from_mhcflurry(
"partial.csv",
fallback=live,
)
predictor = TopiaryPredictor(models=cache)
df = predictor.predict_from_variants(variants) # misses delegate
Semantics:
- Miss routed to
fallback, and the fallback's output is merged into the cache so subsequent queries for the same(peptide, allele, peptide_length)serve locally. Caching hits is always the right default — there's no separate flag. - Same version invariant: the fallback's
(name, version)must equal the cache's (or be inalso_accept_versions), checked on the first fallback call. - Pure read-through mode — empty cache, fallback-only — is
supported:
CachedPredictor(fallback=live). The cache starts empty; identity is discovered from the fallback's first output. - No fallback (default): misses raise
KeyErrorwith the missed peptides listed.
A cached protein-scan occurrence with incompatible flanks or genotype raises
CachedPredictorCoverageError, even when a fallback is attached. The current
fallback path does not retry such occurrences through the fallback's protein
scanner. Context-preserving fallback is tracked in
#307.
Persisting a cache¶
cache.save("cache.parquet") # or .tsv / .tsv.gz / .csv
Writes the cache's internal table using the same schema the loaders
expect. Round-trips cleanly through from_topiary_output.
From the CLI¶
Every loader is exposed through topiary's command-line interface via
--mhc-cache-* flags. The CLI runs the entire prediction pipeline
(variant effects, filtering, ranking, output formatting) against the
cached predictions — the live predictor is never invoked.
--mhc-cache-format is optional for NetMHC-family stdout captures,
mhcflurry CSVs, and topiary's own output — topiary sniffs the format
from file content. Only the generic tsv path needs an explicit
--mhc-cache-format tsv (generic tables don't carry identifying
signatures).
NetMHCpan stdout capture (format sniffed from preamble):
topiary --peptide-csv peptides.csv \
--mhc-cache-file netmhcpan_run.out \
--output-csv results.csv
mhcflurry CSV (format sniffed from column names; predictor_version
auto-composed from the local install):
topiary --peptide-csv peptides.csv \
--mhc-cache-file mhcflurry_predictions.csv \
--output-csv results.csv
Topiary's own saved output (Parquet or TSV round-trip — sniffed):
topiary --peptide-csv peptides.csv \
--mhc-cache-file prior_run.parquet \
--output-csv results.csv
Generic TSV with column mapping (--mhc-cache-format tsv is required
here since generic tables can't be auto-detected). This example requires an
Assay column containing the kind of each measurement, e.g. pMHC_affinity:
topiary --peptide-csv peptides.csv \
--mhc-cache-file third_party.tsv \
--mhc-cache-format tsv \
--mhc-cache-predictor-name my-affinity-model \
--mhc-cache-predictor-version 3.1 \
--mhc-cache-tsv-column affinity=IC50_nM \
--mhc-cache-tsv-column percentile_rank=Rank \
--mhc-cache-tsv-column kind=Assay \
--output-csv results.csv
Sharded — merge every matching file in a directory:
topiary --peptide-csv peptides.csv \
--mhc-cache-directory ./caches \
--mhc-cache-directory-pattern '*.parquet' \
--output-csv results.csv
Full flag reference:
| Flag | Purpose |
|---|---|
--mhc-cache-file PATH |
Single cache file. Format is auto-detected; --mhc-cache-format only needed for generic TSVs. |
--mhc-cache-directory PATH |
Directory of shards; each file loaded via from_topiary_output and concatenated. Alternative to --mhc-cache-file. |
--mhc-cache-directory-pattern GLOB |
Pattern for --mhc-cache-directory. Default *. |
--mhc-cache-format FORMAT |
Optional — sniffed from file content when omitted (see above). One of topiary_output, mhcflurry, tsv, netmhcpan, netmhc, netmhccons, netmhciipan, netmhcstabpan. Only tsv strictly requires the explicit flag. |
--mhc-cache-predictor-name NAME |
Fill missing prediction_method_name in generic TSV, topiary output or directory shards; recorded values must agree. Other formats already identify their predictor. |
--mhc-cache-predictor-version V |
Supply version provenance. For generic TSV, topiary output and shards, fills missing values and rejects conflicts. NetMHC stdout defaults to its preamble; raw MHCflurry defaults to the configured official local bundle. |
--mhc-cache-tsv-column CANONICAL=FILE_COL |
Repeatable. Column-name mapping for tsv format. |
--mhc-cache-tsv-sep SEP |
Separator for tsv. Default tab. |
--mhc-cache-netmhc-version V |
3, 4, or 4.1 for classic NetMHC output. Default 4. |
--mhc-cache-netmhciipan-version V |
legacy, 4, or 4.3 for NetMHCIIpan. Default 4.3. |
--mhc-cache-netmhciipan-mode MODE |
binding_affinity or elution_score (default) for NetMHCIIpan 4+. |
--mhc-predictor is mutually exclusive with cache flags. --mhc-alleles
is optional: omit it to use the cache's allele set, or supply a covered subset.
When not to use¶
- You've never run the predictor on these peptides — just use the live predictor directly; there's nothing to cache yet.
- You want to mix predictions from multiple model versions in one
pipeline. Don't. If you think you need this, revisit whether the
versions are actually interchangeable; if they are, use
also_accept_versions.