Skip to content

Cached predictions

Running an MHC predictor twice on the same peptide is wasted work. CachedPredictor serves predictions from a pre-computed table — an external predictor's output, a previous topiary run, or shards from parallel jobs — and plugs into TopiaryPredictor(models=…) the same way a live predictor does.

Use it when you want to:

  • Iterate on filters / ranking / the DSL without re-running a slow predictor.
  • Pin scores for reproducibility (papers, benchmarks).
  • Predict per-allele / per-sample in parallel jobs, persist each shard, merge later.
  • Ingest predictions from a tool topiary doesn't natively run.

Basic usage

Load a cache, pass it as the model to TopiaryPredictor, go:

from topiary import CachedPredictor, TopiaryPredictor

cache = CachedPredictor.from_topiary_output("run.parquet")
predictor = TopiaryPredictor(models=cache)
df = predictor.predict_from_variants(variants)

The cache answers predict_proteins_dataframe and predict_peptides_dataframe calls from the table, so every predict_from_* method on TopiaryPredictor works unchanged.

For whole-peptide half-life tables, use predict_from_named_peptides to replay the complete peptides originally scored. Known allele-independent kinds retain mhc_dependence="none" without a live fallback; caching a measurement does not make it allele-specific. A fallback's explicit kind metadata takes precedence. The cache does not yet identify different chemistry, assay conditions or model settings beyond the stored method/version: do not mix those experiments in one cache (#288).

Incomplete cache coverage

By default, a coverage gap raises CachedPredictorCoverageError and stops the run. This includes missing peptides and incompatible flank, kind or genotype contexts. raise_on_error=False only affects variant annotation errors.

To retain successful model/input pairs, explicitly supply a report handler:

misses = []
predictor = TopiaryPredictor(models=cache, cache_miss_handler=misses.append)
df = predictor.predict_from_named_sequences(proteins)
# Persist misses alongside df. An empty list means no cache misses were skipped.

A failed sequence loses all its rows for the affected model, including any covered windows. Other sequences and other models continue. Full sequences are used when isolating misses, so flank-dependent predictions keep their original context. Reports identify source_sequence_name, model_key, method/version, stage (protein, peptide, wildtype, or self_nearest), error type and message. WT or self-nearest failures leave their comparator scores missing; they do not remove successful primary predictions. Successful rows remain ordinary DataFrame rows, so retain the separate report even after filtering. PartialPredictionWarning also signals incomplete results. Unrelated exceptions and failures in the handler still raise.

The CLI uses a mandatory JSON sidecar for this opt-in mode:

topiary --fasta proteins.fasta --mhc-cache-file predictions.csv \
    --cache-miss-report missing-predictions.json --output-csv results.csv

The sidecar records complete, prediction_rows (after filters), and failures, including when all inputs fail or all succeed. Exit status 3 means a partial prediction run; 0 means no reported cache misses. A malformed command line exits 2 with the usage text; any other fatal error, including a cache that cannot answer without this option, exits 1 with one line. CSV stdout remains parseable. The report is written before the prediction output, and its path must differ from inputs and outputs. complete concerns cache coverage only, not biological evidence, filtering, or variant annotation skipped under --skip-variant-errors.

Lower-level callers can use the public predict_with_cache_miss_report policy. See the implementation and compatibility notes.

Selecting lengths and retrieving different values

A cache replays stored measurements. Different peptide, allele, prediction-kind, or stored context entries can have different scores. Selecting a different set of entries does not recompute those scores. For example, a table containing SIINFEKL → 11, IINFEKLA → 22, and SIINFEKLA → 33 returns the first two values when scanning 8-mers and the last value when scanning 9-mers. These numbers are illustrative test data, not biological predictions.

cache = CachedPredictor.from_topiary_output("predictions.csv")
print(cache.available_peptide_lengths)  # lengths in the table or its fallback
cache.default_peptide_lengths = [9]
predictor = TopiaryPredictor(models=cache)
df = predictor.predict_from_named_sequences({"protein": "SIINFEKLA"})
cache.default_peptide_lengths = None   # restore all available lengths

Length selection limits the protein windows generated before cache lookup. An unavailable requested length raises an error; available lengths do not guarantee that every peptide/allele/context is covered. An individual miss still raises a coverage error or uses the configured fallback.

The selection leaves the cache table intact, including measurements at other lengths. Explicit peptide inputs (--peptide-csv, --peptide-fasta, or predict_from_named_peptides) query the supplied peptides as-is. Saving and reloading a cache preserves its measurements; scan-length selection is a per-run setting and must be applied again after loading.

From the CLI:

topiary --fasta proteins.fasta --mhc-cache-file predictions.csv \
    --mhc-peptide-lengths 9 --mhc-alleles 'HLA-A*02:01' \
    --subset-output-columns peptide allele kind value score percentile_rank \
    --output-csv -

Omitting the length flag uses all available lengths. The legacy --mhc-epitope-lengths flag also works; --mhc-peptide-lengths takes precedence if both are supplied. Requested length coverage is checked after selecting the requested alleles.

To inspect other metrics, select columns such as value, score, and percentile_rank when present; kind identifies which prediction each row describes. Filters and ranking select or reorder results from the stored measurements. To obtain new predictions for an existing entry, run the intended live predictor and save a new table, then load that table as the cache. --mhc-cache-predictor-version is a provenance label, not a request to run a different model. CLI cache flags and --mhc-predictor are mutually exclusive.

Loaders

From topiary's own prediction output

Round-trip a prior TopiaryPredictor run through Parquet (preferred) or TSV:

# Save
first_run_df.to_parquet("run.parquet", index=False)

# Reload on subsequent iterations
cache = CachedPredictor.from_topiary_output("run.parquet")

Live MHCflurry output now records its package and official model release when the weights are loaded (mhctools 3.44.26+, required by Topiary). That identity survives replay without consulting the MHCflurry installation on the machine reading the cache. Custom or injected weights need an explicit predictor_version when constructing the MHCflurry wrapper; they are not labeled as the default release.

For older output missing provenance, supply the identity of the original run, not whichever model is installed now:

topiary --peptide-csv peptides.csv --mhc-cache-file old-run.csv \
    --mhc-cache-format topiary_output \
    --mhc-cache-predictor-version '2.2.1+release-2.2.0' \
    --output-csv replay.csv

The version above is an example: use it only if it describes the generating run. --mhc-cache-predictor-name likewise fills missing method names. These arguments fill absent columns and unstated cells; conflicting recorded values raise an error, never silently relabel the predictions. The library loaders accept the same prediction_method_name and predictor_version arguments, including from_directory. Numeric-looking versions stay text: 2.10 is not changed into 2.1.

Schema: every column _predict_raw* produces is preserved. Topiary overlay columns (fragment_id, wt_peptide, etc.) are kept in the file but ignored on lookup — the cache is the predictor output, not the full pipeline output.

From NetMHC-family stdout captures

The DTU NetMHC suite (NetMHCpan, NetMHC, NetMHCcons, NetMHCIIpan, NetMHCstabpan) prints binding predictions to stdout when run from the command line. If you've captured that output to a file, load it directly — topiary reuses the parsers shipped by mhctools so every NetMHC version they support is covered here too.

cache = CachedPredictor.from_netmhcpan_stdout("netmhcpan_run.out")
cache = CachedPredictor.from_netmhc_stdout("netmhc_run.out", version="4")
cache = CachedPredictor.from_netmhcpan_cons_stdout("cons_run.out")
cache = CachedPredictor.from_netmhciipan_stdout("ii_run.out", version="4.3")
cache = CachedPredictor.from_netmhcstabpan_stdout("stab_run.out")

Each loader parses the version out of the stdout preamble (e.g. NetMHCpan version 4.1b) and stamps it on predictor_version. Pass predictor_version="..." if you want a different label, or if your capture stripped the preamble.

Multi-kind output: NetMHCpan run with -BA (which produces both eluted-ligand and binding-affinity scores per peptide) is loaded as both kinds in the cache — pMHC_affinity and pMHC_presentation rows for each (peptide, allele). The new 6-tuple index key distinguishes them, and downstream topiary DSL scopes (Affinity.* vs Presentation.*) read the right rows.

The loaders parse the stdout text format. NetMHC's -xlsfile tab-delimited output is a different format and isn't supported today — open an issue if you need it.

From mhcflurry output

cache = CachedPredictor.from_mhcflurry("mhcflurry_predictions.csv")

The loader maps mhcflurry's mhcflurry_affinity / mhcflurry_affinity_percentile / mhcflurry_presentation_score columns onto topiary's canonical affinity / percentile_rank / score. predictor_version is auto-composed from the installed mhcflurry — see mhcflurry version composition below. Automatic composition is appropriate only when that installation and bundle generated the file; otherwise pass its original predictor_version.

Generic TSV / CSV with column mapping

For any tab- or comma-delimited file that doesn't match a format topiary ships a dedicated loader for:

cache = CachedPredictor.from_tsv(
    "third_party.tsv",
    columns={
        "peptide": "Peptide",
        "allele": "HLA",
        "affinity": "IC50_nM",
        "percentile_rank": "Rank%",
    },
    prediction_method_name="netchop",
    predictor_version="3.1",
)

columns maps canonical cache columns to the column names in your file. prediction_method_name and predictor_version are required when the file doesn't embed that identity. They also fill unstated cells and reject conflicts with recorded values.

Every row must carry kind, such as pMHC_affinity, either under that name or mapped from another column (columns={"kind": "Assay"}). The loader does not guess the kind from a score column or predictor label.

Pass sep="," for CSV files.

Sharding: merge multiple caches

Predict per-allele or per-sample in parallel, persist each shard separately, then merge them:

cache = CachedPredictor.concat([shard_a, shard_b, shard_c])

# Or: load every matching file from a directory
cache = CachedPredictor.from_directory(
    "caches/",
    pattern="*.parquet",
)

Every shard must share the same (prediction_method_name, predictor_version) — the core invariant applies across shards the same way it applies inside one. Malformed shards report their filename and the required topiary_output format. Missing or unreadable files report the underlying file-access error.

Overlap resolution (on_overlap=):

  • "raise" (default) — fail if any (peptide, allele, peptide_length) appears in more than one shard. A sample of conflicting keys is included in the error. Use this if shards should be disjoint.
  • "last" — later shard in the input list wins. Useful when the sort order represents "newer overwrites older."
  • "first" — earlier shard wins.
  • callable(row_a, row_b) -> row — custom resolver. Called pairwise per duplicate group. Pattern for "keep stronger binder":
def keep_lower_affinity(a, b):
    return a if a["affinity"] <= b["affinity"] else b

cache = CachedPredictor.concat(shards, on_overlap=keep_lower_affinity)

from_directory passes on_overlap through to concat; file order is sorted lexicographically, so shard_a.tsv is always earlier than shard_b.tsv.

From an in-memory DataFrame

cache = CachedPredictor.from_dataframe(
    df,
    prediction_method_name="custom",
    predictor_version="v7",
)

For programmatic construction — tests, scripts that built predictions some other way, or intermediates in a larger pipeline.

Version invariant

A single CachedPredictor holds predictions from exactly one (prediction_method_name, predictor_version) pair. Scores from different model versions aren't interchangeable — even percentile ranks aren't directly comparable across versions. Mixing them would produce an output that passes every downstream filter invisibly.

Enforcement:

  • On construction: a DataFrame with multiple (name, version) pairs raises. None / NaN / empty-string values are also rejected — silent "I don't know" would mask the invariant.
  • On fallback attachment (below): the fallback's (name, version) must equal the cache's, verified on the first fallback call.
  • On concat / from_directory: every shard must agree.

Explicit opt-in equivalence

Sometimes two similar versions do produce identical predictions — release candidate → final, or a timestamp-only model-data reflash. Opt in explicitly:

cache = CachedPredictor.from_mhcflurry(
    "path.csv",
    predictor_version="2.2.0rc2",
    also_accept_versions={"2.2.0rc1", "2.2.0"},
)

A fallback or shard passes if its version equals the cache's or is in also_accept_versions. Names are still strict — mixing mhcflurry and NetMHCpan is a real type mismatch, not a version wiggle.

mhcflurry version composition

Unlike NetMHCpan (which bakes models into the binary), mhcflurry fetches model weights separately via mhcflurry-downloads fetch. Two systems on the same mhcflurry package version can produce different predictions if they have different model bundles installed. predictor_version must capture both.

topiary.mhcflurry_composite_version() introspects the installed mhcflurry and returns a composite string like "2.2.1+release-2.2.0". CachedPredictor.from_mhcflurry(path) calls it automatically when you omit predictor_version — you never have to enumerate model bundles manually:

# Auto-composed from the local install
cache = CachedPredictor.from_mhcflurry("predictions.csv")

# Explicit override if you want a custom label
cache = CachedPredictor.from_mhcflurry(
    "predictions.csv",
    predictor_version="my-project-20260414",
)

The helper raises with a clear message if mhcflurry isn't installed or no model release is configured.

Other tools (NetMHCpan, NetMHCIIpan, etc.) don't need this — their binary version is their full identity.

Fallback: delegate misses to a live predictor

If you want a cache that falls back to a live predictor for peptides not yet in the table:

from mhctools import MHCflurry
live = MHCflurry(alleles=[...])

cache = CachedPredictor.from_mhcflurry(
    "partial.csv",
    fallback=live,
)

predictor = TopiaryPredictor(models=cache)
df = predictor.predict_from_variants(variants)   # misses delegate

Semantics:

  • Miss routed to fallback, and the fallback's output is merged into the cache so subsequent queries for the same (peptide, allele, peptide_length) serve locally. Caching hits is always the right default — there's no separate flag.
  • Same version invariant: the fallback's (name, version) must equal the cache's (or be in also_accept_versions), checked on the first fallback call.
  • Pure read-through mode — empty cache, fallback-only — is supported: CachedPredictor(fallback=live). The cache starts empty; identity is discovered from the fallback's first output.
  • No fallback (default): misses raise KeyError with the missed peptides listed.

A cached protein-scan occurrence with incompatible flanks or genotype raises CachedPredictorCoverageError, even when a fallback is attached. The current fallback path does not retry such occurrences through the fallback's protein scanner. Context-preserving fallback is tracked in #307.

Persisting a cache

cache.save("cache.parquet")   # or .tsv / .tsv.gz / .csv

Writes the cache's internal table using the same schema the loaders expect. Round-trips cleanly through from_topiary_output.

From the CLI

Every loader is exposed through topiary's command-line interface via --mhc-cache-* flags. The CLI runs the entire prediction pipeline (variant effects, filtering, ranking, output formatting) against the cached predictions — the live predictor is never invoked.

--mhc-cache-format is optional for NetMHC-family stdout captures, mhcflurry CSVs, and topiary's own output — topiary sniffs the format from file content. Only the generic tsv path needs an explicit --mhc-cache-format tsv (generic tables don't carry identifying signatures).

NetMHCpan stdout capture (format sniffed from preamble):

topiary --peptide-csv peptides.csv \
    --mhc-cache-file netmhcpan_run.out \
    --output-csv results.csv

mhcflurry CSV (format sniffed from column names; predictor_version auto-composed from the local install):

topiary --peptide-csv peptides.csv \
    --mhc-cache-file mhcflurry_predictions.csv \
    --output-csv results.csv

Topiary's own saved output (Parquet or TSV round-trip — sniffed):

topiary --peptide-csv peptides.csv \
    --mhc-cache-file prior_run.parquet \
    --output-csv results.csv

Generic TSV with column mapping (--mhc-cache-format tsv is required here since generic tables can't be auto-detected). This example requires an Assay column containing the kind of each measurement, e.g. pMHC_affinity:

topiary --peptide-csv peptides.csv \
    --mhc-cache-file third_party.tsv \
    --mhc-cache-format tsv \
    --mhc-cache-predictor-name my-affinity-model \
    --mhc-cache-predictor-version 3.1 \
    --mhc-cache-tsv-column affinity=IC50_nM \
    --mhc-cache-tsv-column percentile_rank=Rank \
    --mhc-cache-tsv-column kind=Assay \
    --output-csv results.csv

Sharded — merge every matching file in a directory:

topiary --peptide-csv peptides.csv \
    --mhc-cache-directory ./caches \
    --mhc-cache-directory-pattern '*.parquet' \
    --output-csv results.csv

Full flag reference:

Flag Purpose
--mhc-cache-file PATH Single cache file. Format is auto-detected; --mhc-cache-format only needed for generic TSVs.
--mhc-cache-directory PATH Directory of shards; each file loaded via from_topiary_output and concatenated. Alternative to --mhc-cache-file.
--mhc-cache-directory-pattern GLOB Pattern for --mhc-cache-directory. Default *.
--mhc-cache-format FORMAT Optional — sniffed from file content when omitted (see above). One of topiary_output, mhcflurry, tsv, netmhcpan, netmhc, netmhccons, netmhciipan, netmhcstabpan. Only tsv strictly requires the explicit flag.
--mhc-cache-predictor-name NAME Fill missing prediction_method_name in generic TSV, topiary output or directory shards; recorded values must agree. Other formats already identify their predictor.
--mhc-cache-predictor-version V Supply version provenance. For generic TSV, topiary output and shards, fills missing values and rejects conflicts. NetMHC stdout defaults to its preamble; raw MHCflurry defaults to the configured official local bundle.
--mhc-cache-tsv-column CANONICAL=FILE_COL Repeatable. Column-name mapping for tsv format.
--mhc-cache-tsv-sep SEP Separator for tsv. Default tab.
--mhc-cache-netmhc-version V 3, 4, or 4.1 for classic NetMHC output. Default 4.
--mhc-cache-netmhciipan-version V legacy, 4, or 4.3 for NetMHCIIpan. Default 4.3.
--mhc-cache-netmhciipan-mode MODE binding_affinity or elution_score (default) for NetMHCIIpan 4+.

--mhc-predictor is mutually exclusive with cache flags. --mhc-alleles is optional: omit it to use the cache's allele set, or supply a covered subset.

When not to use

  • You've never run the predictor on these peptides — just use the live predictor directly; there's nothing to cache yet.
  • You want to mix predictions from multiple model versions in one pipeline. Don't. If you think you need this, revisit whether the versions are actually interchangeable; if they are, use also_accept_versions.