Skip to content

API Reference

Native Exacto input

read_exacto(path, *, sample_name, schema=None, primary_structures=None, reference_name=None, tag=None, transcript_read_support=None, library_id=None, read_set_id=None) reads native sequence/evidence tables without prediction. read_exacto_fragments(path, **kwargs) uses the same validation and returns optional fragments for explicit new-window scanning. EXACTO_SCHEMAS and EXACTO_SCHEMA_COMMIT describe the tested producer contract. See Reading Exacto for supported files, evidence scope and composition examples.

Fragment identity and RNA outcomes

A fragment record is identified by (sample_name, fragment_id): the ID names the candidate, ProteinFragment.sample_name names the observation. unique_fragments(fragments) returns records in first-occurrence order, coalescing a repeated pair only when its content agrees once passed through fragment IO (a record and its saved copy agree). Conflicting or uncomparable repeats, and one ID naming different candidates (CANDIDATE_FIELDS) in different samples, raise ValueError naming the differing fields. Prediction uses the same validation before model execution.

fragments_for_sample(fragments, sample_name) labels fragments from any source; fragments_from_variants(..., sample_name=) and a frame's sample_name column do the same. Prediction writes the label to the sample_name column, so the default aggregate_evidence_across_samples keys pool one candidate across samples. make_fragment_id(..., qualifiers=[...]) hashes any further values a producer groups records by.

describe_isovar_result(result) returns a JSON-compatible outcome dictionary for a completed Isovar result: variant, separate native read/template counts, protein and edit interval, transcript IDs, filter values and failed-filter names. Missing counts stay null. Status is passing, filtered, filter_status_unavailable, no_usable_reads, no_alt_reads, no_predicted_coding_change or no_protein_sequence. This reports recorded filters, not clinical suitability. Acquisition failures are not biological results. See the all-variant audit.

Amino-acid data

Topiary exposes one encoding shared by sequence consumers:

from topiary import (
    AMINO_ACIDS,
    AMINO_ACID_INDEX,
    BLOSUM62_AMINO_ACIDS,
    ENCODED_AMINO_ACIDS,
    UNKNOWN_AMINO_ACID_INDEX,
    blosum62_distance_matrix,
    blosum62_matrix,
    encode_amino_acids,
)

encoded = encode_amino_acids(["SIINFEKL", "AXX"], length=8)
scores = blosum62_matrix()
distances = blosum62_distance_matrix()

AMINO_ACIDS gives the canonical order, BLOSUM62_AMINO_ACIDS adds NCBI's B/J/Z ambiguity rows, and ENCODED_AMINO_ACIDS adds O/U/X/*. The immutable AMINO_ACID_INDEX mapping follows that order; any other character uses UNKNOWN_AMINO_ACID_INDEX.

blosum62_matrix() returns substitution scores. blosum62_distance_matrix() returns the distances used by SelfProteome: canonical pairs retain Topiary's established behavior, B/J/Z comparisons use their published scores with a symmetric conservative transformation, and O/U/X/* or unrecognized characters receive the symmetric worst-case canonical distance (15). Both functions return fresh read-only int8 arrays, so callers never share mutable matrix state.

Self-match evidence

match_self_peptides(reference, peptides, *, max_mismatches=0, alleles=None, excluded_gene_ids=(), observations=None, predictions=None) returns every same-length Hamming match within the radius, with all source origins and supplied presentation/prediction records. reference.match_candidates(...) delegates to it. self_matches_in_windows(windows, reference, *, peptide_lengths, ...) searches exact peptides at every requested window position. Results use ordinary typed CSV/TSV, with explicit coverage and evidence identity. See self evidence for the input schemas, policy composition and interpretation of missing results.

TopiaryPredictor

Parameter Type Description
models class, instance, or list Predictor model(s). Classes require alleles.
alleles list of str HLA alleles. Used to construct model classes.
filter_by DSLNode or str Boolean filter expression (e.g. Affinity <= 500 or "affinity <= 500 \| el.rank <= 2").
sort_by DSLNode or list of DSLNode Sort expression(s). Lexicographic tiebreakers; NaN falls through.
sort_direction "auto", "asc", or "desc" Direction for sort keys (auto infers per-key).
padding_around_mutation int Residues around mutation for candidate epitopes.
only_novel_epitopes bool Keep known mutation/target/junction overlaps; exclude unknown geometry. Does not establish tumor specificity.
min_gene_expression float Minimum gene FPKM (variant inputs).
min_transcript_expression float Minimum transcript FPKM (variant inputs).
raise_on_error bool Raise on variant-effect errors vs. skip.

Prediction methods

Method Input Behavior
predict_from_named_sequences(dict) {name: sequence} Sliding-window scan
predict_from_named_peptides(dict) {name: peptide} Score as-is
predict_from_peptide_occurrences(table) prediction_id, peptide, optional flanks/coordinates/annotations; unique by ID, peptide, offset and sample Score exact occurrences with their context; no scanning.
predict_from_sequences(list) [sequence, ...] Sliding-window scan
predict_from_fragments(fragments) [ProteinFragment] Universal path — any origin; fragment-level metadata and target_intervals threaded through.
predict_from_variants(variants) VariantCollection Variant pipeline (builds ProteinFragments internally and delegates).
predict_from_mutation_effects(effects) EffectCollection Same as predict_from_variants but starting from pre-computed effects.

predict_peptide_occurrences(table, model, use_flanks=True, predict_wt=False) exposes the same occurrence prediction for one configured mhctools model. It is exported from topiary and is also used by rescore_candidates. See exact occurrence prediction for input identity, unknown flanks, comparator and coverage rules.

prediction_mhc_scope(allele, dependence=..., allele_set=...) returns a canonical MHC identity tuple: one allele, the full genotype, or an allele-free scope. Missing required identity returns None, which must remain unmatchable. Comparator joins use this public rule in both fragment and occurrence prediction; a haplotype's selected presenter is not its genotype identity.

TopiaryResult

TopiaryResult is the semantic result object for Topiary prediction tables. It can ingest either long or wide prediction tables, keeps the active df view for pandas-style access, and materializes cached long_df and wide_df views on demand. The public form value describes the active df view; it is derived from the stored views rather than acting as separate result state. Use result.to_long() / result.to_wide() when you want a new TopiaryResult whose active df is that form.

Wide prediction columns use {model}_{kind}_{value|score|rank}. Wildtype companions use {model}_{kind}_wt_{value|score|rank} and round-trip with their MT prediction row; WT method and version fields do not create additional wide rows.

Topiary has two result-merging operations:

Operation Meaning Use when
topiary.stack_results([a, b]) / a.stack_with(b) More result sets Inputs are separate files, samples, cohorts, or independent result sets.
topiary.combine_predictions([a, b]) / a.combine_predictions(b) More predictions for the same logical identity grid Inputs are separate predictors or predictor shards that should behave like one run.

stack_results is a row-union operation. The inputs do not need to describe the same peptides, alleles, samples, sources, predictors, or score kinds; they are just more Topiary result rows with merged provenance.

combine_predictions is a prediction-union operation. The inputs are pieces of one logical prediction table: separate predictors over the same candidates, or one predictor split into disjoint allele/peptide-length shards. It rejects duplicate (prediction_method_name, kind, identity) rows and, by default, requires every emitted (prediction_method_name, kind) group to cover the same identity grid. Use coverage="partial" only for intentionally sparse prediction unions.

Both operations accept mixed long/wide TopiaryResult inputs and normalize internally. Fresh TopiaryPredictor DataFrame outputs may also be passed to combine_predictions; they are first wrapped as TopiaryResult objects and then validated with the same rules.

Cross-sample evidence

topiary.aggregate_evidence_across_samples(df, group_keys=None) returns one pooled row per prediction identity while leaving the input's per-sample rows unchanged. It collapses repeated prediction rows within a sample before summing canonical RNA and DNA counts, recomputes VAF from pooled alternate and overlapping counts, and reports the number of represented samples as n_samples.

A count is omitted unless every represented sample states it. Poolable counts must share an evidence subject across samples; allele-support counts must also share a derivation method. Coverage-only depths need no invented method. Incompatible units or methods raise rather than being mixed. Expression is not aggregated. By default the standard Topiary candidate identity is inferred, or callers can supply explicit group_keys. These exclude sample_name and the canonical counts/VAFs, which are aggregation values rather than candidate identity.

CachedPredictor

Drop-in replacement for a live predictor that serves scores from a pre-computed table. Pass as models=cache to TopiaryPredictor. See Cached Predictions for full detail.

Loader Source
CachedPredictor.from_topiary_output(path) Parquet / TSV / CSV previously written from a topiary run.
CachedPredictor.from_mhcflurry(path) mhcflurry-predict CSV output. predictor_version auto-composed from the installed mhcflurry when omitted.
CachedPredictor.from_netmhcpan_stdout(path) NetMHCpan stdout capture (auto-detects 2.8 / 3 / 4 / 4.1). Returns every kind present in the output — -BA runs surface both pMHC_affinity and pMHC_presentation rows per (peptide, allele).
CachedPredictor.from_netmhc_stdout(path, version=...) Classic NetMHC stdout (3 / 4 / 4.1).
CachedPredictor.from_netmhcpan_cons_stdout(path) NetMHCcons stdout.
CachedPredictor.from_netmhciipan_stdout(path, version=...) NetMHCIIpan stdout (legacy / 4 / 4.3).
CachedPredictor.from_netmhcstabpan_stdout(path) NetMHCstabpan stdout (pMHC stability).
CachedPredictor.concat([caches], on_overlap=...) Merge shards (all must share name+version). on_overlap: "raise" / "last" / "first" / callable.
CachedPredictor.from_directory(path, pattern="*", on_overlap=...) Glob a dir and concat every matching file.
CachedPredictor.from_tsv(path, columns=..., prediction_method_name=..., predictor_version=...) Generic tab- or comma-delimited.
CachedPredictor.from_dataframe(df, ...) In-memory DataFrame.
CachedPredictor(fallback=live_predictor) Empty cache, lazy identity discovery — pure read-through over a live model.

Constructor-level knobs:

Parameter Description
fallback Live predictor to route misses through. Result is merged back into the cache.
also_accept_versions Set of version strings treated as interchangeable with the cache's own version (opt-in equivalence for rc → final, etc.).

Helper: topiary.mhcflurry_composite_version() returns "<package_version>+release-<model_release>" for the locally-installed mhcflurry. Automatically used by from_mhcflurry when no explicit predictor_version is passed.

Kind accessors

Accessor Kind Default field
Affinity pMHC_affinity .value (IC50 nM)
Presentation pMHC_presentation .value
Stability pMHC_stability .value
Processing antigen_processing .value

Fields

Field Column read Description
.value value Raw prediction value
.rank percentile_rank Percentile rank (lower = better)
.score score Normalized score (higher = better)

Multi-model bracket syntax

Affinity["netmhcpan"].value       # filter to netmhcpan rows
Affinity["mhcflurry"].score       # filter to mhcflurry rows

Method matching is case-insensitive and substring-based. Errors with "Did you mean" on typos.

Column

Reference any DataFrame column in expressions:

Column("charge")                  # reads 'charge' column
Column("cysteine_count") <= 2     # returns a Comparison (boolean DSL node)

Errors with close-match suggestions when column doesn't exist. Raises TypeError for non-numeric columns in arithmetic / ordered comparisons.

Categorical equality and membership

For string / boolean / mixed-dtype columns, Column.eq / Column.ne / Column.isin produce an IsIn node that reads the column raw (no float cast):

Form Behavior
Column("mhc_class").eq("I") True for rows whose mhc_class equals "I"
Column("mhc_class").ne("II") True for mhc_class != "II" (NaN survives, pandas-style)
Column("mhc_class").isin(["I", "II"]) True for either value
~Column("mhc_class").isin([...]) Negate via ~

Column("x") == "y" is Python equality, not one of these: DSLNode keeps ordinary ==, so it evaluates to a plain bool. Every expression argument refuses a bool with a TypeError that names .eq(), rather than filtering or sorting on a constant.

In the string DSL, string literals are legal on the RHS of == / != only:

parse('mhc_class == "I"')
parse('affinity.value <= 500 & mhc_class != "II"')

Pre-built shortcuts

Symbol Definition
class_i IsIn("mhc_class", ["I"])
class_ii IsIn("mhc_class", ["II"])

Both require a mhc_class column. read_pvacseq() produces one; for fresh TopiaryPredictor results, derive it first:

from topiary import derive_mhc_class
df["mhc_class"] = derive_mhc_class(df["allele"])

DSL tree

Every DSL expression is a DSLNode. The tree is composed of:

Leaf Produces
Const(v) Constant scalar
Column(name) Per-group first-row of that column
IsIn(name, values, negate=False) Per-group boolean membership test against scalar values (no float cast); reachable via Column.eq / Column.ne / Column.isin
Field(kind, field, method=None, version=None, scope="") Per-group first-row of {scope}{field} for rows matching kind (and method / version if given)
Len(scope="") Peptide length
Count(chars, scope="") Amino-acid character count
Composite Purpose
BinOp(left, right, op) Elementwise +, -, *, /, **
UnaryOp(inner, fn) abs, log, log2, log10, log1p, exp, sqrt
NormExpr, SurvivalExpr, LogisticExpr, ClipExpr, AggExpr Gaussian CDF / survival, logistic, clip, aggregate
Comparison(left, op, right) <=, >=, <, >, ==, != → boolean Series
BoolOp(op, children) &, \|, ~ over boolean children

Boolean and numeric nodes compose freely — (Affinity <= 500) * Affinity.score is valid.

wt (scope prefix)

Wildtype scope prefix. Reads wt_* columns for ranking expressions. For predict_from_fragments() and variant-derived predictions, TopiaryPredictor(predict_wt=True) populates WT prediction columns by scoring wt_peptide with the configured MHC model(s):

# Python API (capitalized kind names)
wt.Affinity.value                 # reads wt_value
wt.Affinity.score                 # reads wt_score
wt.Affinity["netmhcpan"].score    # qualified WT

# String DSL (lowercase kind names)
# wt.affinity.value
# wt.affinity.score
# wt.affinity["netmhcpan"].score

For ranking expressions only (not filters). Returns NaN when WT columns are absent or the row has no length-compatible WT peptide.

shuffled and self (reserved scope prefixes)

Two more column-prefix scopes, reserved for values a producer computes elsewhere. shuffled.affinity reads shuffled_value, self.affinity.score reads self_score, and so on, exactly as wt. reads wt_*:

from topiary import Column, apply_filter, parse

# Keep peptides that bind better than their shuffled decoy.
parse("affinity.value < shuffled.affinity.value")

Topiary never populates these columns. Supply them on the frame and the scope reads them; leave them out and every expression under the scope evaluates to NaN, which filters to False and sorts as a missing key. That is the difference from wt. (populated by predict_wt=True) and from self_nearest. (populated by SelfProteome / predict_self_nearest).

self. is not nearest-self: it is a whole separate prefix kept for a producer's own definition of a self-match. Use self_nearest. for the columns Topiary computes.

len and count()

Peptide-level expressions that compose with scope prefixes:

# String DSL
len                           # peptide length (reads peptide_length column)
count('C')                    # cysteine count (reads from peptide column)
wt.len                        # wildtype peptide length (reads wt_peptide_length)
wt.count('C')                 # wildtype cysteine count (reads from wt_peptide)

Expr transforms

Method Description
.ascending_cdf(mean, std) Gaussian CDF: higher input → higher output. Alias: .norm()
.descending_cdf(mean, std) 1-CDF: lower input → higher output (for IC50, rank)
.logistic(midpoint, width) Logistic sigmoid: 1 / (1 + exp((x - midpoint) / width))
.clip(lo, hi) Clamp to range
.hinge() max(0, x) — zeroes out negative values
.log() / .log2() / .log10() Logarithm
.log1p() log(1 + x), accurate for small x
.exp() Exponential
.sqrt() Square root
abs(expr) Absolute value
expr ** n Power
+, -, *, / Arithmetic between expressions and scalars

Aggregation functions

Combine multiple expressions, skipping NaN values:

Function Description
mean(a, b, ...) Arithmetic mean
geomean(a, b, ...) Geometric mean (skips non-positive)
minimum(a, b, ...) Minimum value
maximum(a, b, ...) Maximum value
median(a, b, ...) Median (mean of middle two for even count)
from topiary import mean, geomean, minimum

# Average binding across models
mean(Affinity["netmhcpan"].logistic(350, 150),
     Affinity["mhcflurry"].logistic(350, 150))

# Best binding across models
minimum(Affinity["netmhcpan"].value, Affinity["mhcflurry"].value)

Building the long form

Callers holding mhctools.Prediction objects — a report reader, a cache, anything that didn't run a TopiaryPredictor end to end — should build the frame with from_predictions rather than by hand:

from topiary import from_predictions

df = from_predictions(predictions)                       # Prediction objects
df = from_predictions(model.predict_dataframe(peptides)) # or mhctools' own rows

df = from_predictions(
    predictions,
    allele_set={"pMHC_presentation": patient_alleles},   # genotype-level kinds
)

Rows come back in input order, one per prediction, which is what lets positional data line up:

df = from_predictions(
    predictions,
    extra_columns={"prediction_id": ids, "peptide_offset": offsets},
    allele_set=lambda prediction: attribution_for(prediction),
)

extra_columns carries identity a prediction doesn't itself have — a provenance key the consumer groups by, or an offset belonging to the candidate a peptide came from rather than to the prediction. A scalar fills the column; a sequence is positional and must match the input's length. allele_set also takes a callable, for attribution decided per peptide, where two peptides in one frame legitimately get different sets for the same kind.

It renames to topiary's column vocabulary, fills the context columns a producer may not have set, derives peptide_length, and backfills affinity for affinity rows — the same normalization TopiaryPredictor applies to its own output, so both paths produce one long form. A hand-written builder is a copy of that schema topiary can't see and can't migrate: a column added here never reaches it.

Filter expressions

Expression Creates
Affinity <= 500 Comparison(Field(pMHC_affinity, "value"), <=, 500)
Affinity.rank <= 2 Comparison(Field(pMHC_affinity, "percentile_rank"), <=, 2)
Affinity.score >= 0.5 Comparison(Field(pMHC_affinity, "score"), >=, 0.5)
Column("x") <= 2 Comparison(Column("x"), <=, 2)
(A) \| (B) BoolOp(\|, [A, B])
(A) & (B) BoolOp(&, [A, B])

Apply with apply_filter(df, node) and apply_sort(df, [nodes]).

Every argument that takes an expression -- apply_filter, apply_sort, evaluate_scores, rank_candidates, TopiaryResult.filter_by / sort_by and TopiaryPredictor(filter_by=..., sort_by=...) -- accepts the same three forms: a DSL string, a node, or a kind accessor such as Affinity, which stands for its .value. Sort arguments also take a list, primary key first. They all route through as_dsl_node / as_dsl_nodes, which you can call directly to validate an expression before using it. Anything else raises TypeError; a bool gets a message pointing at .eq(), since it is almost always Column("x") == "y" (see below).

Context options

For portable policy definitions, SelectionPolicy stores explicit filter/score strings and model/version selections. resolve_selection_policy(mapping) fills authoring defaults after a consumer has composed its overrides; SelectionPolicy.from_dict(mapping) strictly reloads a complete definition. to_dict() and sha256 expose the effective settings and content identity. read_selection_policy(path) / write_selection_policy(policy, path) provide JSON IO; writes refuse existing files. rank_with_policy(result, policy, provenance=None) wraps rank_candidates and returns a TopiaryResult carrying the effective definition and separate provenance/execution metadata. See the policy workflow and consumer boundary.

apply_filter, apply_sort and evaluate_scores share five keyword-only options, all forwarded to EvalContext:

Option Meaning
group_keys Explicit group identity columns; default infers them from the DataFrame
default_methods Per-kind default prediction_method_name for unqualified references
default_versions Per-(kind, model) default predictor_version, when one model has several
kind_support Per-(model, kind) metadata from TopiaryPredictor.kind_support
alleles Alleles peptides are evaluated against, so allele-free evidence reaches a genotype. Sequence (all peptides), mapping, or callable (per peptide)

See Group identity for when to pass group_keys.

All three also accept context=, a prebuilt EvalContext used in place of the five options above — see Sharing a context.

EvalContext attributes

What a node reads off the ctx handed to its eval() — see Writing your own node:

Attribute Meaning
ctx.df The prediction rows to group. Identity keys are normalized, so a groupby(ctx.group_keys) over it cannot produce a key group_index lacks
ctx.group_keys The identity columns, in order
ctx.group_index Unique group keys — a flat Index for one key, a MultiIndex for several. Every eval() returns a Series indexed by this
ctx.key_frame Just the group-key columns of ctx.df
ctx.row_group_codes() Each row's position in group_index, for mapping a per-group result back onto rows
ctx.row_group_tuples() Each row's group key, aligned to ctx.df.index
ctx.empty_series(fill) A group_index-shaped Series, for a node with nothing to say
ctx.is_built_on(df) Whether this context was built on df itself — the identity check apply_* make before accepting a context=
ctx.derive(**opts) A context on the same frame with options changed, reusing the grouping

String parsing

Kind aliases

Alias Kind
affinity, ba, aff, ic50 pMHC_affinity
presentation, el pMHC_presentation
stability pMHC_stability
processing, antigen_processing antigen_processing

Tool-qualified kinds

Prefer colon syntax: netmhcpan:affinity, mhcflurry:el, netmhcpan:ba. The older underscore aliases also work: netmhcpan_affinity, mhcflurry_el, netmhcpan_ba.

parse

The unified DSL parser. Returns a DSLNode.

String Result
"affinity <= 500" Comparison(Field(pMHC_affinity, "value"), <=, 500)
"netmhcpan_ba <= 500" Comparison(Field(pMHC_affinity, "value", method="netmhcpan"), <=, 500)
"netmhcpan:affinity <= 500" Comparison(Field(pMHC_affinity, "value", method="netmhcpan"), <=, 500)
"affinity:netmhcpan <= 500" Comparison(Field(pMHC_affinity, "value", method="netmhcpan"), <=, 500)
"netmhcpan.affinity <= 500" Comparison(Field(pMHC_affinity, "value", method="netmhcpan"), <=, 500)
"netmhcpan[4.1b]:affinity.score" Field(pMHC_affinity, "score", method="netmhcpan", version="4.1b")
"netmhcpan-4.1b:affinity.score" Field(pMHC_affinity, "score", method="netmhcpan", version="4.1b")
"netmhcpan[release-2.2.0]:affinity.score" Field(pMHC_affinity, "score", method="netmhcpan", version="release-2.2.0")
"affinity['netmhcpan', '4.1b'].value <= 500" Comparison(Field(..., method="netmhcpan", version="4.1b"), <=, 500)
"column(cysteine_count) <= 2" Comparison(Column("cysteine_count"), <=, 2)
"a <= 500 \| b <= 2" BoolOp(\|, [Comparison(a), Comparison(b)])
"a <= 500 & b <= 2" BoolOp(&, [Comparison(a), Comparison(b)])

Mixing | and & follows standard precedence (& binds tighter than |); use parentheses for the other grouping.

Input functions

Function Input Returns
read_fasta(path) FASTA file {name: sequence}
read_peptide_fasta(path) FASTA of peptides {name: peptide}
read_peptide_csv(path) CSV with peptide col {name: peptide}
read_sequence_csv(path) CSV with sequence col {name: sequence}
read_tsv(path) / read_csv(path) Topiary-format table with comment-block metadata TopiaryResult
read_lens(path, binding_metrics=None) LENS report (v1.4 / v1.5.1 / v1.9) TopiaryResult (wide form)
read_pvacseq(path) pVACseq aggregated or all_epitopes TSV (MHC-I or MHC-II) TopiaryResult (long form)
read_isovar_hypotheses(data, tag=None) Isovar hypothesis JSON export or mapping TopiaryResult of comparison observations; no generated candidates
normalize_isovar_rna_support(support, evidence_sets=None, evidence_scope=None) Legacy or current Isovar RNA support mapping Validated copy with canonical read fields, optional label counts/completeness, and preserved provenance
melt_pvacseq_algorithms(result) Loaded pVACseq all_epitopes result TopiaryResult with one row per (peptide, allele, algorithm)
parse_prediction_metric(model_name, metric_name) External predictor and metric labels PredictionMetric(method, kind, field, sequence) or None when ambiguous
derive_mhc_class(allele_series) Allele Series (mhcgnomes-normalized or raw) Series of "I" / "II" / pd.NA
slice_regions(seqs, regions) Sequences + intervals {name:start-end: subseq}
exclude_by(df, ref, mode) DataFrame + ref sequences Filtered DataFrame

Custom metadata in TSV and CSV

TopiaryResult.extra stores custom metadata. Put dataset fields under a custom key, for example extra={"dataset": {"source": "sha256 digest", "form": "reads"}}. Nested dictionaries and lists use the existing JSON comment encoding. Strings containing line breaks, surrounding whitespace or a json: prefix also use that encoding, as JSON strings, so their exact text and type survive writing and reading. Text under kind_support is always quoted to distinguish it from legacy dictionary syntax. Ordinary text keeps its readable form; existing files, including old structured kind_support comments, remain readable. Non-JSON objects retain their fallback text representation, with the same quoting when needed.

Writers reject top-level extra keys topiary_version, form, source, filter_by, sort_by, topiary_flank_encoding, topiary_text_encoding, topiary_json_encoding, and any key starting with model:. These comment keys belong to the result's built-in metadata; using them in extra previously changed their meaning on read. Set the corresponding metadata field when that is intended, or nest the value under a custom key.

Extra keys must be nonempty strings without surrounding whitespace, line breaks or =. Invalid keys raise ValueError before the output file is opened, including when writing over an existing file. TSV and CSV use the same checks; reading existing files retains the previous behavior.

The writer derives topiary_flank_encoding automatically to preserve empty terminal flanks separately from unknown context, and topiary_json_encoding so columns of dicts and lists such as measurement_context read back as the same values rather than as their text. See saving and reloading flanking context for the format and the limitations of legacy blank cells.

Correcting a LENS binding-column mapping

read_lens first applies the same prediction vocabulary as read_pvacseq to new predictor-shaped columns. EL means presentation, BA / Aff / Affinity mean affinity, and IM means immunogenicity; explicit quantity words take precedence. It warns when the remaining tool/metric pair is still ambiguous. Pass binding_metrics to close that gap without waiting for a topiary release:

read_lens(path, binding_metrics={
    ("newtool", "opaque_metric"): ("affinity", "value"),
    ("sometool", "noisy_metric"): None,   # not a prediction
})

Overrides merge over the built-in table, so one column can be patched without restating the rest, and a built-in mapping can be corrected as well as a missing one added.

Two columns of one tool cannot share a (kind, field) — that would put a duplicate column in the frame, making one set of values unreachable and to_long() fail. A mapping that would do it is refused, naming both source columns; map one to a different field, or to None.

Keys are (tool, metric) — the pair the warning itself names — and carry no version, so one entry covers a tool however a file spells its release. Values are (kind, field) with field one of value / score / rank, validated up front; None declares the column a non-prediction, silencing the warning while leaving the column in place as an annotation column.

LENS's normalized wide schema describes the mutant peptide and has no WT-scoped prediction fields. A newly inferred WT-specific column is therefore left under its original name with a warning; Topiary does not relabel it as an MT value.

Source functions

Function Returns
sequences_from_gene_names(names) {GENE\|TRANSCRIPT: seq}
sequences_from_gene_ids(ids) {GENE\|TRANSCRIPT: seq}
sequences_from_transcript_ids(ids) {GENE\|TRANSCRIPT: seq}
sequences_from_transcript_names(names) {GENE\|TRANSCRIPT: seq}
tissue_expressed_sequences(tissues) {GENE\|TRANSCRIPT: seq}
tissue_expressed_gene_ids(tissues) set of Ensembl gene IDs
cta_gene_ids(source, tier) set of CTA Ensembl gene IDs
cta_sequences() CTA protein sequences
non_cta_sequences() Non-CTA protein sequences
ensembl_proteome() All Ensembl proteins
available_tissues() List of tissue names

Which genes count as cancer-testis antigens

cta_gene_ids() is the one implementation, used by cta_sequences(), non_cta_sequences() and SelfProteome's include="non_cta" scope, so they cannot disagree about which genes leave a self proteome.

oncoref is the authority. source="pirlygenes" (the human default) and "tsarina" re-export oncoref's own functions — is identity, not a copy — so all three spellings select the same table, and a proteome's reference_version records the oncoref version rather than the shim it was reached through.

tier= chooses the membership, because the spread is wide enough that the choice belongs to the caller:

Tier Genes Meaning
default 293 oncoref's canonical CTA set
filtered 302 before the audited demotions
unfiltered 390 the full candidate universe
testis_restricted 248 restriction subset
placental_restricted 11 restriction subset

oncoref's complements (excluded, never_expressed) and its clinical-target list are deliberately not accepted: passing one where a CTA definition is expected would exclude the wrong genes from a self proteome. Counts are from oncoref 1.8.194 and move with its table; the gene IDs, not the counts, are the contract.

Peptide properties

from topiary.properties import add_peptide_properties, available_properties

# Add properties to predictions DataFrame
df = add_peptide_properties(df)                              # all
df = add_peptide_properties(df, groups=["core"])             # named group
df = add_peptide_properties(df, include=["charge"])          # specific
df = add_peptide_properties(df, peptide_column="wt_peptide", prefix="wt_")

# List available properties and their groups
available_properties()

Groups: "core", "manufacturability", "immunogenicity". See Peptide Properties for details.

Occurrence policy and criterion evaluation

  • evaluate_selection_policy(result, policy, group_keys=..., alleles=..., source_contexts=...) returns PolicyEvaluation with complete evidence, occurrences, selected and criterion audit views.
  • replay_selection_policy(evidence) repeats evaluation from the embedded definition and materialized runtime context without prediction.
  • select_policy_representatives(evaluation, candidate_keys=..., strata=...) selects actual occurrences and links alternatives without adding support.
  • describe_evaluation_context(context) materializes grouping, model choices and per-occurrence genotype declarations for replay.
  • candidate_identifier(sample, peptide, allele) shares combined/projected pMHC identity, using mhcgnomes for allele canonicalization.
  • SelectionCriterion, RankingTerm, and resolve_selection_expression define and expand named DSL criteria with role/cycle validation.
  • evaluate_selection_criteria(policy, context, references=...) exposes named measurements and reasons; evaluate_filter(frame, expression, context=...) exposes group decisions without discarding rows.

See the complete source-policy workflow for schema compatibility, source-local defaults, unknown handling and persistence.