API Reference¶
Native Exacto input¶
read_exacto(path, *, sample_name, schema=None, primary_structures=None,
reference_name=None, tag=None, transcript_read_support=None, library_id=None,
read_set_id=None) reads native sequence/evidence tables without prediction.
read_exacto_fragments(path, **kwargs) uses the same validation and returns
optional fragments for explicit new-window scanning. EXACTO_SCHEMAS and
EXACTO_SCHEMA_COMMIT describe the tested producer contract. See Reading
Exacto for supported files, evidence scope and composition examples.
Fragment identity and RNA outcomes¶
A fragment record is identified by (sample_name, fragment_id): the ID
names the candidate, ProteinFragment.sample_name names the observation.
unique_fragments(fragments) returns records in first-occurrence order,
coalescing a repeated pair only when its content agrees once passed through
fragment IO (a record and its saved copy agree). Conflicting or uncomparable
repeats, and one ID naming different candidates (CANDIDATE_FIELDS) in
different samples, raise ValueError naming the differing fields. Prediction
uses the same validation before model execution.
fragments_for_sample(fragments, sample_name) labels fragments from any
source; fragments_from_variants(..., sample_name=) and a frame's
sample_name column do the same. Prediction writes the label to the
sample_name column, so the default aggregate_evidence_across_samples
keys pool one candidate across samples. make_fragment_id(...,
qualifiers=[...]) hashes any further values a producer groups records by.
describe_isovar_result(result) returns a JSON-compatible outcome dictionary
for a completed Isovar result: variant, separate native read/template counts,
protein and edit interval, transcript IDs, filter values and failed-filter
names. Missing counts stay null. Status is passing, filtered,
filter_status_unavailable, no_usable_reads, no_alt_reads,
no_predicted_coding_change or no_protein_sequence. This reports recorded
filters, not clinical suitability. Acquisition failures are not biological
results. See the all-variant audit.
Amino-acid data¶
Topiary exposes one encoding shared by sequence consumers:
from topiary import (
AMINO_ACIDS,
AMINO_ACID_INDEX,
BLOSUM62_AMINO_ACIDS,
ENCODED_AMINO_ACIDS,
UNKNOWN_AMINO_ACID_INDEX,
blosum62_distance_matrix,
blosum62_matrix,
encode_amino_acids,
)
encoded = encode_amino_acids(["SIINFEKL", "AXX"], length=8)
scores = blosum62_matrix()
distances = blosum62_distance_matrix()
AMINO_ACIDS gives the canonical order, BLOSUM62_AMINO_ACIDS adds NCBI's
B/J/Z ambiguity rows, and ENCODED_AMINO_ACIDS adds O/U/X/*. The immutable
AMINO_ACID_INDEX mapping follows that order; any other character uses
UNKNOWN_AMINO_ACID_INDEX.
blosum62_matrix() returns substitution scores. blosum62_distance_matrix()
returns the distances used by SelfProteome: canonical pairs retain Topiary's
established behavior, B/J/Z comparisons use their published scores with a
symmetric conservative transformation, and O/U/X/* or unrecognized characters
receive the symmetric worst-case canonical distance (15). Both functions
return fresh read-only int8 arrays, so callers never share mutable matrix state.
Self-match evidence¶
match_self_peptides(reference, peptides, *, max_mismatches=0, alleles=None,
excluded_gene_ids=(), observations=None, predictions=None) returns every
same-length Hamming match within the radius, with all source origins and
supplied presentation/prediction records. reference.match_candidates(...)
delegates to it. self_matches_in_windows(windows, reference, *, peptide_lengths,
...) searches exact peptides at every requested window position. Results use
ordinary typed CSV/TSV, with explicit coverage and evidence identity.
See self evidence for the input schemas,
policy composition and interpretation of missing results.
TopiaryPredictor¶
| Parameter | Type | Description |
|---|---|---|
models |
class, instance, or list | Predictor model(s). Classes require alleles. |
alleles |
list of str | HLA alleles. Used to construct model classes. |
filter_by |
DSLNode or str | Boolean filter expression (e.g. Affinity <= 500 or "affinity <= 500 \| el.rank <= 2"). |
sort_by |
DSLNode or list of DSLNode | Sort expression(s). Lexicographic tiebreakers; NaN falls through. |
sort_direction |
"auto", "asc", or "desc" | Direction for sort keys (auto infers per-key). |
padding_around_mutation |
int | Residues around mutation for candidate epitopes. |
only_novel_epitopes |
bool | Keep known mutation/target/junction overlaps; exclude unknown geometry. Does not establish tumor specificity. |
min_gene_expression |
float | Minimum gene FPKM (variant inputs). |
min_transcript_expression |
float | Minimum transcript FPKM (variant inputs). |
raise_on_error |
bool | Raise on variant-effect errors vs. skip. |
Prediction methods¶
| Method | Input | Behavior |
|---|---|---|
predict_from_named_sequences(dict) |
{name: sequence} |
Sliding-window scan |
predict_from_named_peptides(dict) |
{name: peptide} |
Score as-is |
predict_from_peptide_occurrences(table) |
prediction_id, peptide, optional flanks/coordinates/annotations; unique by ID, peptide, offset and sample |
Score exact occurrences with their context; no scanning. |
predict_from_sequences(list) |
[sequence, ...] |
Sliding-window scan |
predict_from_fragments(fragments) |
[ProteinFragment] |
Universal path — any origin; fragment-level metadata and target_intervals threaded through. |
predict_from_variants(variants) |
VariantCollection | Variant pipeline (builds ProteinFragments internally and delegates). |
predict_from_mutation_effects(effects) |
EffectCollection | Same as predict_from_variants but starting from pre-computed effects. |
predict_peptide_occurrences(table, model, use_flanks=True, predict_wt=False)
exposes the same occurrence prediction for one configured mhctools model.
It is exported from topiary and is also used by rescore_candidates.
See exact occurrence prediction
for input identity, unknown flanks, comparator and coverage rules.
prediction_mhc_scope(allele, dependence=..., allele_set=...) returns a canonical
MHC identity tuple: one allele, the full genotype, or an allele-free scope.
Missing required identity returns None, which must remain unmatchable.
Comparator joins use this public rule in both fragment and occurrence prediction;
a haplotype's selected presenter is not its genotype identity.
TopiaryResult¶
TopiaryResult is the semantic result object for Topiary prediction tables.
It can ingest either long or wide prediction tables, keeps the active df
view for pandas-style access, and materializes cached long_df and wide_df
views on demand. The public form value describes the active df view; it is
derived from the stored views rather than acting as separate result state. Use
result.to_long() / result.to_wide() when you want a new TopiaryResult
whose active df is that form.
Wide prediction columns use {model}_{kind}_{value|score|rank}. Wildtype
companions use {model}_{kind}_wt_{value|score|rank} and round-trip with their
MT prediction row; WT method and version fields do not create additional wide
rows.
Topiary has two result-merging operations:
| Operation | Meaning | Use when |
|---|---|---|
topiary.stack_results([a, b]) / a.stack_with(b) |
More result sets | Inputs are separate files, samples, cohorts, or independent result sets. |
topiary.combine_predictions([a, b]) / a.combine_predictions(b) |
More predictions for the same logical identity grid | Inputs are separate predictors or predictor shards that should behave like one run. |
stack_results is a row-union operation. The inputs do not need to describe
the same peptides, alleles, samples, sources, predictors, or score kinds; they
are just more Topiary result rows with merged provenance.
combine_predictions is a prediction-union operation. The inputs are pieces
of one logical prediction table: separate predictors over the same candidates,
or one predictor split into disjoint allele/peptide-length shards. It rejects
duplicate (prediction_method_name, kind, identity) rows and, by default,
requires every emitted (prediction_method_name, kind) group to cover the
same identity grid. Use coverage="partial" only for intentionally sparse
prediction unions.
Both operations accept mixed long/wide TopiaryResult inputs and normalize
internally. Fresh TopiaryPredictor DataFrame outputs may also be passed to
combine_predictions; they are first wrapped as TopiaryResult objects and
then validated with the same rules.
Cross-sample evidence¶
topiary.aggregate_evidence_across_samples(df, group_keys=None) returns one
pooled row per prediction identity while leaving the input's per-sample rows
unchanged. It collapses repeated prediction rows within a sample before
summing canonical RNA and DNA counts, recomputes VAF from pooled alternate and
overlapping counts, and reports the number of represented samples as
n_samples.
A count is omitted unless every represented sample states it. Poolable counts
must share an evidence subject across samples; allele-support counts must also
share a derivation method. Coverage-only depths need no invented method.
Incompatible units or methods raise rather than being mixed. Expression is not
aggregated. By default the standard Topiary candidate identity is inferred, or
callers can supply explicit group_keys. These exclude sample_name and the
canonical counts/VAFs, which are aggregation values rather than candidate
identity.
CachedPredictor¶
Drop-in replacement for a live predictor that serves scores from a
pre-computed table. Pass as models=cache to TopiaryPredictor. See
Cached Predictions for full detail.
| Loader | Source |
|---|---|
CachedPredictor.from_topiary_output(path) |
Parquet / TSV / CSV previously written from a topiary run. |
CachedPredictor.from_mhcflurry(path) |
mhcflurry-predict CSV output. predictor_version auto-composed from the installed mhcflurry when omitted. |
CachedPredictor.from_netmhcpan_stdout(path) |
NetMHCpan stdout capture (auto-detects 2.8 / 3 / 4 / 4.1). Returns every kind present in the output — -BA runs surface both pMHC_affinity and pMHC_presentation rows per (peptide, allele). |
CachedPredictor.from_netmhc_stdout(path, version=...) |
Classic NetMHC stdout (3 / 4 / 4.1). |
CachedPredictor.from_netmhcpan_cons_stdout(path) |
NetMHCcons stdout. |
CachedPredictor.from_netmhciipan_stdout(path, version=...) |
NetMHCIIpan stdout (legacy / 4 / 4.3). |
CachedPredictor.from_netmhcstabpan_stdout(path) |
NetMHCstabpan stdout (pMHC stability). |
CachedPredictor.concat([caches], on_overlap=...) |
Merge shards (all must share name+version). on_overlap: "raise" / "last" / "first" / callable. |
CachedPredictor.from_directory(path, pattern="*", on_overlap=...) |
Glob a dir and concat every matching file. |
CachedPredictor.from_tsv(path, columns=..., prediction_method_name=..., predictor_version=...) |
Generic tab- or comma-delimited. |
CachedPredictor.from_dataframe(df, ...) |
In-memory DataFrame. |
CachedPredictor(fallback=live_predictor) |
Empty cache, lazy identity discovery — pure read-through over a live model. |
Constructor-level knobs:
| Parameter | Description |
|---|---|
fallback |
Live predictor to route misses through. Result is merged back into the cache. |
also_accept_versions |
Set of version strings treated as interchangeable with the cache's own version (opt-in equivalence for rc → final, etc.). |
Helper: topiary.mhcflurry_composite_version() returns
"<package_version>+release-<model_release>" for the locally-installed
mhcflurry. Automatically used by from_mhcflurry when no explicit
predictor_version is passed.
Kind accessors¶
| Accessor | Kind | Default field |
|---|---|---|
Affinity |
pMHC_affinity |
.value (IC50 nM) |
Presentation |
pMHC_presentation |
.value |
Stability |
pMHC_stability |
.value |
Processing |
antigen_processing |
.value |
Fields¶
| Field | Column read | Description |
|---|---|---|
.value |
value |
Raw prediction value |
.rank |
percentile_rank |
Percentile rank (lower = better) |
.score |
score |
Normalized score (higher = better) |
Multi-model bracket syntax¶
Affinity["netmhcpan"].value # filter to netmhcpan rows
Affinity["mhcflurry"].score # filter to mhcflurry rows
Method matching is case-insensitive and substring-based. Errors with "Did you mean" on typos.
Column¶
Reference any DataFrame column in expressions:
Column("charge") # reads 'charge' column
Column("cysteine_count") <= 2 # returns a Comparison (boolean DSL node)
Errors with close-match suggestions when column doesn't exist. Raises TypeError for non-numeric columns in arithmetic / ordered comparisons.
Categorical equality and membership¶
For string / boolean / mixed-dtype columns, Column.eq / Column.ne / Column.isin produce an IsIn node that reads the column raw (no float cast):
| Form | Behavior |
|---|---|
Column("mhc_class").eq("I") |
True for rows whose mhc_class equals "I" |
Column("mhc_class").ne("II") |
True for mhc_class != "II" (NaN survives, pandas-style) |
Column("mhc_class").isin(["I", "II"]) |
True for either value |
~Column("mhc_class").isin([...]) |
Negate via ~ |
Column("x") == "y" is Python equality, not one of these: DSLNode keeps
ordinary ==, so it evaluates to a plain bool. Every expression argument
refuses a bool with a TypeError that names .eq(), rather than
filtering or sorting on a constant.
In the string DSL, string literals are legal on the RHS of == / != only:
parse('mhc_class == "I"')
parse('affinity.value <= 500 & mhc_class != "II"')
Pre-built shortcuts¶
| Symbol | Definition |
|---|---|
class_i |
IsIn("mhc_class", ["I"]) |
class_ii |
IsIn("mhc_class", ["II"]) |
Both require a mhc_class column. read_pvacseq() produces one; for fresh TopiaryPredictor results, derive it first:
from topiary import derive_mhc_class
df["mhc_class"] = derive_mhc_class(df["allele"])
DSL tree¶
Every DSL expression is a DSLNode. The tree is composed of:
| Leaf | Produces |
|---|---|
Const(v) |
Constant scalar |
Column(name) |
Per-group first-row of that column |
IsIn(name, values, negate=False) |
Per-group boolean membership test against scalar values (no float cast); reachable via Column.eq / Column.ne / Column.isin |
Field(kind, field, method=None, version=None, scope="") |
Per-group first-row of {scope}{field} for rows matching kind (and method / version if given) |
Len(scope="") |
Peptide length |
Count(chars, scope="") |
Amino-acid character count |
| Composite | Purpose |
|---|---|
BinOp(left, right, op) |
Elementwise +, -, *, /, ** |
UnaryOp(inner, fn) |
abs, log, log2, log10, log1p, exp, sqrt |
NormExpr, SurvivalExpr, LogisticExpr, ClipExpr, AggExpr |
Gaussian CDF / survival, logistic, clip, aggregate |
Comparison(left, op, right) |
<=, >=, <, >, ==, != → boolean Series |
BoolOp(op, children) |
&, \|, ~ over boolean children |
Boolean and numeric nodes compose freely — (Affinity <= 500) * Affinity.score is valid.
wt (scope prefix)¶
Wildtype scope prefix. Reads wt_* columns for ranking expressions. For predict_from_fragments() and variant-derived predictions, TopiaryPredictor(predict_wt=True) populates WT prediction columns by scoring wt_peptide with the configured MHC model(s):
# Python API (capitalized kind names)
wt.Affinity.value # reads wt_value
wt.Affinity.score # reads wt_score
wt.Affinity["netmhcpan"].score # qualified WT
# String DSL (lowercase kind names)
# wt.affinity.value
# wt.affinity.score
# wt.affinity["netmhcpan"].score
For ranking expressions only (not filters). Returns NaN when WT columns are absent or the row has no length-compatible WT peptide.
shuffled and self (reserved scope prefixes)¶
Two more column-prefix scopes, reserved for values a producer computes
elsewhere. shuffled.affinity reads shuffled_value, self.affinity.score
reads self_score, and so on, exactly as wt. reads wt_*:
from topiary import Column, apply_filter, parse
# Keep peptides that bind better than their shuffled decoy.
parse("affinity.value < shuffled.affinity.value")
Topiary never populates these columns. Supply them on the frame and the scope
reads them; leave them out and every expression under the scope evaluates to
NaN, which filters to False and sorts as a missing key. That is the
difference from wt. (populated by predict_wt=True) and from
self_nearest. (populated by SelfProteome / predict_self_nearest).
self. is not nearest-self: it is a whole separate prefix kept for a
producer's own definition of a self-match. Use self_nearest. for the
columns Topiary computes.
len and count()¶
Peptide-level expressions that compose with scope prefixes:
# String DSL
len # peptide length (reads peptide_length column)
count('C') # cysteine count (reads from peptide column)
wt.len # wildtype peptide length (reads wt_peptide_length)
wt.count('C') # wildtype cysteine count (reads from wt_peptide)
Expr transforms¶
| Method | Description |
|---|---|
.ascending_cdf(mean, std) |
Gaussian CDF: higher input → higher output. Alias: .norm() |
.descending_cdf(mean, std) |
1-CDF: lower input → higher output (for IC50, rank) |
.logistic(midpoint, width) |
Logistic sigmoid: 1 / (1 + exp((x - midpoint) / width)) |
.clip(lo, hi) |
Clamp to range |
.hinge() |
max(0, x) — zeroes out negative values |
.log() / .log2() / .log10() |
Logarithm |
.log1p() |
log(1 + x), accurate for small x |
.exp() |
Exponential |
.sqrt() |
Square root |
abs(expr) |
Absolute value |
expr ** n |
Power |
+, -, *, / |
Arithmetic between expressions and scalars |
Aggregation functions¶
Combine multiple expressions, skipping NaN values:
| Function | Description |
|---|---|
mean(a, b, ...) |
Arithmetic mean |
geomean(a, b, ...) |
Geometric mean (skips non-positive) |
minimum(a, b, ...) |
Minimum value |
maximum(a, b, ...) |
Maximum value |
median(a, b, ...) |
Median (mean of middle two for even count) |
from topiary import mean, geomean, minimum
# Average binding across models
mean(Affinity["netmhcpan"].logistic(350, 150),
Affinity["mhcflurry"].logistic(350, 150))
# Best binding across models
minimum(Affinity["netmhcpan"].value, Affinity["mhcflurry"].value)
Building the long form¶
Callers holding mhctools.Prediction objects — a report reader, a cache, anything that didn't run a TopiaryPredictor end to end — should build the frame with from_predictions rather than by hand:
from topiary import from_predictions
df = from_predictions(predictions) # Prediction objects
df = from_predictions(model.predict_dataframe(peptides)) # or mhctools' own rows
df = from_predictions(
predictions,
allele_set={"pMHC_presentation": patient_alleles}, # genotype-level kinds
)
Rows come back in input order, one per prediction, which is what lets positional data line up:
df = from_predictions(
predictions,
extra_columns={"prediction_id": ids, "peptide_offset": offsets},
allele_set=lambda prediction: attribution_for(prediction),
)
extra_columns carries identity a prediction doesn't itself have — a provenance key the consumer groups by, or an offset belonging to the candidate a peptide came from rather than to the prediction. A scalar fills the column; a sequence is positional and must match the input's length. allele_set also takes a callable, for attribution decided per peptide, where two peptides in one frame legitimately get different sets for the same kind.
It renames to topiary's column vocabulary, fills the context columns a producer may not have set, derives peptide_length, and backfills affinity for affinity rows — the same normalization TopiaryPredictor applies to its own output, so both paths produce one long form. A hand-written builder is a copy of that schema topiary can't see and can't migrate: a column added here never reaches it.
Filter expressions¶
| Expression | Creates |
|---|---|
Affinity <= 500 |
Comparison(Field(pMHC_affinity, "value"), <=, 500) |
Affinity.rank <= 2 |
Comparison(Field(pMHC_affinity, "percentile_rank"), <=, 2) |
Affinity.score >= 0.5 |
Comparison(Field(pMHC_affinity, "score"), >=, 0.5) |
Column("x") <= 2 |
Comparison(Column("x"), <=, 2) |
(A) \| (B) |
BoolOp(\|, [A, B]) |
(A) & (B) |
BoolOp(&, [A, B]) |
Apply with apply_filter(df, node) and apply_sort(df, [nodes]).
Every argument that takes an expression -- apply_filter, apply_sort,
evaluate_scores, rank_candidates, TopiaryResult.filter_by / sort_by
and TopiaryPredictor(filter_by=..., sort_by=...) -- accepts the same three
forms: a DSL string, a node, or a kind accessor such as Affinity, which
stands for its .value. Sort arguments also take a list, primary key first.
They all route through as_dsl_node / as_dsl_nodes, which you can call
directly to validate an expression before using it. Anything else raises
TypeError; a bool gets a message pointing at .eq(), since it is almost
always Column("x") == "y" (see below).
Context options¶
For portable policy definitions, SelectionPolicy stores explicit filter/score
strings and model/version selections. resolve_selection_policy(mapping) fills
authoring defaults after a consumer has composed its overrides;
SelectionPolicy.from_dict(mapping) strictly reloads a complete definition.
to_dict() and sha256 expose the effective settings and content identity.
read_selection_policy(path) / write_selection_policy(policy, path) provide
JSON IO; writes refuse existing files. rank_with_policy(result, policy,
provenance=None) wraps rank_candidates and returns a TopiaryResult carrying
the effective definition and separate provenance/execution metadata. See the
policy workflow and consumer boundary.
apply_filter, apply_sort and evaluate_scores share five keyword-only
options, all forwarded to EvalContext:
| Option | Meaning |
|---|---|
group_keys |
Explicit group identity columns; default infers them from the DataFrame |
default_methods |
Per-kind default prediction_method_name for unqualified references |
default_versions |
Per-(kind, model) default predictor_version, when one model has several |
kind_support |
Per-(model, kind) metadata from TopiaryPredictor.kind_support |
alleles |
Alleles peptides are evaluated against, so allele-free evidence reaches a genotype. Sequence (all peptides), mapping, or callable (per peptide) |
See Group identity for when to pass group_keys.
All three also accept context=, a prebuilt EvalContext used in place
of the five options above — see Sharing a
context.
EvalContext attributes¶
What a node reads off the ctx handed to its eval() — see Writing your
own node:
| Attribute | Meaning |
|---|---|
ctx.df |
The prediction rows to group. Identity keys are normalized, so a groupby(ctx.group_keys) over it cannot produce a key group_index lacks |
ctx.group_keys |
The identity columns, in order |
ctx.group_index |
Unique group keys — a flat Index for one key, a MultiIndex for several. Every eval() returns a Series indexed by this |
ctx.key_frame |
Just the group-key columns of ctx.df |
ctx.row_group_codes() |
Each row's position in group_index, for mapping a per-group result back onto rows |
ctx.row_group_tuples() |
Each row's group key, aligned to ctx.df.index |
ctx.empty_series(fill) |
A group_index-shaped Series, for a node with nothing to say |
ctx.is_built_on(df) |
Whether this context was built on df itself — the identity check apply_* make before accepting a context= |
ctx.derive(**opts) |
A context on the same frame with options changed, reusing the grouping |
String parsing¶
Kind aliases¶
| Alias | Kind |
|---|---|
affinity, ba, aff, ic50 |
pMHC_affinity |
presentation, el |
pMHC_presentation |
stability |
pMHC_stability |
processing, antigen_processing |
antigen_processing |
Tool-qualified kinds¶
Prefer colon syntax: netmhcpan:affinity, mhcflurry:el,
netmhcpan:ba. The older underscore aliases also work:
netmhcpan_affinity, mhcflurry_el, netmhcpan_ba.
parse¶
The unified DSL parser. Returns a DSLNode.
| String | Result |
|---|---|
"affinity <= 500" |
Comparison(Field(pMHC_affinity, "value"), <=, 500) |
"netmhcpan_ba <= 500" |
Comparison(Field(pMHC_affinity, "value", method="netmhcpan"), <=, 500) |
"netmhcpan:affinity <= 500" |
Comparison(Field(pMHC_affinity, "value", method="netmhcpan"), <=, 500) |
"affinity:netmhcpan <= 500" |
Comparison(Field(pMHC_affinity, "value", method="netmhcpan"), <=, 500) |
"netmhcpan.affinity <= 500" |
Comparison(Field(pMHC_affinity, "value", method="netmhcpan"), <=, 500) |
"netmhcpan[4.1b]:affinity.score" |
Field(pMHC_affinity, "score", method="netmhcpan", version="4.1b") |
"netmhcpan-4.1b:affinity.score" |
Field(pMHC_affinity, "score", method="netmhcpan", version="4.1b") |
"netmhcpan[release-2.2.0]:affinity.score" |
Field(pMHC_affinity, "score", method="netmhcpan", version="release-2.2.0") |
"affinity['netmhcpan', '4.1b'].value <= 500" |
Comparison(Field(..., method="netmhcpan", version="4.1b"), <=, 500) |
"column(cysteine_count) <= 2" |
Comparison(Column("cysteine_count"), <=, 2) |
"a <= 500 \| b <= 2" |
BoolOp(\|, [Comparison(a), Comparison(b)]) |
"a <= 500 & b <= 2" |
BoolOp(&, [Comparison(a), Comparison(b)]) |
Mixing | and & follows standard precedence (& binds tighter than |); use parentheses for the other grouping.
Input functions¶
| Function | Input | Returns |
|---|---|---|
read_fasta(path) |
FASTA file | {name: sequence} |
read_peptide_fasta(path) |
FASTA of peptides | {name: peptide} |
read_peptide_csv(path) |
CSV with peptide col |
{name: peptide} |
read_sequence_csv(path) |
CSV with sequence col |
{name: sequence} |
read_tsv(path) / read_csv(path) |
Topiary-format table with comment-block metadata | TopiaryResult |
read_lens(path, binding_metrics=None) |
LENS report (v1.4 / v1.5.1 / v1.9) | TopiaryResult (wide form) |
read_pvacseq(path) |
pVACseq aggregated or all_epitopes TSV (MHC-I or MHC-II) |
TopiaryResult (long form) |
read_isovar_hypotheses(data, tag=None) |
Isovar hypothesis JSON export or mapping | TopiaryResult of comparison observations; no generated candidates |
normalize_isovar_rna_support(support, evidence_sets=None, evidence_scope=None) |
Legacy or current Isovar RNA support mapping | Validated copy with canonical read fields, optional label counts/completeness, and preserved provenance |
melt_pvacseq_algorithms(result) |
Loaded pVACseq all_epitopes result |
TopiaryResult with one row per (peptide, allele, algorithm) |
parse_prediction_metric(model_name, metric_name) |
External predictor and metric labels | PredictionMetric(method, kind, field, sequence) or None when ambiguous |
derive_mhc_class(allele_series) |
Allele Series (mhcgnomes-normalized or raw) | Series of "I" / "II" / pd.NA |
slice_regions(seqs, regions) |
Sequences + intervals | {name:start-end: subseq} |
exclude_by(df, ref, mode) |
DataFrame + ref sequences | Filtered DataFrame |
Custom metadata in TSV and CSV¶
TopiaryResult.extra stores custom metadata. Put dataset fields under a custom
key, for example extra={"dataset": {"source": "sha256 digest", "form": "reads"}}.
Nested dictionaries and lists use the existing JSON comment encoding.
Strings containing line breaks, surrounding whitespace or a json: prefix
also use that encoding, as JSON strings, so their exact text and type survive
writing and reading. Text under kind_support is always quoted to distinguish
it from legacy dictionary syntax. Ordinary text keeps its readable form;
existing files, including old structured kind_support comments, remain
readable. Non-JSON objects retain their fallback text representation, with the
same quoting when needed.
Writers reject top-level extra keys topiary_version, form, source,
filter_by, sort_by, topiary_flank_encoding, topiary_text_encoding,
topiary_json_encoding, and any key starting with model:. These comment keys
belong to the result's built-in metadata; using them in extra previously
changed their meaning on read. Set the corresponding metadata field when
that is intended, or nest the value under a custom key.
Extra keys must be nonempty strings without surrounding whitespace, line
breaks or =. Invalid keys raise ValueError before the output file is opened,
including when writing over an existing file. TSV and CSV use the same checks;
reading existing files retains the previous behavior.
The writer derives topiary_flank_encoding automatically to preserve empty
terminal flanks separately from unknown context, and topiary_json_encoding
so columns of dicts and lists such as measurement_context read back as the
same values rather than as their text. See
saving and reloading flanking context
for the format and the limitations of legacy blank cells.
Correcting a LENS binding-column mapping¶
read_lens first applies the same prediction vocabulary as read_pvacseq to
new predictor-shaped columns. EL means presentation, BA / Aff /
Affinity mean affinity, and IM means immunogenicity; explicit quantity
words take precedence. It warns when the remaining tool/metric pair is still
ambiguous. Pass binding_metrics to close that gap without waiting for a
topiary release:
read_lens(path, binding_metrics={
("newtool", "opaque_metric"): ("affinity", "value"),
("sometool", "noisy_metric"): None, # not a prediction
})
Overrides merge over the built-in table, so one column can be patched without restating the rest, and a built-in mapping can be corrected as well as a missing one added.
Two columns of one tool cannot share a (kind, field) — that would put
a duplicate column in the frame, making one set of values unreachable
and to_long() fail. A mapping that would do it is refused, naming both
source columns; map one to a different field, or to None.
Keys are (tool, metric) — the pair the warning itself names — and
carry no version, so one entry covers a tool however a file spells its
release. Values are (kind, field) with field one of
value / score / rank, validated up front; None declares the
column a non-prediction, silencing the warning while leaving the column
in place as an annotation column.
LENS's normalized wide schema describes the mutant peptide and has no WT-scoped prediction fields. A newly inferred WT-specific column is therefore left under its original name with a warning; Topiary does not relabel it as an MT value.
Source functions¶
| Function | Returns |
|---|---|
sequences_from_gene_names(names) |
{GENE\|TRANSCRIPT: seq} |
sequences_from_gene_ids(ids) |
{GENE\|TRANSCRIPT: seq} |
sequences_from_transcript_ids(ids) |
{GENE\|TRANSCRIPT: seq} |
sequences_from_transcript_names(names) |
{GENE\|TRANSCRIPT: seq} |
tissue_expressed_sequences(tissues) |
{GENE\|TRANSCRIPT: seq} |
tissue_expressed_gene_ids(tissues) |
set of Ensembl gene IDs |
cta_gene_ids(source, tier) |
set of CTA Ensembl gene IDs |
cta_sequences() |
CTA protein sequences |
non_cta_sequences() |
Non-CTA protein sequences |
ensembl_proteome() |
All Ensembl proteins |
available_tissues() |
List of tissue names |
Which genes count as cancer-testis antigens¶
cta_gene_ids() is the one implementation, used by cta_sequences(),
non_cta_sequences() and SelfProteome's include="non_cta" scope, so they
cannot disagree about which genes leave a self proteome.
oncoref is the authority. source="pirlygenes" (the human default) and
"tsarina" re-export oncoref's own functions — is identity, not a copy — so
all three spellings select the same table, and a proteome's
reference_version records the oncoref version rather than the shim it was
reached through.
tier= chooses the membership, because the spread is wide enough that the
choice belongs to the caller:
| Tier | Genes | Meaning |
|---|---|---|
default |
293 | oncoref's canonical CTA set |
filtered |
302 | before the audited demotions |
unfiltered |
390 | the full candidate universe |
testis_restricted |
248 | restriction subset |
placental_restricted |
11 | restriction subset |
oncoref's complements (excluded, never_expressed) and its clinical-target
list are deliberately not accepted: passing one where a CTA definition is
expected would exclude the wrong genes from a self proteome. Counts are from
oncoref 1.8.194 and move with its table; the gene IDs, not the counts, are the
contract.
Peptide properties¶
from topiary.properties import add_peptide_properties, available_properties
# Add properties to predictions DataFrame
df = add_peptide_properties(df) # all
df = add_peptide_properties(df, groups=["core"]) # named group
df = add_peptide_properties(df, include=["charge"]) # specific
df = add_peptide_properties(df, peptide_column="wt_peptide", prefix="wt_")
# List available properties and their groups
available_properties()
Groups: "core", "manufacturability", "immunogenicity". See Peptide Properties for details.
Occurrence policy and criterion evaluation¶
evaluate_selection_policy(result, policy, group_keys=..., alleles=..., source_contexts=...)returnsPolicyEvaluationwith completeevidence,occurrences,selectedand criterionauditviews.replay_selection_policy(evidence)repeats evaluation from the embedded definition and materialized runtime context without prediction.select_policy_representatives(evaluation, candidate_keys=..., strata=...)selects actual occurrences and links alternatives without adding support.describe_evaluation_context(context)materializes grouping, model choices and per-occurrence genotype declarations for replay.candidate_identifier(sample, peptide, allele)shares combined/projected pMHC identity, using mhcgnomes for allele canonicalization.SelectionCriterion,RankingTerm, andresolve_selection_expressiondefine and expand named DSL criteria with role/cycle validation.evaluate_selection_criteria(policy, context, references=...)exposes named measurements and reasons;evaluate_filter(frame, expression, context=...)exposes group decisions without discarding rows.
See the complete source-policy workflow for schema compatibility, source-local defaults, unknown handling and persistence.