Combine source tables and re-score on demand¶
Topiary combines normalized source tables, preserves their original values, and uses the existing DSL to filter and rank reported candidates. Models run only when explicitly requested. Raw reads and protein reconstruction are not prerequisites for combining tables.
from topiary import (
combine_sources, read_lens, read_pvacseq, melt_pvacseq_algorithms,
rank_candidates,
)
combined = combine_sources({
"lens": read_lens("lens.tsv"),
"pvacseq": melt_pvacseq_algorithms(read_pvacseq("all_epitopes.tsv")),
}, sample_name="patient-01")
ranked = rank_candidates(
combined, "affinity['netmhcpan'].value", ascending=True,
duplicates="best",
)
per_source = rank_candidates(
combined, "affinity['netmhcpan'].value", ascending=True,
duplicates="best", strata=["source_label", "candidate_mhc_class"],
)
combined.to_tsv("all-evidence.tsv")
ranked.to_csv("ranked-candidates.tsv", sep="\t", index=False)
Inputs may be TopiaryResult objects or normalized DataFrames. Existing readers
normalize LENS and pVACseq; native Exacto tables use read_exacto.
Already normalized ORF/RNA DataFrames are also supported. Aggregated reports contribute only the
candidates they actually report.
A single input needs no model name, version or per-row source metadata. For example, this table can be ranked directly after combination:
import pandas as pd
from topiary import combine_sources, rank_candidates
simple = combine_sources({"input": pd.DataFrame({
"peptide": ["SIINFEKL"], "allele": ["HLA-A*02:01"],
"kind": ["pMHC_affinity"], "value": [50.0],
})}, sample_name="patient-01")
ranked = rank_candidates(simple, "affinity.value", ascending=True)
Unknown predictor names and versions remain unknown. The ordinary expression
affinity.value continues to work; adding provenance does not require a
version-qualified expression for this simple case.
Named selection policies¶
SelectionPolicy saves the exact candidate filter, score, model/version
selections, ranking direction, duplicate policy and strata under a stable name.
It has no implicit scientific recipe. Vaxrank owns the published
builtin:openvax-v1 bundle, shipped with experimental overlays in
Vaxrank 3.36.0 / PR #565.
Topiary evaluates its policy subtree and tests against that released bundle;
it does not maintain a second definition. Give changed policies new names.
from topiary import (
SelectionPolicy, read_selection_policy, write_selection_policy,
rank_with_policy, read_tsv,
)
# Illustrative settings, not a calibrated biological model.
policy = SelectionPolicy(
name="example-v1", score_by="1 / affinity.value",
filter_by="n_rna_alt >= 5", duplicates="best",
)
write_selection_policy(policy, "example-v1.json") # refuses to overwrite
combined.to_tsv("all-evidence.tsv")
ranked = rank_with_policy(combined, policy)
ranked.to_tsv("ranked.tsv")
replayed = rank_with_policy(
read_tsv("all-evidence.tsv"), read_selection_policy("example-v1.json"),
)
The immutable policy has to_dict() / from_dict() for complete, strict,
schema-versioned mappings and sha256 for the canonical definition, including
its name. Every default is saved explicitly. Changing an expression or model
selection changes the digest even if the name remains openvax-v1. Runtime
versions and input facts live separately in
ranked.extra['selection_policy']['execution']; authoring history can be passed
as provenance= and is stored alongside the definition, outside its digest.
Record actual base digests and ordered config-file hashes there when available.
rank_with_policy delegates to rank_candidates: it filters first, then scores
surviving evidence and chooses one representative per candidate/stratum. Unknown
filter results exclude groups under the existing DSL's evidence-retention
rules. Missing scores remain unranked; zero scores remain zero. It performs no
predictor calls, automatic method/version preference selection, or score filling.
Explicit model selections follow the same DSL rules as direct calls. To choose
input-dependent defaults, call resolve_default_methods and
resolve_default_versions explicitly and include their results in the effective
policy before saving. Unstated historical versions remain unknown. Distinct
source-local selections need separate effective policies/invocations; do not
apply one source's choice globally.
CSV/TSV exports retain the full definition, digest, derivation and execution record in long and wide form. Numeric measurements and annotations round-trip exactly as binary floats. Preserve the full evidence as well: a filtered ranking cannot restore excluded observations or discarded alternatives. Replay also needs the recorded Topiary version when exact execution semantics matter.
Proteasome cleavage and whole-peptide half-life can already be named explicitly,
for example peptide_view(proteasome_cleavage.score) and
peptide_view(serum_half_life.value). Saving a policy does not enable these
predictors or add processing weights to existing defaults. Typed extracellular
site-level evidence remains #288.
Composed consumer configuration¶
Vaxrank owns YAML composition and window/construct configuration. Its repeated
--config builtin:openvax-v1 --config overrides.yaml workflow merges mappings
left to right and replaces lists/scalars. The Topiary integration boundary is
the already-composed policy subtree:
from topiary import resolve_selection_policy
# `merged` comes from the consumer's config loader, after all overrides.
policy = resolve_selection_policy(merged["selection_policy"])
saved_settings = {
"selection_policy": policy.to_dict(),
"selection_policy_sha256": policy.sha256,
"selection_policy_provenance": config_provenance,
"vaccine_settings": effective_vaccine_settings,
}
# Reload complete definitions strictly; never apply newer authoring defaults.
policy = SelectionPolicy.from_dict(saved_settings["selection_policy"])
resolve_selection_policy accepts a JSON- or YAML-decoded mapping with required
name and score_by. It fills omitted optional settings only after composition.
An omitted filter in an override inherits the base filter; an explicit YAML
filter_by: null removes it. Replacing a filter expression does not AND it with
the old expression. Model selections are mappings; version selections are lists
of {kind, method, version} records, so an override replaces that entire list.
Topiary does not merge files or parse unrelated consumer settings.
Vaxrank 3.36.0 consumes this subtree in its shipped configuration bundle.
Its post-score minimum gate, missing-score fill, window rules, RNA weighting
and construct settings belong to that bundle too. The consumer fixture in
scripts/check_vaxrank_candidates.py loads builtin:openvax-v1 through
Vaxrank's public config loader and compares its numeric scores with
evaluate_selection_policy. A cutoff inside a score expression must not become
a destructive pre-filter when replaying that baseline.
The representative ranking above is one view. evaluate_selection_policy
retains every input row and evaluates explicit occurrence identities before
representative selection or vaccine-window construction:
from topiary import (
SelectionPolicy, evaluate_selection_policy, replay_selection_policy,
select_policy_representatives, read_tsv,
)
policy = SelectionPolicy(
"example-baseline", "(affinity.value < 5000) * affinity.value.logistic_normalized(350, 150)",
score_fill=0.0, min_score=1e-5, duplicates="best",
)
evaluated = evaluate_selection_policy(combined, policy)
occurrences = evaluated.occurrences # includes excluded and unscorable alternatives
eligible = evaluated.selected # no candidate collapse
representatives = select_policy_representatives(evaluated)
evaluated.evidence.to_tsv("complete-evaluation.tsv")
replayed = replay_selection_policy(read_tsv("complete-evaluation.tsv"))
This example illustrates a scoring convention; use Vaxrank's published bundle
for openvax-v1. Neither expression is a calibrated probability of immunogenicity. The cutoff
remains inside the score, with a separate inclusive minimum-score gate.
raw_score preserves missingness even when score_fill=0 supplies an effective
zero. Pre-filtered groups are not scored and cannot be restored by filling.
With no minimum gate, missing scores remain eligible but unranked, matching the
existing candidate-ranker convention.
For direct Vaxrank frames, pass
group_keys=["prediction_id", "peptide", "peptide_offset", "allele"] and the
consumer's per-occurrence alleles mapping/callback. The callback is evaluated
once per peptide identity and saved as concrete declarations, never Python code.
Projected allele groups carry supporting_rows links, not duplicated RNA counts.
evidence_rows links the group's original observations. These positions refer
to evaluated.evidence.long_df; retain that complete long-form result when
saving an evaluation. Input columns and metadata remain available there.
Use source_contexts={label: {...}} when sources have independent model defaults.
Every source label must be named; each mapping may replace default_methods,
default_versions, kind_support, and alleles. Filtering and scoring share
those choices. Definitions remain separate from runtime contexts, which are
stored alongside decisions in extra["policy_evaluation"]. Replay uses the
complete evidence and stored contexts and never invokes prediction. Model
defaults resolve ambiguity in the existing DSL; they are not a requirement that
all input measurements come from the selected model.
select_policy_representatives chooses an actual eligible occurrence, without
summing support, and records all alternative occurrence IDs, including excluded
alternatives. It records the choice and its runtime keys in the evaluation
evidence metadata; replay repeats that explicit choice and audit links
criteria to the selected occurrence and rationale. Direct consumers supply
their candidate_keys and strata if
combined-source candidate columns are absent. Missing values sort last; stable
input order breaks exact ties. Topiary owns these generic decisions; Vaxrank
still owns window geometry, source admission and construct assembly.
Named criteria and audit decisions¶
Criteria reuse existing DSL expressions through an explicit namespace. This example is synthetic policy content, not a recommended processing weight:
from topiary import SelectionCriterion, RankingTerm
binding = SelectionCriterion("binding", "affinity.value < 500", "eligibility")
processing = SelectionCriterion(
"processing", "peptide_view(proteasome_cleavage.score)", "score",
)
policy = SelectionPolicy(
"example-processing", 'criterion("processing")',
filter_by='criterion("binding")', criteria=(binding, processing),
ranking_by=(RankingTerm("n_rna_alt", ascending=False),),
unknown="exclude", duplicates="best",
)
evaluated = evaluate_selection_policy(combined, policy)
audit = evaluated.audit
Eligibility criteria compose with explicit &, |, and ~; score terms
combine through explicit arithmetic; ranking_by lists ordered expressions
and directions after the primary score. YAML overrides do none of this
implicitly. Numeric references can be compared to form predicates, and predicates can
participate in explicit score arithmetic. A bare numeric reference is not an
eligibility predicate; a bare predicate is not a named numeric score term.
An input column called binding remains a column; only criterion("binding")
means the named criterion. Unknown/cyclic references, duplicate names, wrong
roles and non-boolean eligibility outputs raise. Optional applies_to is an
eligibility expression; false applicability is recorded as not_applicable,
and its reference remains unknown rather than becoming an observed failure.
Audit rows retain occurrence identity, raw row links, criterion name and role, value, status and reason:
| Status | Meaning |
|---|---|
pass / fail |
An applicable predicate evaluated true / false |
unknown |
Missing column/input, missing/ambiguous/conflicting model evidence, unknown applicability, or an out-of-domain calculation |
value |
An observed numeric term, including zero |
not_applicable |
The applicability predicate was observed false |
not_evaluated |
Unreferenced, or scoring skipped after a pre-filter |
Named predicates use three-valued boolean logic. unknown="exclude",
"include", or "error" controls the decision on an unknown final eligibility
result, while audit values remain unknown. Direct-expression policies without
criteria keep historical DSL comparison behavior. A vector expression with an
ambiguous/conflicting model is unknown for that evaluation context; separate
source contexts prevent one source's model selection from being imposed on
another. Unexpected programming/DSL errors still raise.
Schema 2 persists criteria, their complete expansions, references, ordered
terms, fill/gate settings and unknown handling in the policy digest. Original
schema-1 definitions retain their original serialization and digest. Definitions
are self-contained; a mutable registry cannot change a saved recipe. Export
evaluated.evidence, not a filtered ranking, to reproduce rejected alternatives.
The Vaxrank consumer test carries these records through native dataset save/load
and verifies frozen scores, selected windows, peptide constructs and mRNA
constructs, plus a changed criterion that changes the selected window.
Coverage and policy comparisons¶
Coverage does not depend on whether a policy fills a missing score with zero, includes unknown eligibility, or rejects a candidate. Use retained evaluations:
from topiary import policy_coverage, summarize_policy_coverage, compare_policy_evaluations
coverage = policy_coverage(evaluated)
summary = summarize_policy_coverage(coverage)
reasons = summarize_policy_coverage(coverage, by=["level", "criterion", "reason"])
comparison = compare_policy_evaluations(evaluated, another_evaluation)
shared = comparison.df.query("assessment_set == 'both'")
coverage.to_tsv("coverage.tsv")
comparison.to_tsv("policy-comparison.tsv")
policy_coverage emits one score and each named criterion per occurrence. Native
pass, fail and numeric value are all assessed; unknown becomes missing.
Not-applicable and not-evaluated remain distinct. A measured zero is assessed;
a zero from score_fill retains its missing raw score. The original status,
reason, criterion value, raw/effective score and eligibility remain visible.
By default the denominator contains evaluated groups. Supply universe, a
DataFrame of unique identity keys that includes every evaluated group, to keep
expected inputs that produced no evaluation. Absent groups are not evaluated.
Keys default to the retained occurrence grouping; explicit keys must uniquely
identify occurrences. Do not drop flanks, offsets, samples or sources when that
would merge different occurrences.
Returned prediction rows alone cannot reveal failed or unattempted requests.
For model coverage, pass an explicit prediction_requests DataFrame: one row
per expected occurrence/model/field assessment, with the identity keys plus
canonical kind, prediction_method_name, predictor_version, and numeric
field (for example score). A null version matches only an unstated version.
Include allele_set for joint-genotype evidence. An optional diagnostic status
(missing, failed, not_applicable, or not_evaluated), reason, and detail
can preserve explanations such as insufficient_c_terminal_context or a backend
failure when no row exists. Diagnostics cannot override observed measurements.
Topiary does not infer a context/domain failure from an absent value.
Model records describe availability, not which model contributed to a criterion. They match exact model/version, occurrence context and canonical MHC scope; allele-free measurements can project, while allele-credited and joint measurements retain their scope. A criterion can combine multiple models or annotations, and this report does not invent lineage for that expression.
The summary separates requested assessment counts from unique occurrences,
unique peptide/allele/genotype candidates, and the union of referenced raw
prediction rows. Repeated source discoveries can describe one candidate;
allele projections can reference the same raw row. Finite raw fields are counted
separately as n_observed_prediction_rows, even if conflicting values make the
assessment unknown. Requests with no output contribute to assessment counts,
not raw-row counts. assessed_fraction explicitly uses all assessments as
its denominator, including inapplicable and unevaluated requests. Group by
reason or any occurrence key to inspect the cause or location of missingness.
Comparison preserves both memberships, eligibility, raw/effective scores and
criterion decisions. assessment_set separates both, left_only,
right_only, and neither; a finite raw score plus no unknown referenced
criterion is required for assessability. Unused and inapplicable criteria do
not disqualify an otherwise assessed score. Prefiltered groups have no assessed
score. raw_score_delta is right minus left only on the shared assessable
set; it does not normalize different policy scales or add scientific weights.
Reports retain definitions, digests, contexts, provenance and explicit reporting
inputs in extra["policy_coverage"] or extra["policy_comparison"]. Save the
underlying evaluated.evidence too: replay it with replay_selection_policy,
then regenerate coverage using the saved keys, universe and
prediction_requests (convert the latter two record lists to DataFrames when
not null). CSV/TSV readers may add their normal per-row file source label;
coverage counts and replay are unchanged. Reports run no predictors and change
no selection. Vaxrank owns vaccine-window/construct choices and its frozen
openvax-v1 bundle.
Identities and source evidence¶
| Column | Meaning |
|---|---|
source_label, source_row |
Source label and zero-based row in its normalized long table |
source_observation_id |
One source's sequence/context/annotation observation; differing abundance stays distinct |
candidate_id |
Same sample, peptide and canonical HLA allele across sources |
candidate_sample, candidate_allele |
Resolved identity, without replacing original sample/allele cells |
protein_sequence_id |
Exact full amino-acid sequence explicitly supplied in protein_sequence |
source_prediction_kind, source_prediction_method, source_predictor_version, source_prediction_run_name |
Original prediction identity, including for sparse wide-file round trips |
source_prediction_mhc_dependence |
Original allele-free, per-allele or joint-genotype scope |
Rows lacking sample identity require sample_name=. Already named samples
retain their identities. Never assign different patients the same label.
Original source metadata is retained under extra['combined_sources'].
Observations are retained without summing counts, averaging predictions or
choosing a source silently. Identical predictor/version names in different
pipelines can therefore have distinct attributable measurements. The DSL groups
combined frames by source observation, peptide and allele, with its usual
sample/genotype context. Existing frames keep their existing grouping. Source
annotations and new features work through Column(...) and string expressions;
peptide_view(...) still reads allele-independent evidence in pMHC scoring.
Incompatible allele scopes require filtering to compatible sources before a
kind expression can score them together. Contradictory predictions within the
same source/observation/model/version slot require distinct source labels or
prediction_run_name values, so the DSL cannot silently select the first row.
rank_candidates chooses one representative observation per candidate per
stratum. Default duplicates="error" requires scores to agree, including
whether they are missing. "best" and "worst" explicitly select an observation
without combining its abundance with another source's. candidate_observations
links all contributing observations in that selected view. MHC classes rank
separately by default. Missing scores have ranking_status="missing_score"
and no rank. Filters narrow the returned view, preserving original evidence.
A shared table does not calibrate incompatible scores. Select compatible measurements or an explicit composite expression; use stratified lists where a pooled policy cannot score every candidate. Tumor specificity requires its own evidence and eligibility decision.
RNA measurements need their own units and biological subject. For example, LENS reports transcript TPM, fusion fragments per million, splice-expression and viral read-count measures in different workflows. A common column name does not make them interchangeable. pVACseq reports also distinguish transcript expression from tumor RNA depth/VAF and report predictions per algorithm. Preserve these distinctions when defining a policy.
Isovar hypotheses for comparison¶
read_isovar_hypotheses reads isovar.protein_hypotheses.v1 and v2
JSON exports or the mapping returned by export_protein_hypotheses. Version 2
is produced by Isovar 1.37+. The companion
Isovar TSV lacks the evidence sets and full provenance; supply JSON here.
from topiary import read_isovar_hypotheses, combine_sources, rank_candidates
hypotheses = read_isovar_hypotheses("tumor-1.hypotheses.json")
combined = combine_sources({"reported": reported_candidates, "isovar": hypotheses})
alternatives = combined.filter_by("isovar_rank > 1")
ranked = rank_candidates(combined, "affinity.value", ascending=True)
Each imported row is one translation, including synonymous nucleotide sequences
and lower-ranked alternatives. A hypothesis with no translations still has a
protein-only observation. The reader creates no peptides, alleles or scores;
these rows have no candidate_id after combination. Adding them leaves the
reported candidates, their scores/ranks, and exact-peptide re-scoring calls
unchanged. Existing top-protein fragment selection, reconstruction settings and
filter defaults also remain unchanged. Using alternative ORFs to generate new
candidates requires a separate, explicit choice.
isovar_rank, representative and passes_all_filters describe the producer's
results; they do not establish default eligibility. Filtered and uncertain
outcomes stay visible for comparison. protein_hypotheses_complete and
protein_sequence_limit report whether the upstream export may be truncated.
RNA support and reconstructed sequence alone do not establish tumor specificity.
protein_hypothesis_sequence includes partial translated windows.
protein_sequence is populated only for translations explicitly starting at the
annotated start codon whose protein ends at a stop codon. Only these full
sequences enter protein_evidence_view; a partial window is never promoted to a
full ORF. The producer's exact sequence identifier is kept as
isovar_protein_sequence_id, while the combined table reserves
protein_sequence_id for its full-protein grouping.
The complete original export is retained in
hypotheses.extra['isovar_hypotheses'], including events without proteins,
reference contexts, observed edits, filters and RNA evidence sets. After
combination it is under
combined.extra['combined_sources']['isovar']['extra']['isovar_hypotheses'].
Topiary CSV/TSV saving retains that metadata in both long and wide form.
Literal sample IDs and sequences also survive reload: for example, 001 stays
a string and the amino-acid window NA stays a sequence, not a missing value.
Read and fragment counts have separate protein_* and translation_* columns.
Both export versions populate protein_reads and translation_reads; the
older *_segments columns remain aliases for compatibility. Optional *_umis
and *_cells are label counts with separate *_umis_complete and
*_cells_complete flags. Null means unassessed, false means the count is not
exact, and true means the producer assessed complete labels. *_unlabeled_reads
and *_unknown_library_reads explain incomplete measurements; detailed
label_statuses and allele-level support stay in the original metadata.
These are not TPM or independent molecule counts.
For example, the existing filter DSL can restrict comparisons to complete
UMI measurements with combined.filter_by("protein_umis_complete & protein_umis >= 2").
This filters comparison observations; it does not admit alternatives as candidates.
To combine read support explicitly with Isovar 1.37+:
from isovar import union_rna_support
from topiary import normalize_isovar_rna_support
export = hypotheses.extra["isovar_hypotheses"]
keys = hypotheses.df.loc[hypotheses.df.isovar_rank <= 2, "protein_evidence_set_id"]
support = union_rna_support([
normalize_isovar_rna_support(export["evidence_sets"][key]) for key in keys])
This counts shared reads once, including support repeated across synonymous
translation rows. Normalization translates legacy segments/segment_ids
into reads/read_ids and validates counts and membership without changing the
saved export. The union refuses incompatible evidence scopes. It does not union
UMI/cell counts: exports do not include the label identities needed for that.
Missing evidence-set
IDs cannot be resolved or safely combined from counts alone. The comparison
reader itself does not import Isovar or run any reconstruction or predictor.
Full ORFs and RNA-only tables¶
import pandas as pd
from topiary import protein_evidence_view
orf_table = pd.DataFrame({
"event_id": ["GRCh38:event-123", "GRCh38:event-123"],
"protein_sequence": ["MAAASIINFEKL", "MQQQSIINFEKL"],
"transcript_expression": [11.0, 22.0],
"expression_unit": ["TPM", "TPM"],
})
combined_orfs = combine_sources({"exacto_normalized": orf_table}, sample_name="patient-01")
proteins = protein_evidence_view(combined_orfs)
No peptide, HLA or prediction columns are required here. The protein view links
equal full sequences within a shared sample and explicit event_id, retaining
different sequences as alternative ORFs. Its observation links lead back to each
tool's RNA measurements; abundance is never automatically averaged or summed.
Equal protein products do not establish equal nucleotide ORFs. Distinct
orf_id, coding_sequence, transcript, frame or completeness annotations
remain independent source observations even when the translated sequence is
identical. Use reconcile_evidence below to establish explicit ORF relationships
while retaining those observations.
event_id must be explicitly harmonized, including assembly/variant identity
where appropriate. Native labels from different tools are not assumed equal.
Without a common event ID, observations stay separate while sequence IDs still
show exact sequence agreement. Local pep_context or sequence is not promoted
to a full ORF. A peptide-only table does not establish its generating full ORF.
To use another source's RNA measurement for an explicitly matching ORF, first
reconcile the input and join the selected measurement as a named feature. join_annotations refuses ambiguous annotation keys:
from topiary import join_annotations, reconcile_evidence
combined = reconcile_evidence(combined)
# Require an explicit ORF match; equal proteins alone are insufficient.
keys = ["candidate_sample", "orf_hypothesis_id"]
rna = combined.df.loc[
combined.df.source_label.eq("exacto_normalized"),
[*keys, "transcript_expression"],
]
enriched = join_annotations(
combined, rna, on=keys, prefix="exacto",
provenance={"source": "exacto_normalized", "unit": "TPM",
"policy": "same sample and reconciled ORF; retain transcript-level unit"},
)
ranked = rank_candidates(
enriched, "exacto_transcript_expression / affinity['netmhcpan'].value",
duplicates="best",
)
This expression illustrates a policy, not a calibrated biological model. The alternative ORF's abundance cannot attach just because its variant or gene agrees. Missing or ambiguous identity requires explicit resolution.
Reconcile ORFs and RNA observations (5.87.0+)¶
reconcile_evidence(combined) returns a copy with stable links between biological
entities. It neither chooses a caller nor changes the candidate universe.
evidence_views(combined) exposes events, orfs, proteins, occurrences,
candidates, rna_observations, links and candidate_occurrences as DataFrames; supply
source_labels=["lens"] for a source-stratified view.
| Input assertion | Identity rule |
|---|---|
event_id or list-valued event_ids |
Same sample, explicit reference_name and normalized event name; absent reference remains source-local |
orf_id |
Caller-local ID; contradictory descriptors raise, absent descriptors can be supplied by another row with the same local ID |
coding_sequence and transcript_path, or transcript_id plus orf_start/orf_end |
Cross-caller ORF agreement requires the same sample, reference and all supplied descriptors |
reading_frame, orf_completeness, start/stop flags, linked_variants |
Retained in ORF identity; differing assertions remain alternative hypotheses |
protein_sequence |
Exact full product, separate from nucleotide ORF identity |
protein_hypothesis_sequence |
May be partial; never promoted to a full protein |
peptide_start/peptide_end |
Zero-based half-open occurrence in the supplied protein; validated against its sequence |
| Missing peptide coordinates | Source-local occurrence; no guessed equivalence from a shared peptide |
ORF bounds are also zero-based half-open. transcript_path is an explicit ordered
JSON path in the caller's normalized reference convention. Topiary does not lift
over assemblies, normalize native genomic variant strings or infer a path from
gene names. Incomplete records remain useful without establishing equivalence.
Reference scope follows the distinction between sequence and annotation identity
in NCBI's feature documentation.
rank_candidates retains the chosen row's orf_hypothesis_id and
peptide_occurrence_id. candidate_observations and the relational links
retain alternative support. Rediscovery does not multiply candidate scores.
Use an explicit duplicate policy or source filter when measurements disagree.
candidate_occurrences links all same-sample peptide occurrences to existing
pMHC queries, including sequence reports with no HLA assignment. Its
reported_candidate flag distinguishes a reported peptide-HLA pair from a
sequence match. These links transfer no scores, expression, specificity or
eligibility between observations, and create no new candidates.
RNA observations can be supplied as a list of records in rna_observations:
measurement = {
"sample_name": "patient-01", "entity_type": "transcript",
"entity_id": "ENST-example", "quantity": "count", "unit": "reads",
"value": 2, "library_id": "rna-library-1", "read_set_id": "alignment-1",
"evidence_unit_ids": ["read-a", "read-b"], "method": "caller", "version": "1",
}
normalize_rna_observation validates this shape. entity_type may be gene,
transcript, orf or variant. A TPM measurement uses quantity="abundance",
unit="TPM" and no evidence-unit list. Unknown values remain null. Original
quantifier metadata and extra fields survive; gene/transcript expression never
becomes ORF expression automatically. Each caller's observations remain distinct.
union_rna_observations([measurement, other_measurement]) unions identified
read, fragment, UMI or cell memberships only within one sample, library, read-set
namespace, measured entity and unit. Shared members count once. Missing membership,
different namespaces and TPM quantities raise: neither different file names nor
different callers establish independent evidence. An explicitly empty membership
is measured zero; an absent membership is unknown. Scalar expression/count
columns already present in source tables stay untouched and are never implicitly
converted into identified evidence sets.
All identity columns, descriptors, alternative observations and metadata survive Topiary CSV/TSV save/reload in long and wide forms. Reconciliation can be repeated after reloading. The composed workflow test covers original and additive scoring, retained hypothesis selection, and count union without running a predictor during assembly.
Predict exact peptide occurrences¶
Topiary 5.91.0 adds predict_peptide_occurrences and
TopiaryPredictor.predict_from_peptide_occurrences for a selected peptide
universe. Supply one record per (prediction_id, peptide, peptide_offset) within
each sample (when sample_name or candidate_sample is supplied). An ID may name
several peptide windows in the same source. Use the same occurrence
identity Vaxrank and the policy evaluator consume. Different occurrences of one
peptide remain separate even when they share a gene, coordinate or HLA candidate.
import pandas as pd
from topiary import (
TopiaryPredictor, TopiaryResult, SelectionPolicy, evaluate_selection_policy,
)
# model is an explicitly configured mhctools predictor with patient alleles.
occurrences = pd.DataFrame([
dict(prediction_id="gene-a:12", peptide="SIINFEKL", peptide_offset=12,
n_flank="AAA", c_flank="GGG", gene="gene-a"),
dict(prediction_id="gene-b:20", peptide="SIINFEKL", peptide_offset=20,
n_flank="TTT", c_flank="CCC", gene="gene-b"),
])
predictor = TopiaryPredictor(models=model)
predictions = predictor.predict_from_peptide_occurrences(occurrences)
evaluation = evaluate_selection_policy(
TopiaryResult(predictions), SelectionPolicy("binding", "1 / affinity.value"),
group_keys=["prediction_id", "peptide", "peptide_offset", "allele"],
kind_support=predictor.kind_support,
)
evaluation.evidence.to_tsv("occurrence-evidence.tsv")
The model scores only the supplied peptides against its configured alleles.
Within a shared sample/flank/genotype context, distinct peptides are batched and
identical inference inputs are scored once, then linked to each original
occurrence. Source annotations and read support are copied, never summed.
prediction_mhc_dependence records each kind's scope; haplotype predictions carry
the configured allele_set. An explicitly supplied haplotype set must match the
configured model. Missing kind/allele coverage and ambiguous outputs raise.
Flank-dependent models require both n_flank and c_flank. An empty string is a
known terminus; a missing value is unknown. use_flanks=False explicitly requests
peptide-only inference while preserving the original context. The output's
prediction_flanks_supplied records whether flanks were supplied in the model call; it does not imply that every returned prediction kind depends on them.
Missing source offsets remain missing. Model peptide-length/domain checks still
apply; no padding, comparator sequence, or additional peptide window is inferred.
With TopiaryPredictor(predict_wt=True) (or predict_wt=True on the standalone
function), supplied wt_peptide values are scored using their own
wt_n_flank / wt_c_flank. Flank-dependent comparator prediction requires those
flanks or explicit use_flanks=False; primary-peptide context is not borrowed.
Absent comparators keep missing scores. Explicit filter/sort settings operate on
occurrences; only_novel_epitopes does not reinterpret this already selected
peptide universe. Cache failures are strict by default; an explicit
cache_miss_handler / on_miss retains the existing reported-partial-batch
contract, with occurrence IDs in the failure records.
Prediction inputs contain context and annotations, not historical measurement columns. Keep imported measurements in their original table and use the additive operation below. Context-aware live prediction requires mhctools' flank-aware API; sparse historical tables with unknown model versions remain the separate prediction-source work in #368.
Add prediction features on demand¶
from mhctools import MHCflurry
from topiary import rescore_candidates
model = MHCflurry(
alleles=["HLA-A*02:01"], presentation_allele_mode="per_allele",
)
enriched = rescore_candidates(
combined, model, prefix="fresh",
select="candidate_allele == 'HLA-A*02:01'",
use_flanks=False, # explicitly request peptide-only predictions
)
original_policy = rank_candidates(
enriched, "affinity['netmhcpan'].value", ascending=True, duplicates="best",
)
new_policy = rank_candidates(
enriched, "fresh__mhcflurry__pMHC_presentation__score", duplicates="best",
)
Features such as fresh__mhcflurry__pMHC_affinity__value are appended. Original
values and predictor names keep their meaning. Adding features does not switch
the ranking policy. They work in filter_by, sort_by, evaluate_scores, and
Vaxrank DSL expressions. extra['candidate_rescoring'] records the new producer,
models, versions, measurement semantics, UTC creation time, selection and discovery-source labels.
Discovery source and prediction producer are separate facts.
Configured mhctools models supply predict_dataframe and kind_support().
This path reuses predict_peptide_occurrences to batch compatible requests and
preserve source identity; it appends primary-peptide features and leaves all
historical comparator measurements intact.
Supplied flanks are forwarded by default; flank-dependent models require both
flanks or explicit use_flanks=False. Haplotype models require a matching
stated allele_set. Existing candidate identities are preserved; ORF-only rows
remain unscored. Scanning new windows or adding HLA candidates is a separate
operation, with novelty/context handling tracked in
#364.
Every declared allele-dependent prediction kind must cover the selected candidate. An allele-free processing output cannot substitute for missing affinity or presentation; processing-only models remain supported.
Save and reload flanking context¶
An empty n_flank or c_flank string states a known protein terminus; a missing
value states that the context is unknown. Use Topiary's to_tsv/to_csv and
read_tsv/read_csv together to preserve that distinction in both long and wide
tables. Context-dependent re-scoring after reload then receives the same flanks
as before saving; unknown context still requires an explicit resolution.
Files with canonical flank columns include
#topiary_flank_encoding=escaped-v1. In those columns, missing values are
written as <NA> and known-empty strings remain empty cells. Ordinary sequence
text, including NA, remains literal text. A literal <NA> or a string starting
with a backslash gains one leading backslash, which the reader removes.
Other columns containing only strings and missing values use the same cell
encoding. The #topiary_text_encoding metadata records escaped-v1 and the
affected column names, preserving literal identifiers and sequence text.
Columns whose stated cells are all dicts or lists, such as
measurement_context, are written as one JSON document per cell, with missing
cells as <NA>. The #topiary_json_encoding metadata records json-v1 and the
affected column names; the reader decodes the cells back into dicts and lists.
Values follow JSON's data model, so tuples come back as lists and non-string
keys as strings. A cell JSON cannot represent raises TypeError before the
output file is opened.
Numeric measurements retain their normal numeric representation and inference.
Legacy files without this flag retain the earlier missing-value interpretation.
Their blank flank cells cannot distinguish termini from unknown context. Recover
that distinction from the original source; do not replace all missing flanks
with empty strings. Ordinary pandas.read_csv also cannot recover the distinction
without decoding this format.
Vaxrank¶
Pass the full enriched frame to Vaxrank's
attach_per_allele_scores(..., topiary_df=...). Rebuilding it from prediction
objects alone loses annotation features. Its current grouping key is
prediction_id: map source_observation_id to that name in a separate consumer
view, preserving original IDs in the combined evidence. Explicitly select
representative candidates before construction so repeated discovery does not
multiply vaccine targets. Construct generation still requires suitable sequence
context, targetability and tumor-specificity admission evidence.
scripts/check_vaxrank_candidates.py tests actual DSL scoring and
VaccinePeptide construction against released Vaxrank for mutation, fusion,
splice, CTA, ERV and viral synthetic examples. Original, re-scored and
RNA-enriched policies select different construct orderings. Generalized
Vaxrank file/CLI ingestion and version adoption are tracked in
Vaxrank #497.