Reading native Exacto output¶
read_exacto imports native Exacto TSV files into TopiaryResult without
running Exacto, reconstructing reads or executing a prediction model. Supply
sample_name: these files do not identify the patient/sample themselves.
from topiary import read_exacto
exacto = read_exacto(
"sample_exacto_peptide_variants.tsv",
sample_name="patient-01",
primary_structures="sample_exacto_primary_structures.tsv",
tag="exacto-run-01",
)
Supported native schemas¶
The tested producer is Exacto 0.4.6, commit
307c086.
Topiary identifies required column sets rather than assuming that a filename
or the producer's package version establishes a schema. EXACTO_SCHEMAS
exports these column sets; schema= can select one explicitly. Input may be a
path, gzip-compressed path or open text stream. Extra columns and original
records remain in result.extra["exacto"], including companion tables.
| Schema | Native output | Normalized observations |
|---|---|---|
peptide_variants |
call-peptide-vars TSV |
Reported peptide sequences and parent IDs; optional primary structures supply exact occurrences and context |
primary_structures |
translate-structs TSV |
One translated ORF hypothesis per native peptide ID, with producer-annotated changed intervals |
translations |
translate-seqs TSV/TSV.gz |
Every reported RNA/ORF translation, including partial alternatives |
transcript_read_support is a companion table, supplied with the argument of
that name. It is accepted with primary structures, or peptide variants plus
primary structures. Provide library_id and read_set_id; stream inputs also
require a tag identifying the Exacto run. Local transcript model IDs are
scoped by this tag, or by the input path. Use the same run tag when importing
different files from one run. Distinct read names supply a transcript-level
read count with explicit membership. Missing models remain unknown. These
counts are not variant support or ORF-specific abundance; repeated peptide
rows do not create independent RNA observations.
The native corpus in tests/data/exacto pins source paths and SHA-256 digests:
228 peptide records, six primary structures and 77 translations. Unsupported
headers, primary record types, sequence-bearing event rows, invalid codons or
inconsistent peptide/parent geometry raise ValueError. This reader does not
import arbitrary Exacto tables or reconstruct genomic events from local IDs.
What the evidence means¶
Native DNA and RNA call IDs, transcript model IDs, reference transcript IDs and all source records are retained. Local numeric IDs do not establish agreement with another caller's genomic events. Translation ORF bounds are converted from zero-based inclusive coordinates to zero-based half-open coordinates; original values remain in metadata. Translation IDs include both RNA and peptide IDs, since the producer reuses peptide IDs across RNA records.
With primary structures, peptide coordinates and flanks are checked against
the reconstructed translation. Primary indices count base and event records,
so they are not treated as amino-acid coordinates. Changed amino-acid intervals
come from Exacto's explicit amino_acid_change labels. Without primary
structures, context and novelty intervals remain unknown and the reported
peptide is still available.
sequence_source="predicted_from_observed_rna" distinguishes inferred protein
sequence from a protein-expression measurement. A complete protein_sequence
requires a start codon and terminal stop. Partial translations remain in
protein_hypothesis_sequence. Native files do not establish matched-normal
specificity or comparator protein sequences: these stay unknown. A changed
sequence is not automatically tumor-specific.
Combine with reported candidates¶
from topiary import (
combine_sources, evidence_views, melt_pvacseq_algorithms, rank_candidates,
read_lens, read_pvacseq, reconcile_evidence,
)
combined = reconcile_evidence(combine_sources({
"exacto": exacto,
"lens": read_lens("lens.tsv"),
"pvacseq": melt_pvacseq_algorithms(read_pvacseq("all_epitopes.tsv")),
}, sample_name="patient-01"))
ranked = rank_candidates(
combined, "affinity['netmhcpan'].value", ascending=True, duplicates="best",
)
by_source = rank_candidates(
combined, "affinity['netmhcpan'].value", ascending=True, duplicates="best",
strata=["source_label", "candidate_mhc_class"],
)
combined.to_tsv("combined-evidence.tsv")
links = evidence_views(combined)["candidate_occurrences"]
Exacto's native files contain no HLA assignments or pMHC predictions. Therefore
its rows retain null allele, value and candidate_id; ORF-only rows also
have no peptide. The ranking universe comes from already reported pMHCs.
candidate_occurrences links same-sample, identical peptide occurrences to
those queries, with reported_candidate=False for a sequence match. This does
not transfer expression, scores or eligibility to a different observation.
All source records and native metadata survive long/wide CSV and TSV.
Choose compatible original measurements explicitly. Combined-source
ranking explains conflicts, missing scores and additive
rescore_candidates calls; rescoring alone does not invent HLA combinations.
Optional new windows¶
read_exacto_fragments(path, **kwargs) accepts the same inputs and produces
ProteinFragment records with the recovered sequence, novelty intervals and
native provenance. It calls no model. Empty native tables return an empty list.
Pass these fragments to TopiaryPredictor.predict_from_fragments to explicitly
scan new peptide windows. This changes the candidate universe and is separate
from rescoring reported candidates. Select lengths/alleles deliberately and
use overlaps_target to filter recovered changed intervals. Missing intervals
remain unknown. Fragment conversion preserves each native observation, so
several reported peptides from one parent may carry the same parent sequence.
The general reader uses no Varcode fusion reconstruction: that adapter handles selected DNA-SV-linked linear RNA models, whereas these native translated sequences can be imported directly without genomic anchors or raw reads.