Osteosarc RNA: indels and rearrangements¶
This is a source-pinned reconstruction audit, not a claim of peptide binding, tumor specificity, clinical suitability, or coverage of every website variant. Historical pVACseq reports and newer RNA evidence remain separate.
Historical pVACseq annotation gap¶
The 21 archived final reports represent 20 genomic alleles. They lack RNA depth, RNA VAF and gene expression already in the source. The retained April 27 variant input has 63 rows, all missing those annotations before prediction. RNA sequencing exists; those quantities were not populated into this historical prediction input. Recovering the original annotated VCF and checking sample FORMAT tags and gene/transcript matching is tracked in Topiary #339. See pVACseq coverage for the offline selected-column tests. New January-2025 RNA counts and expression can now be used alongside the historical predictions through the separate RNA overlay. It preserves the original missing values and does not depend on resolving the old annotation-step history.
Exact small indels¶
All coordinates below are GRCh38 VCF anchors; counts are RG/QNAME templates, not proven independent molecules. RNA products and samples are not pooled.
| Candidate | Exact allele | Evidence and status |
|---|---|---|
| GTF3C5 | chr9:133057893 GGAGGAGGAGGAA>G | Present in historical pVAC reports and existing RNA regression fixtures. Additional default-policy T1/T2 ONT support is 117/128; short-read support 7/12. |
| RNF213 | chr17:80327830 ATAC>A | Present in historical pVAC reports. Additional T1/T2 ONT support 5/14; short-read support 42/11. |
| GLIS3 | chr9:3856149 CTGATGTGG>C | Absent from historical pVAC reports. Nine alternate T1 short-read templates; independently checked 46-aa RNA context. |
| KTN1 | chr14:55627965 G>GTT | Absent from historical pVAC reports. T2 ONT 10 alternate templates, 29-aa context; T2 short reads 3 under both placement policies, 18-aa context (Isovar 1.18.1). |
GLIS3 and these KTN1 reconstructions pass default result filters. Isovar 1.18.1 fixes insertion-boundary assignments on the unchanged input records: KTN1 ONT alternate/other becomes 10/3 rather than 10/8, and default short RNA becomes 3/0 rather than 2/1. The earlier rejections were consequences of that counting defect. A reconstructed sequence is nevertheless not automatically an accepted result; the full-index audit retains filtered examples. A short context can supply mutation-overlapping 9-mers without being long enough for a 25-mer. No default threshold has been relaxed.
Actual rearrangement alignment paths¶
The extraction correction already shipped in Isovar #291 / 1.17.4. It uses observed partner SAM records and their full CIGARs. An SA tag declares an alignment relationship; it is not a substitute for a missing record, nor a reliable source of exact base-level mapping. The SAM specification defines primary/supplementary records, strand flags and clipping.
The follow-up matched every original record for 13 complete paths against the retained tagged-BAM regional acquisitions, verifying each acquisition hash. The three selected input files are pinned byte-for-byte with their original SAM/partner SAM records, source-query offsets and Ensembl 87 reference models.
| Rearrangement / sample | All complete paths | Distinct cell/UMI labels | Selected identical RNA window | Coding result |
|---|---|---|---|---|
| GABBR1–SLC29A1 T1 | 2 | 2 | 120 nt; direct junction | unresolved frame |
| GABBR1–SLC29A1 T2 | 3 | 3 | 120 nt; direct junction | unresolved frame |
| OTUD7A–FMN1 T2 | 8 | 6 | 122 nt; AG insertion; 4 selected observations | unresolved frame |
The selected OTUD7A window contains AG in its displayed minus/plus orientation. Do not substitute the DNA catalogue's CT spelling without accounting for orientation and the observed RNA sequence. The four observations supporting that exact selected window are not the eight total complete paths. Cell/UMI labels are not independently validated molecule counts.
Both inputs have unresolved frames, with reason
no_exact_collinear_annotated_donor, and no translations under the pinned
Ensembl 87 models. Isovar's v1 fusion schema calls this unresolved_frame;
the v2 schema (Isovar 1.31+) calls it unresolved and places translations in
paths[0]. This means the observed RNA junction is not yet connected to
a supported coding frame; it does not mean the RNA junction is absent.
Do not guess a reading frame or translate all frames and label the result
RNA-supported. A next coding-reconstruction step needs a validated transcript
path and CDS/frame relationship.
Deduplicated ONT predecessors can retain SA declarations without the actual partner records. Their lack of a complete observed path cannot establish absence of fusion RNA. Tagged and deduplicated products are processing-related and must not be counted as independent additional support.
Offline checks and provenance¶
The checked-in fixture manifest pins source-BAM receipts, 13 path identities and original-record hashes, and three selected JSON input hashes. CI checks the nine selected observations: paired-record identity, cell/UMI agreement, source sequence at the recorded offset, junction insertion, evidence counts and continued absence of a coding translation. This is not a rerun of all-source BAM extraction in CI. Upstream full-path tests recount the pinned original paths and validate their actual CIGAR-derived windows.
The local regeneration script verifies full source acquisitions before copying inputs. No sequence is invented and no historical pVAC row is backfilled:
python tests/data/osteosarc_rearrangements/regenerate.py /path/to/2026-09-17_05-32-18-330588Z
Larger leads remain separate¶
DLG5's event affects the canonical start region; its full DNA junction contains additional sequence, so the nominal 79.5-kb deletion is insufficient to specify the complete allele. The completed sequence-resolved follow-up supports the DNA junction but leaves the mutant RNA junction and coding frame unresolved; the expanded audit queried 14 RNA products without pooling them. AFF3's 128-bp and KEAP1's 211-bp deletions do not overlap coding-transcript exons in the Ensembl 87 models checked. Ordinary RNA splice skips across those intronic intervals do not distinguish deleted from wild-type alleles. None should receive a guessed coding consequence from these footprints.