SV coordinates and fusion rules¶
For loading and result access, start with Structural variants. This reference describes how the current implementation handles junctions and transcript boundaries. The biological examples use Ensembl 95.
Alleles, coordinates, and exports¶
Structural ref/alt and original_ref/original_alt expose the actual record
alleles, including symbolic or breakend ALT. Small-edit flags (is_snv,
is_indel, is_insertion, is_deletion, is_transition, is_transversion)
are false; use is_structural and sv_type for event classification.
SV-containing tables add sv_type, end, mate_contig, mate_start,
affected_start, and affected_end. These CSVs are summaries, not lossless
structural archives: from_csv rejects them. Retain original VCF/JSON variants
and their RNA evidence; point-only and empty table columns are unchanged.
Symbolic VCF span records retain POS as the padding/junction coordinate;
affected_start..affected_end is the inclusive changed interval POS+1..END.
This applies to DEL, DUP, INV and CNV (including CN0/CN3). This parser requires
a nonempty span with END > POS. Direct StructuralVariant(...) construction
keeps its existing explicit-coordinate defaults: pass affected_start
separately when start includes padding. Paired breakends already supply it.
A directly constructed BND takes mate_contig, mate_start and
mate_orientation from a breakend ALT when they aren't passed, and raises
ValueError when a passed one contradicts it.
Junction orientation¶
An SV creates one or more junctions, each joining two breakpoints. At
each breakpoint the junction keeps either the bases to the left (up to
and including the position) or those to the right.
StructuralVariant.junctions lists them.
For a breakend record, the ALT says which sides are kept:
| ALT | This record keeps | Mate keeps |
|---|---|---|
t[p[ |
left | right |
t]p] |
left | left |
]p]t |
right | left |
[p[t |
right | right |
For the other SV types the junctions follow from the type (start is the
base before the event, as in VCF):
| Type | Junctions |
|---|---|
DEL |
start (keeps left) joined to end + 1 (keeps right) |
DUP |
end (keeps left) joined to start + 1 (keeps right) |
INV (symbolic) |
two: start–end keeping left, and start + 1–end + 1 keeping right |
INV built by pair_breakends |
only the junction its breakend pair observed |
INS, CNV, single breakend |
none |
A breakend with no readable ALT has unknown sides, except that
mate_orientation "[[" / "]]" still gives the mate's side (right /
left).
Fusion partners¶
Keeping a side of a transcript keeps either its 5' end or its 3' end, depending on its strand:
| Strand | Keeps left | Keeps right |
|---|---|---|
forward (+) |
its 5' end → 5' partner | its 3' end → 3' partner |
reverse (−) |
its 3' end → 3' partner | its 5' end → 5' partner |
A junction makes a GeneFusion for a transcript when all of these hold:
- The transcript contains one end of the junction.
- Its gene doesn't also contain the other end (that would be intragenic).
- At the other end there's a protein-coding transcript in a different gene, whose gene doesn't contain this end either.
- The two roles are opposite: one 5' partner, one 3' partner. Two 5' ends (head to head) or two 3' ends (tail to tail) can't be read through, so they don't fuse.
Every partner transcript satisfying these rules is retained as a GeneFusion
candidate, across all compatible junctions. Junctions retaining the annotated
transcript's 5' end come first, followed by pyensembl's partner order. The first
candidate stays the primary result for compatibility; this is not a likelihood
ranking. There is no count cap, and distinct isoforms remain separate even if
they predict the same protein. Repeated copies of the same junction/partner
record are deduplicated.
Each candidate has its own transcript pair and reference-derived cDNA/protein,
when sequence is available. A missing sequence leaves that candidate unresolved.
This enumerates the current annotated-isoform model, not novel splicing, phased
multi-SV paths, or all possible mature RNAs. A supplied alt_assembly still takes
precedence over reference-derived sequence. See candidate access.
GeneFusion.transcript is the transcript being annotated, which can be
either partner. five_prime_transcript and three_prime_transcript say
which is which, and partner_transcript is the other one.
Partners starting just past the other end¶
When the annotated transcript is the 5' partner and no transcript at the other
end needs to span it, a protein-coding transcript can still be joined whole: one
that starts on the other end's kept side, at most
MAX_UPSTREAM_PARTNER_DISTANCE (10 kb) past it, on the strand the junction
continues along, and that has a second exon. Its first exon has no splice
acceptor, so the fused cDNA continues from its second exon. This follows LINX's
rule for 3' breakends upstream of a gene, with LINX's limit for fusions other
than known pairs.
These fusions are further candidates after the effect's primary classification,
never the primary one, because RNA need not splice that way. In the Sid
osteosarcoma, the GABBR1--SLC29A1 junction lands 597 bases before SLC29A1's
first transcript starts. Isovar's long reads there run from GABBR1 intronic
sequence into the intergenic stretch and end in poly(A) before SLC29A1. So a
breakend there stays TranslocationToIntergenic, with one fusion candidate per
SLC29A1 isoform. The model's evidence records three_prime_breakend="upstream"
and three_prime_upstream_distance. A DEL, DUP or INV gains the same candidates
behind its local consequence. Its most_likely_effect is unchanged, but, like
any structural effect set, its priority follows its most disruptive candidate.
Breakend outcomes¶
Effects on the transcripts at the record's own breakpoint:
| This breakpoint | Mate | Effect |
|---|---|---|
| Intergenic | anything | Intergenic |
| In a coding gene | intergenic | TranslocationToIntergenic, plus fusion candidates for partners starting just past the mate |
| In a coding gene | in a non-coding gene only (e.g. MALAT1) | TranslocationToIntergenic |
| In a coding gene | in a coding gene, opposite roles | GeneFusion |
| In a coding gene | in a coding gene, head to head or tail to tail | TranslocationToIntergenic |
| In a coding gene | in the same gene | the local DEL, DUP or INV its kept sides describe (see Same gene below) |
| In a coding gene | single breakend (no mate) | TranslocationToIntergenic |
| In a coding gene | ALT isn't a breakend | GeneFusion treating this transcript as the 5' partner, with a warning |
TranslocationToIntergenic therefore means "a breakend that doesn't form
a gene fusion", not only "the mate is intergenic".
Its reference-derived mutant_transcript contains only the retained local
5' prefix or 3' suffix, including the base at an exonic breakpoint. Intronic
breakpoints retain the corresponding annotated exons. A. keeps genomic left;
.A keeps genomic right, even without a mate. Transcript cDNA is already
strand-oriented, so the segment itself has strand="+" on either gene strand.
The fragment is labeled evidence["sequence_status"] = "retained_reference_fragment";
full cdna_sequence and mutant_protein_sequence remain None. Concatenating this
one fragment does not establish the full allele or prove a mature RNA/protein.
Unknown local orientation (including mate-orientation metadata alone) leaves
mutant_transcript None. A supplied alt_assembly still takes precedence.
Intergenic ↔ gene. The record at the intergenic end reports
Intergenic; the gene is annotated only from its own record, which
reports TranslocationToIntergenic. Fusions with a transcript starting
just past the intergenic end are its only further candidates; varcode
doesn't model an intergenic promoter or enhancer driving a gene.
Same gene. A junction with both ends in one gene is intragenic, so it isn't a fusion for that gene's transcripts. A transcript read across it carries the DEL, DUP or INV its kept sides describe, so varcode annotates each of those transcripts as that local event:
- keeping the left side of the lower position and the right side of the higher one is a DEL between them;
- keeping the right side of the lower and the left side of the higher is a tandem DUP;
- keeping the same side at both is an INV.
This is also what a labeled breakend pair gives once pair_breakends types it
by the caller's SVTYPE. A single record annotates exactly as its labeled pair
would, but its effects still report the record. Their evidence gives
sv_type="BND" and junction_event, the local event it was read as.
For example, openvax-v2's ATP5MG--KMT2A junction, chr11:118,401,717 to
118,468,775, is exactly intron 1 of Ensembl's ATP5MG-KMT2A readthrough
transcript AP001267.5-202. A breakend there reports that transcript as
Intronic, and the ATP5MG isoforms, whose gene doesn't reach KMT2A, as fusions.
Direction. Both records of a junction report the same fusion
direction, each on its own gene. For CFTR (chr7:117,485,000, +, keeping
left) joined to BRCA1 (chr17:43,120,000, −, keeping left):
| Record | Annotated on | Result |
|---|---|---|
N]17:43120000] at chr7:117,485,000 |
CFTR | GeneFusion, 5' = CFTR, 3' = BRCA1 |
N]7:117485000] at chr17:43,120,000 |
BRCA1 | GeneFusion, 5' = CFTR, 3' = BRCA1 |
N[17:43120000[ at chr7:117,485,000 |
CFTR | TranslocationToIntergenic (head to head) |
]17:43120000]N at chr7:117,485,000 |
CFTR | TranslocationToIntergenic (tail to tail) |
The two records' primary proteins can differ (712 aa vs 1,854 aa here),
because their first transcript pairs differ. Both records now retain other
compatible isoforms in .candidates; matching the same 5'/3' transcript pair
gives the same predicted protein from either end.
Deletions, duplications and inversions¶
Whether a span fuses the genes at its two ends depends on their strands:
| Type | Genes on the same strand | Genes on opposite strands |
|---|---|---|
DEL |
fusion: the gene upstream in transcription is 5' | no fusion |
DUP |
fusion: the gene downstream in transcription is 5' | no fusion |
INV (symbolic) |
no fusion | reciprocal fusion candidates; each gene's own 5' contribution comes first |
Examples from chr7, with CFTR (+), LSM8 (+, downstream) and CTTNBP2
(−, downstream):
| Event | On CFTR | On the other gene |
|---|---|---|
DEL CFTR..LSM8 |
GeneFusion CFTR → LSM8 |
GeneFusion CFTR → LSM8 |
DUP CFTR..LSM8 |
GeneFusion LSM8 → CFTR |
GeneFusion LSM8 → CFTR |
INV CFTR..LSM8 |
Unresolved candidate |
Unresolved candidate |
DEL CFTR..CTTNBP2 |
deletion consequence | deletion consequence |
DUP CFTR..CTTNBP2 |
Unresolved candidate |
Unresolved candidate |
INV CFTR..CTTNBP2 |
GeneFusion CFTR → CTTNBP2 |
GeneFusion CTTNBP2 → CFTR |
On the reverse strand the same rules give TMPRSS2 → ERG for the chr21 deletion behind that prostate-cancer fusion.
Per transcript:
| Transcript | Effect |
|---|---|
| Wholly inside the span | ExonLoss of every exon for a deletion; unresolved DUP/INV model |
| Holds both ends | translated local consequence, unresolved model, or Intronic if no exon overlaps |
| Holds one end, fusion conditions met | GeneFusion, with the span's effect as a further candidate |
| Holds one end, no partner | the span's consequence (e.g. FrameShift or StartLoss) |
When a boundary transcript gets a GeneFusion, what the span does to its
exons isn't lost:
from varcode.effects import StructuralVariantEffect
fusion = deletion.effect_on_transcript(tmprss2)
(span_effect,) = [c.effect for c in fusion.candidates
if type(c.effect) is StructuralVariantEffect]
span_effect.affected_exons
INS and CNV never fuse. When they overlap exons, unspecified sequence or
copy-number structure yields an Unresolved candidate; it does not establish a
tandem duplication.
Local DUP/INV transcript models require every junction end to lie in the
selected transcript. An event extending outside it (including a span enclosing
the entire transcript) does not establish duplicated or inverted cDNA within
that transcript. Without a fusion partner or supplied alt_assembly, the
result contains an Unresolved candidate and mutant_transcript is None.
This means unresolved sequence, not an unchanged protein or a proven truncation.
The same restriction applies to a span candidate attached to a fusion; it does
not remove the fusion's own transcript model. A longer isoform of the same gene
does not establish the shorter isoform's structure.
Breakpoint position and protein sequence¶
The fused cDNA is the 5' partner's cDNA up to its breakpoint followed by the 3' partner's from its breakpoint on.
For an exonic join, bases inserted in the VCF breakend ALT are retained
between the two transcript segments, in transcriptional orientation.
Reciprocal records contribute the insertion once, including after
pair_breakends. variant.junction_inserted_sequence(junction) returns
the sequence while traversing the ordered junction, "" for no insertion,
or None when the records are unreadable or disagree. This follows
VCF 4.5 sections 5.4–5.4.1.
At two intronic breakpoints, the model excludes the insert under reference
splicing and records junction_insertion_status="excluded_by_reference_splicing"
in its evidence. With one exonic and one intronic breakpoint, insertion
retention requires RNA evidence: both cDNA and protein are None, with
sequence_status="unresolved_insertion_retention". Unknown/conflicting
insert sequence similarly gives sequence_status="unresolved_junction_insertion".
A supplied alt_assembly still takes precedence.
| Breakpoint | What's kept |
|---|---|
| In an intron | The 5' side ends with the last complete exon before it; the 3' side starts with the first exon after it. For CFTR intron 1 that's exon 1's 185 bases. |
| In an exon | The base at the breakpoint stays on the side that keeps it, so the cut is mid-exon: 51 bases for a breakpoint 50 bases into CFTR exon 1. |
The protein is translated from the 5' partner's start codon, through the
junction, to the first stop. That gives these cases, all still reported
as GeneFusion:
| 5' partner breakpoint | mutant_transcript.mutant_protein_sequence |
|---|---|
| Before its start codon (5' UTR or an early intron) | None. BRCA1's start codon is in exon 2, so a break in intron 1 loses it. |
| Within its coding sequence, in frame with the 3' partner | 5' partner's N-terminus plus the 3' partner's C-terminus (CPEB2::FAM193A, 1,500 aa) |
| Within its coding sequence, out of frame | 5' partner's N-terminus, then the 3' partner read in a shifted frame to the first stop (OTX1::KIF3C, 35 aa) |
| After its stop codon (3' UTR) | the 5' partner's unchanged protein (CFTR, 1,480 aa) |
There is no separate in-frame flag; compare the protein with the partners'
reference proteins. A 5' partner transcript without a complete coding
sequence also gives None.
Breakpoints near exon boundaries add candidates to any SV effect:
- Splice outcomes, when the SV's
startorendis within 6 bp of an exon edge (source="varcode_splice"). - Cryptic exons, scanned in ±500 bp windows around every breakpoint
(
source="varcode_motif"), from a chromosome FASTA when one is attached.