Experimental annotators¶
Most users don't need this page. The default annotator already predicts protein sequences and handles structural variants; use effects() without selecting an implementation.
Varcode also includes two opt-in implementations, protein_diff and
transcript_model, for comparing predictions. They do not support every input
the default handles.
Supported inputs¶
| Selection | Method | Inputs |
|---|---|---|
Default (omit annotator=) |
Predicts small-edit and structural consequences; compares small edits against a patient baseline when supplied | Small variants and SVs. Germline/phase context for small edits, but not general SV-plus-germline combinations. |
annotator="protein_diff" |
Edits and translates transcripts, then compares proteins; shares the default's splice, location, and germline helpers | Small variants. SVs are unsupported. |
annotator="transcript_model" |
Builds transcripts for phase/splice hypotheses, compares each with the patient baseline, and merges equivalent results | Small variants and local DEL/DUP/INV. Insertions need alt_assembly; CNVs are unsupported. |
The transcript-model experiment delegates BNDs to the structural helper; it does not resolve BND-plus-haplotype combinations. DUP/INV events that extend beyond its layout are left unresolved; see limitations.
The default's registry name is fast; it appears in provenance but need not be
passed explicitly. protein_diff is an alternative implementation, not an
option you need for protein output.
Comparing annotators¶
Given a variant and one of its transcript objects, selection is explicit:
effect = variant.effect_on_transcript(transcript) # ordinary use
comparison = variant.effect_on_transcript(transcript, annotator="protein_diff")
experimental = variant.effect_on_transcript(transcript, annotator="transcript_model")
The transcript model can use germline= and phase_resolver= context. Canonical
and exon-skip paths need only transcript annotation; a genome with reference
FASTA additionally supplies sequence for intron retention and cryptic splice
sites. Without a calibrated scorer, mechanism preference is an ordering rule,
not a probability. Selecting this experiment does not guarantee that every
structural or combined input is supported.
Through effects(germline=...), the transcript model uses the same phase cap
as the default path (GermlineContext.max_phase_hypotheses, default 8). It
also caps the combined phase × splice outcomes at max_hypotheses=64, which is
the phase cap too when predict_transcript_model_effect is called directly.
Exceeding either returns a HypothesisLimit
rather than a partial set of candidates.
Mixed inherited and somatic haplotypes¶
predict_transcript_model_effect(variants, transcript, germline_variants=...)
and annotate_haplotype take a known-cis group. Members matching the supplied
germline are included once in the patient baseline. Only novel alleles are
applied to that baseline to obtain the mutant; effect.variant is the first
novel member and effect.variants retains the supplied group. Matching uses
reference assembly, contig, normalized position, REF and ALT. A group with no
novel allele returns GermlineAlleleOverlap, without inferring LOH.
The group itself supplies cis constraints. Phase can propagate from any member
to other germline alleles; unlinked alleles still produce separate hypotheses.
Homozygous alleles are on both haplotypes and do not bridge their phase.
As with other phase constraints, conflicting resolver answers are logged and
ignored. Each hypothesis retains its germline cis/trans assignments, phase
state, resolver source and novel somatic_variants in its evidence.
VariantCollection.effects constructs these groups only when the resolver
supports cis. Unknown or trans relationships retain individual predictions;
listing alleles in a collection alone does not assert that they are in cis.
Transcript-model results¶
Use the sequence accessors for a single effect, including the same checks for missing sequences and partial structures.
The ordinary and experimental candidate wrappers are not interchangeable:
| Candidate type | Shared access | Additional information |
|---|---|---|
EffectCandidate (ordinary splice/SV/RNA outcome sets) |
candidate.effect |
source, evidence; sequences are on the effect or its optional mutant_transcript |
RealizedEffectCandidate (transcript-model hypothesis pipeline) |
candidate.effect |
outcomes, hypotheses, probability, ordinal_key; each outcome has baseline and mutant products with cdna_sequence, protein_sequence and evidence |
The experiment returns an ordinary top effect with candidates attached, not
necessarily a MultiOutcomeEffect. Its BND delegation uses ordinary candidates;
early noncoding, incomplete or unresolved results may have no candidates at all.
Inspect the returned shape rather than assuming it from the annotator name:
from varcode import EffectCandidate
from varcode.effect_hypotheses import RealizedEffectCandidate
print(experimental.short_description)
for candidate in getattr(experimental, "candidates", ()):
print(candidate.effect.short_description)
if isinstance(candidate, RealizedEffectCandidate):
# A merged candidate can retain several hypotheses and their products.
for outcome in candidate.outcomes:
print(outcome.baseline.protein_sequence, outcome.mutant.protein_sequence)
print(outcome.mutant.cdna_sequence, outcome.hypothesis.evidence)
elif isinstance(candidate, EffectCandidate):
print(candidate.source, candidate.evidence)
print(candidate.effect.mutant_protein_sequence)
Known limitations¶
- Combined variants: the selected annotator makes known-cis joint predictions
through
VariantCollection.effects(phase_resolver=...). The default andprotein_diffbuild point-edit haplotypes; joint germline and structural composition requires the opt-intranscript_model. Unsupported groups remain explicit unresolved results. See the joint contract. - Rearrangements beyond the layout: the transcript model leaves an overlapping DUP/INV unresolved when either boundary extends beyond its finite genomic layout, because a clipped copy of the interval cannot establish the complete rearranged transcript or an unchanged protein (#449). Fully represented local events still use the layout model, and supplied transcript assemblies keep their sequence, with unmapped CDS/translation uncertainty preserved.
Unresolved results report None, not False, for sequence-change flags.
drop_silent_and_noncoding() retains them by default; see
SV filtering.
Previous names¶
realized, RealizedEffectAnnotator, and predict_realized_effect are aliases
for the transcript-model implementation. Use the transcript_model names in
new code. The RealizedEffectCandidate wrapper still has a different interface
from ordinary EffectCandidate, as shown above.
To implement your own, see Writing an annotator.