Skip to content

Experimental annotators

Most users don't need this page. The default annotator already predicts protein sequences and handles structural variants; use effects() without selecting an implementation.

Varcode also includes two opt-in implementations, protein_diff and transcript_model, for comparing predictions. They do not support every input the default handles.

Supported inputs

Selection Method Inputs
Default (omit annotator=) Predicts small-edit and structural consequences; compares small edits against a patient baseline when supplied Small variants and SVs. Germline/phase context for small edits, but not general SV-plus-germline combinations.
annotator="protein_diff" Edits and translates transcripts, then compares proteins; shares the default's splice, location, and germline helpers Small variants. SVs are unsupported.
annotator="transcript_model" Builds transcripts for phase/splice hypotheses, compares each with the patient baseline, and merges equivalent results Small variants and local DEL/DUP/INV. Insertions need alt_assembly; CNVs are unsupported.

The transcript-model experiment delegates BNDs to the structural helper; it does not resolve BND-plus-haplotype combinations. DUP/INV events that extend beyond its layout are left unresolved; see limitations.

The default's registry name is fast; it appears in provenance but need not be passed explicitly. protein_diff is an alternative implementation, not an option you need for protein output.

Comparing annotators

Given a variant and one of its transcript objects, selection is explicit:

effect = variant.effect_on_transcript(transcript)  # ordinary use
comparison = variant.effect_on_transcript(transcript, annotator="protein_diff")
experimental = variant.effect_on_transcript(transcript, annotator="transcript_model")

The transcript model can use germline= and phase_resolver= context. Canonical and exon-skip paths need only transcript annotation; a genome with reference FASTA additionally supplies sequence for intron retention and cryptic splice sites. Without a calibrated scorer, mechanism preference is an ordering rule, not a probability. Selecting this experiment does not guarantee that every structural or combined input is supported.

Through effects(germline=...), the transcript model uses the same phase cap as the default path (GermlineContext.max_phase_hypotheses, default 8). It also caps the combined phase × splice outcomes at max_hypotheses=64, which is the phase cap too when predict_transcript_model_effect is called directly. Exceeding either returns a HypothesisLimit rather than a partial set of candidates.

Mixed inherited and somatic haplotypes

predict_transcript_model_effect(variants, transcript, germline_variants=...) and annotate_haplotype take a known-cis group. Members matching the supplied germline are included once in the patient baseline. Only novel alleles are applied to that baseline to obtain the mutant; effect.variant is the first novel member and effect.variants retains the supplied group. Matching uses reference assembly, contig, normalized position, REF and ALT. A group with no novel allele returns GermlineAlleleOverlap, without inferring LOH.

The group itself supplies cis constraints. Phase can propagate from any member to other germline alleles; unlinked alleles still produce separate hypotheses. Homozygous alleles are on both haplotypes and do not bridge their phase. As with other phase constraints, conflicting resolver answers are logged and ignored. Each hypothesis retains its germline cis/trans assignments, phase state, resolver source and novel somatic_variants in its evidence.

VariantCollection.effects constructs these groups only when the resolver supports cis. Unknown or trans relationships retain individual predictions; listing alleles in a collection alone does not assert that they are in cis.

Transcript-model results

Use the sequence accessors for a single effect, including the same checks for missing sequences and partial structures.

The ordinary and experimental candidate wrappers are not interchangeable:

Candidate type Shared access Additional information
EffectCandidate (ordinary splice/SV/RNA outcome sets) candidate.effect source, evidence; sequences are on the effect or its optional mutant_transcript
RealizedEffectCandidate (transcript-model hypothesis pipeline) candidate.effect outcomes, hypotheses, probability, ordinal_key; each outcome has baseline and mutant products with cdna_sequence, protein_sequence and evidence

The experiment returns an ordinary top effect with candidates attached, not necessarily a MultiOutcomeEffect. Its BND delegation uses ordinary candidates; early noncoding, incomplete or unresolved results may have no candidates at all. Inspect the returned shape rather than assuming it from the annotator name:

from varcode import EffectCandidate
from varcode.effect_hypotheses import RealizedEffectCandidate

print(experimental.short_description)
for candidate in getattr(experimental, "candidates", ()):
    print(candidate.effect.short_description)
    if isinstance(candidate, RealizedEffectCandidate):
        # A merged candidate can retain several hypotheses and their products.
        for outcome in candidate.outcomes:
            print(outcome.baseline.protein_sequence, outcome.mutant.protein_sequence)
            print(outcome.mutant.cdna_sequence, outcome.hypothesis.evidence)
    elif isinstance(candidate, EffectCandidate):
        print(candidate.source, candidate.evidence)
        print(candidate.effect.mutant_protein_sequence)

Known limitations

  • Combined variants: the selected annotator makes known-cis joint predictions through VariantCollection.effects(phase_resolver=...). The default and protein_diff build point-edit haplotypes; joint germline and structural composition requires the opt-in transcript_model. Unsupported groups remain explicit unresolved results. See the joint contract.
  • Rearrangements beyond the layout: the transcript model leaves an overlapping DUP/INV unresolved when either boundary extends beyond its finite genomic layout, because a clipped copy of the interval cannot establish the complete rearranged transcript or an unchanged protein (#449). Fully represented local events still use the layout model, and supplied transcript assemblies keep their sequence, with unmapped CDS/translation uncertainty preserved.

Unresolved results report None, not False, for sequence-change flags. drop_silent_and_noncoding() retains them by default; see SV filtering.

Previous names

realized, RealizedEffectAnnotator, and predict_realized_effect are aliases for the transcript-model implementation. Use the transcript_model names in new code. The RealizedEffectCandidate wrapper still has a different interface from ordinary EffectCandidate, as shown above.

To implement your own, see Writing an annotator.