Skip to content

Feature objects

Gene and transcript coordinates refer to the selected annotation, with one-based inclusive bounds. The guides show how to find genes and transcripts and get protein and transcript sequences.

Gene

Gene(gene_id, gene_name, contig, start, end, strand, biotype, genome)

contig instance-attribute

contig = normalize_chromosome(contig)

strand instance-attribute

strand = normalize_strand(strand)

start instance-attribute

start = start

end instance-attribute

end = end

length property

length

on_forward_strand property

on_forward_strand

on_positive_strand property

on_positive_strand

on_backward_strand property

on_backward_strand

on_negative_strand property

on_negative_strand

genome instance-attribute

genome = genome

biotype instance-attribute

biotype = biotype

db property

db

Resolve annotation only when needed, preserving metadata-only use.

is_protein_coding property

is_protein_coding

True iff this entry's biotype is the canonical "protein_coding". Conservative by design - this is what downstream effect predictors (e.g. varcode) read to decide whether a variant lands in a translatable transcript. See :attr:is_protein_coding_extended for a wider definition that also covers IG/TR gene segments and translated pseudogenes, and :attr:is_translated for the widest definition that additionally includes NMD/NSD targets.

is_protein_coding_extended property

is_protein_coding_extended

True for any biotype that makes a stable, functional protein product:

  • protein_coding (canonical)
  • IG_{C,D,J,V}_gene and TR_{C,D,J,V}_gene (immunoglobulin and T-cell receptor gene segments — produce protein after V(D)J recombination)
  • polymorphic_pseudogene (codes in some individuals)
  • translated_{processed,unprocessed}_pseudogene (pseudogenes with translation evidence)

Excludes nonsense_mediated_decay and non_stop_decay — those biotypes are translated but the product is targeted for degradation rather than stable accumulation. Use :attr:is_translated if you want those.

is_translated property

is_translated

True for any biotype that gets translated on a ribosome, even transiently. Equivalent to :attr:is_protein_coding_extended plus nonsense_mediated_decay and non_stop_decay.

Useful when you care about whether a variant lands in a translated frame at all (e.g. picking a top variant effect, RNA-seq peptide analysis) rather than whether the encoded protein product is stably expressed.

gene_id instance-attribute

gene_id = gene_id

gene_name instance-attribute

gene_name = gene_name

id property

id

Alias for gene_id necessary for backwards compatibility.

name property

name

Alias for gene_name necessary for backwards compatibility.

version property

version

Alias for :attr:gene_version.

versioned_gene_id property

versioned_gene_id

gene_id.gene_version when a version is available, else gene_id.

versioned_id property

versioned_id

Alias for :attr:versioned_gene_id.

to_tuple

to_tuple()

offset

offset(position)

Offset of given position from stranded start of this locus.

For example, if a Locus goes from 10..20 and is on the negative strand, then the offset of position 13 is 7, whereas if the Locus is on the positive strand, then the offset is 3.

offset_range

offset_range(start, end)

Database start/end entries are always ordered such that start < end. This makes computing a relative position (e.g. of a stop codon relative to its transcript) complicated since the "end" position of a backwards locus is actually earlir on the strand. This function correctly selects a start vs. end value depending on this locuses's strand and determines that position's offset from the earliest position in this locus.

on_contig

on_contig(contig)

on_strand

on_strand(strand)

can_overlap

can_overlap(contig, strand=None)

Is this locus on the same contig and (optionally) on the same strand?

distance_to_interval

distance_to_interval(start, end)

Find the distance between intervals [start1, end1] and [start2, end2]. If the intervals overlap then the distance is 0.

distance_to_locus

distance_to_locus(other, ignore_strand=False)

Distance between two loci. Returns infinity for loci on different contigs, and (by default) for loci on opposite strands of the same contig. Pass ignore_strand=True to compute a strand-invariant distance between loci on the same contig.

overlaps

overlaps(contig, start, end, strand=None)

Does this locus overlap with a given range of positions?

Since locus position ranges are inclusive, we should make sure that e.g. chr1:10-10 overlaps with chr1:10-10

overlaps_locus

overlaps_locus(other_locus)

overlap_length

overlap_length(other, ignore_strand=False)

Number of overlapping base positions between this locus and other. Returns 0 when the loci do not overlap, are on different contigs, or (by default) are on opposite strands. Pass ignore_strand=True to count overlap regardless of strand.

intersect

intersect(other, ignore_strand=False)

Return a new :class:Locus covering the inclusive-inclusive overlap between this locus and other, or None if they do not overlap. Mirrors bedtools' intersect for a single pair.

Different contigs always return None. Different strands return None unless ignore_strand=True is passed, in which case the result is reported on this locus's strand.

contains

contains(contig, start, end, strand=None)

contains_locus

contains_locus(other_locus)

to_dict

to_dict()

gene_version

gene_version()

Ensembl annotation version for this gene (int) or None if the GTF did not provide a gene_version attribute.

transcripts

transcripts()

Property which dynamically construct transcript objects for all transcript IDs associated with this gene.

exons

exons()

Transcript

Transcript(transcript_id, transcript_name, contig, start, end, strand, biotype, gene_id, genome, support_level=None)

Transcript encompasses the locus, exons, and sequence of a transcript.

Lazily fetches sequence in case we"re constructing many Transcripts and not using the sequence, avoid the memory/performance overhead of fetching and storing sequences from a FASTA file.

contig instance-attribute

contig = normalize_chromosome(contig)

strand instance-attribute

strand = normalize_strand(strand)

start instance-attribute

start = start

end instance-attribute

end = end

length property

length

on_forward_strand property

on_forward_strand

on_positive_strand property

on_positive_strand

on_backward_strand property

on_backward_strand

on_negative_strand property

on_negative_strand

genome instance-attribute

genome = genome

biotype instance-attribute

biotype = biotype

db property

db

Resolve annotation only when needed, preserving metadata-only use.

is_protein_coding property

is_protein_coding

True iff this entry's biotype is the canonical "protein_coding". Conservative by design - this is what downstream effect predictors (e.g. varcode) read to decide whether a variant lands in a translatable transcript. See :attr:is_protein_coding_extended for a wider definition that also covers IG/TR gene segments and translated pseudogenes, and :attr:is_translated for the widest definition that additionally includes NMD/NSD targets.

is_protein_coding_extended property

is_protein_coding_extended

True for any biotype that makes a stable, functional protein product:

  • protein_coding (canonical)
  • IG_{C,D,J,V}_gene and TR_{C,D,J,V}_gene (immunoglobulin and T-cell receptor gene segments — produce protein after V(D)J recombination)
  • polymorphic_pseudogene (codes in some individuals)
  • translated_{processed,unprocessed}_pseudogene (pseudogenes with translation evidence)

Excludes nonsense_mediated_decay and non_stop_decay — those biotypes are translated but the product is targeted for degradation rather than stable accumulation. Use :attr:is_translated if you want those.

is_translated property

is_translated

True for any biotype that gets translated on a ribosome, even transiently. Equivalent to :attr:is_protein_coding_extended plus nonsense_mediated_decay and non_stop_decay.

Useful when you care about whether a variant lands in a translated frame at all (e.g. picking a top variant effect, RNA-seq peptide analysis) rather than whether the encoded protein product is stably expressed.

transcript_id instance-attribute

transcript_id = transcript_id

transcript_name instance-attribute

transcript_name = transcript_name

gene_id instance-attribute

gene_id = gene_id

support_level instance-attribute

support_level = support_level

id property

id

Alias for transcript_id necessary for backward compatibility.

name property

name

Alias for transcript_name necessary for backward compatibility.

gene property

gene

gene_name property

gene_name

exons property

exons

version property

version

Alias for :attr:transcript_version.

versioned_transcript_id property

versioned_transcript_id

transcript_id.transcript_version when available, else transcript_id.

versioned_id property

versioned_id

Alias for :attr:versioned_transcript_id.

fasta_version property

fasta_version

Integer version that the cDNA FASTA header carried for this transcript, or None if the FASTA didn't carry one (older Ensembl releases shipped bare headers) or no transcript FASTA is attached to the genome.

Differs from :attr:transcript_version (which comes from the GTF's transcript_version attribute). When the two disagree, the FASTA version is the authoritative source-of-truth for the bytes returned by :attr:sequence.

to_tuple

to_tuple()

offset

offset(position)

Offset of given position from stranded start of this locus.

For example, if a Locus goes from 10..20 and is on the negative strand, then the offset of position 13 is 7, whereas if the Locus is on the positive strand, then the offset is 3.

offset_range

offset_range(start, end)

Database start/end entries are always ordered such that start < end. This makes computing a relative position (e.g. of a stop codon relative to its transcript) complicated since the "end" position of a backwards locus is actually earlir on the strand. This function correctly selects a start vs. end value depending on this locuses's strand and determines that position's offset from the earliest position in this locus.

on_contig

on_contig(contig)

on_strand

on_strand(strand)

can_overlap

can_overlap(contig, strand=None)

Is this locus on the same contig and (optionally) on the same strand?

distance_to_interval

distance_to_interval(start, end)

Find the distance between intervals [start1, end1] and [start2, end2]. If the intervals overlap then the distance is 0.

distance_to_locus

distance_to_locus(other, ignore_strand=False)

Distance between two loci. Returns infinity for loci on different contigs, and (by default) for loci on opposite strands of the same contig. Pass ignore_strand=True to compute a strand-invariant distance between loci on the same contig.

overlaps

overlaps(contig, start, end, strand=None)

Does this locus overlap with a given range of positions?

Since locus position ranges are inclusive, we should make sure that e.g. chr1:10-10 overlaps with chr1:10-10

overlaps_locus

overlaps_locus(other_locus)

overlap_length

overlap_length(other, ignore_strand=False)

Number of overlapping base positions between this locus and other. Returns 0 when the loci do not overlap, are on different contigs, or (by default) are on opposite strands. Pass ignore_strand=True to count overlap regardless of strand.

intersect

intersect(other, ignore_strand=False)

Return a new :class:Locus covering the inclusive-inclusive overlap between this locus and other, or None if they do not overlap. Mirrors bedtools' intersect for a single pair.

Different contigs always return None. Different strands return None unless ignore_strand=True is passed, in which case the result is reported on this locus's strand.

contains

contains(contig, start, end, strand=None)

contains_locus

contains_locus(other_locus)

to_dict

to_dict()

contains_start_codon

contains_start_codon()

Does this transcript have an annotated start_codon entry?

contains_stop_codon

contains_stop_codon()

Does this transcript have an annotated stop_codon entry?

start_codon_complete

start_codon_complete()

Does the start codon span 3 genomic positions?

stop_codon_complete

stop_codon_complete()

Does the stop codon span 3 genomic positions?

start_codon_positions

start_codon_positions()

Chromosomal positions of nucleotides in start codon.

stop_codon_positions

stop_codon_positions()

Chromosomal positions of nucleotides in stop codon.

exon_intervals

exon_intervals()

List of (start,end) tuples for each exon of this transcript, in the order specified by the 'exon_number' column of the exon table.

spliced_offset

spliced_offset(position)

Convert from an absolute chromosomal position to the offset into this transcript's spliced mRNA.

Returns a zero-based index into sequence. Position must be inside an exon; otherwise raises ValueError.

start_codon_unspliced_offsets

start_codon_unspliced_offsets()

Offsets from start of unspliced pre-mRNA transcript of nucleotides in start codon.

stop_codon_unspliced_offsets

stop_codon_unspliced_offsets()

Offsets from start of unspliced pre-mRNA transcript of nucleotides in stop codon.

start_codon_spliced_offsets

start_codon_spliced_offsets()

Offsets from start of spliced mRNA transcript of nucleotides in start codon.

stop_codon_spliced_offsets

stop_codon_spliced_offsets()

Offsets from start of spliced mRNA transcript of nucleotides in stop codon.

coding_sequence_position_ranges

coding_sequence_position_ranges()

Return absolute chromosome position ranges for CDS fragments of this transcript, including the stop codon (which Ensembl encodes as a separate feature from the CDS).

complete

complete()

Consider a transcript complete if it has three-base start and stop codons and a coding sequence whose length is divisible by 3

sequence

sequence()

Spliced cDNA sequence of transcript (includes 5" UTR, coding sequence, and 3" UTR)

first_start_codon_spliced_offset

first_start_codon_spliced_offset()

Offset of first nucleotide in start codon into the spliced mRNA (excluding introns)

last_stop_codon_spliced_offset

last_stop_codon_spliced_offset()

Offset of last nucleotide in stop codon into the spliced mRNA (excluding introns)

coding_sequence

coding_sequence()

cDNA coding sequence (from start codon to stop codon, without any introns)

five_prime_utr_sequence

five_prime_utr_sequence()

cDNA sequence of 5' UTR (untranslated region at the beginning of the transcript)

three_prime_utr_sequence

three_prime_utr_sequence()

cDNA sequence of 3' UTR (untranslated region at the end of the transcript)

transcript_version

transcript_version()

Ensembl annotation version for this transcript (int) or None if the GTF did not provide a transcript_version attribute.

protein_id

protein_id()

protein

protein()

:class:Protein view object for this transcript's encoded protein, or None if this transcript is non-coding.

protein_sequence

protein_sequence()

Exon

Exon(exon_id, contig, start, end, strand, gene_name, gene_id, exon_version=None)

contig instance-attribute

contig = normalize_chromosome(contig)

strand instance-attribute

strand = normalize_strand(strand)

start instance-attribute

start = start

end instance-attribute

end = end

length property

length

on_forward_strand property

on_forward_strand

on_positive_strand property

on_positive_strand

on_backward_strand property

on_backward_strand

on_negative_strand property

on_negative_strand

exon_id instance-attribute

exon_id = exon_id

gene_name instance-attribute

gene_name = gene_name

gene_id instance-attribute

gene_id = gene_id

exon_version instance-attribute

exon_version = exon_version

id property

id

Alias for exon_id necessary for backward compatibility.

version property

version

Alias for :attr:exon_version.

versioned_exon_id property

versioned_exon_id

exon_id.exon_version when available, else exon_id.

versioned_id property

versioned_id

Alias for :attr:versioned_exon_id.

to_tuple

to_tuple()

offset

offset(position)

Offset of given position from stranded start of this locus.

For example, if a Locus goes from 10..20 and is on the negative strand, then the offset of position 13 is 7, whereas if the Locus is on the positive strand, then the offset is 3.

offset_range

offset_range(start, end)

Database start/end entries are always ordered such that start < end. This makes computing a relative position (e.g. of a stop codon relative to its transcript) complicated since the "end" position of a backwards locus is actually earlir on the strand. This function correctly selects a start vs. end value depending on this locuses's strand and determines that position's offset from the earliest position in this locus.

on_contig

on_contig(contig)

on_strand

on_strand(strand)

can_overlap

can_overlap(contig, strand=None)

Is this locus on the same contig and (optionally) on the same strand?

distance_to_interval

distance_to_interval(start, end)

Find the distance between intervals [start1, end1] and [start2, end2]. If the intervals overlap then the distance is 0.

distance_to_locus

distance_to_locus(other, ignore_strand=False)

Distance between two loci. Returns infinity for loci on different contigs, and (by default) for loci on opposite strands of the same contig. Pass ignore_strand=True to compute a strand-invariant distance between loci on the same contig.

overlaps

overlaps(contig, start, end, strand=None)

Does this locus overlap with a given range of positions?

Since locus position ranges are inclusive, we should make sure that e.g. chr1:10-10 overlaps with chr1:10-10

overlaps_locus

overlaps_locus(other_locus)

overlap_length

overlap_length(other, ignore_strand=False)

Number of overlapping base positions between this locus and other. Returns 0 when the loci do not overlap, are on different contigs, or (by default) are on opposite strands. Pass ignore_strand=True to count overlap regardless of strand.

intersect

intersect(other, ignore_strand=False)

Return a new :class:Locus covering the inclusive-inclusive overlap between this locus and other, or None if they do not overlap. Mirrors bedtools' intersect for a single pair.

Different contigs always return None. Different strands return None unless ignore_strand=True is passed, in which case the result is reported on this locus's strand.

contains

contains(contig, start, end, strand=None)

contains_locus

contains_locus(other_locus)

to_dict

to_dict()

Protein

Protein(protein_id, protein_version=None, genome=None)

Lightweight view object exposing the protein identity for a transcript. Accessed via :attr:Transcript.protein.

protein_id instance-attribute

protein_id = protein_id

protein_version instance-attribute

protein_version = protein_version

genome instance-attribute

genome = genome

id property

id

Alias for :attr:protein_id.

version property

version

Alias for :attr:protein_version.

versioned_protein_id property

versioned_protein_id

protein_id.protein_version when available, else protein_id.

versioned_id property

versioned_id

Alias for :attr:versioned_protein_id.

fasta_version property

fasta_version

Integer version that the protein FASTA header carried for this protein_id, or None if the FASTA didn't carry one (older Ensembl releases shipped bare headers) or the genome has no protein FASTA attached.

Differs from :attr:protein_version (which comes from the GTF's protein_version attribute). When the two disagree, the FASTA version is the authoritative source-of-truth for the bytes returned by :meth:Transcript.protein_sequence.

to_dict

to_dict()

Locus

Locus(contig, start, end, strand)

Base class for any entity which can be localized at a range of positions on a particular strand of a chromosome/contig.

contig : str Chromosome or other sequence name in the reference assembly

start : int Start position of locus on the contig

end : int Inclusive end position on the contig

strand : str Should we read the locus forwards ('+') or backwards ('-')?

contig instance-attribute

contig = normalize_chromosome(contig)

strand instance-attribute

strand = normalize_strand(strand)

start instance-attribute

start = start

end instance-attribute

end = end

length property

length

on_forward_strand property

on_forward_strand

on_positive_strand property

on_positive_strand

on_backward_strand property

on_backward_strand

on_negative_strand property

on_negative_strand

to_tuple

to_tuple()

to_dict

to_dict()

offset

offset(position)

Offset of given position from stranded start of this locus.

For example, if a Locus goes from 10..20 and is on the negative strand, then the offset of position 13 is 7, whereas if the Locus is on the positive strand, then the offset is 3.

offset_range

offset_range(start, end)

Database start/end entries are always ordered such that start < end. This makes computing a relative position (e.g. of a stop codon relative to its transcript) complicated since the "end" position of a backwards locus is actually earlir on the strand. This function correctly selects a start vs. end value depending on this locuses's strand and determines that position's offset from the earliest position in this locus.

on_contig

on_contig(contig)

on_strand

on_strand(strand)

can_overlap

can_overlap(contig, strand=None)

Is this locus on the same contig and (optionally) on the same strand?

distance_to_interval

distance_to_interval(start, end)

Find the distance between intervals [start1, end1] and [start2, end2]. If the intervals overlap then the distance is 0.

distance_to_locus

distance_to_locus(other, ignore_strand=False)

Distance between two loci. Returns infinity for loci on different contigs, and (by default) for loci on opposite strands of the same contig. Pass ignore_strand=True to compute a strand-invariant distance between loci on the same contig.

overlaps

overlaps(contig, start, end, strand=None)

Does this locus overlap with a given range of positions?

Since locus position ranges are inclusive, we should make sure that e.g. chr1:10-10 overlaps with chr1:10-10

overlaps_locus

overlaps_locus(other_locus)

overlap_length

overlap_length(other, ignore_strand=False)

Number of overlapping base positions between this locus and other. Returns 0 when the loci do not overlap, are on different contigs, or (by default) are on opposite strands. Pass ignore_strand=True to count overlap regardless of strand.

intersect

intersect(other, ignore_strand=False)

Return a new :class:Locus covering the inclusive-inclusive overlap between this locus and other, or None if they do not overlap. Mirrors bedtools' intersect for a single pair.

Different contigs always return None. Different strands return None unless ignore_strand=True is passed, in which case the result is reported on this locus's strand.

contains

contains(contig, start, end, strand=None)

contains_locus

contains_locus(other_locus)

find_nearest_locus

find_nearest_locus(start, end, loci)

Finds nearest locus (object with method distance_to_interval) to the interval defined by the given start and end positions. Returns the distance to that locus, along with the locus object itself.