Skip to content

Get protein and transcript sequences

These examples go further than the home page with the TP53 transcript TP53-201. They use the human GRCh38 / Ensembl release 93 data installed on the home page and run in one Python session.

Protein sequence

from pyensembl import EnsemblRelease

data = EnsemblRelease(93, species="human")
transcript = data.transcript_by_id("ENST00000269305")
protein = transcript.protein_sequence
print(transcript.versioned_id, transcript.protein.versioned_id, len(protein))
ENST00000269305.8 ENSP00000269305.4 393

protein_sequence is the annotated translation from Ensembl's peptide file. It normally has no stop symbol, but a few peptides, mostly from polymorphic pseudogenes, contain * stop symbols. versioned_id adds this release's version to a stable ID; transcript.protein_id is the stable protein ID without it.

When you start from a protein ID, look up its sequence, transcript or gene directly:

print(data.protein_sequence(transcript.protein_id) == protein)
print(data.transcript_by_protein_id("ENSP00000269305").name)
print(data.gene_by_protein_id("ENSP00000269305").name)
True
TP53-201
TP53

Each lookup also accepts a versioned ID such as ENSP00000269305.4. A version that differs from this release's, such as ENSP00000269305.3, raises ValueError instead of returning this release's protein.

Noncoding transcripts have no protein, so their protein, protein_id and protein_sequence are None:

noncoding = data.transcript_by_id("ENST00000505014")
print(noncoding.name, noncoding.biotype, noncoding.protein_sequence)
TP53-210 retained_intron None

Coding sequence

cds = transcript.coding_sequence
print(len(cds), cds[:9], cds[-3:])
print(transcript.complete)
1182 ATGGAGGAG TGA
True

The coding sequence runs from the first base of the start codon through the stop codon, so its 1,182 bases are 393 codons plus the stop. complete is True when the transcript has annotated three-base start and stop codons and a coding length divisible by three.

Transcript sequence and UTRs

transcript.sequence is the full spliced cDNA, oriented 5′ to 3′. For a transcript with annotated start and stop codons, it is the 5′ UTR, the coding sequence and the 3′ UTR joined together:

utr5 = transcript.five_prime_utr_sequence
utr3 = transcript.three_prime_utr_sequence
print(len(transcript.sequence), len(utr5), len(cds), len(utr3))
print(transcript.sequence == utr5 + cds + utr3)
2579 190 1182 1207
True

To include introns or flanking sequence, read genomic DNA from the same assembly.

Incomplete transcripts

A protein_coding biotype does not guarantee a complete coding sequence. Some transcripts are fragments without an annotated start or stop codon:

fragment = data.transcript_by_id("ENST00000576024")
print(fragment.name, fragment.biotype, fragment.complete)
print(fragment.contains_start_codon, fragment.coding_sequence)
print(fragment.protein_sequence[:10])
TP53-216 protein_coding False
False None
XSPQPKKKPL

Without an annotated start codon, coding_sequence and five_prime_utr_sequence are None. The peptide file can still hold a partial translation, so check complete before analyzing codons or reading frames.

Exons and genomic coordinates

Exons

print(transcript.contig, transcript.start, transcript.end, transcript.strand)
for exon in transcript.exons[:2]:
    print(exon.id, exon.start, exon.end)
print(len(transcript.exons))
17 7668402 7687538 -
ENSE00001146308 7687377 7687538
ENSE00002667911 7676521 7676622
11

The transcript's span includes its introns. transcript.exons is in transcription order, so on the minus strand the first exon has the highest coordinates. Every start is still the lower coordinate, and coordinates are one-based and inclusive.

Codon positions and coding intervals

print(transcript.start_codon_positions)
print(transcript.stop_codon_positions)
print(transcript.coding_sequence_position_ranges[:2])
[7676592, 7676593, 7676594]
[7669609, 7669610, 7669611]
[(7669609, 7669690), (7670609, 7670715)]

These are genomic positions in ascending order. On the minus strand, translation starts at the highest position of the start codon, 7676594. coding_sequence_position_ranges lists the coding part of each exon, including the stop codon, which Ensembl annotates separately from the CDS.

Map a genomic position to the transcript

spliced_offset converts a genomic position inside an exon to a zero-based index into transcript.sequence:

offset = transcript.spliced_offset(7676594)
print(offset, transcript.sequence[offset:offset + 3])
data.close()
190 ATG

The first base of the start codon is at offset 190, immediately after the 5′ UTR. Positions in introns or outside the transcript raise ValueError. The Transcript reference lists every attribute.