Skip to content

PyEnsembl

PyEnsembl lets you find where genes are located, which transcripts they produce, and their sequences in Python. It uses Ensembl annotations or your own files. Once the data is installed, queries run locally.

Install

Use Python 3.9 or later. Run these commands in a terminal:

python -m pip install pyensembl
pyensembl install --release 93 --species human

The second command downloads and indexes human annotation, transcript and protein data. It can take several minutes. These examples use GRCh38 and Ensembl release 93 so you can reproduce the output; for your own analysis, choose the reference that matches your data.

Look up a gene

Run this in Python after installation. It retrieves TP53 by its gene ID:

from pyensembl import EnsemblRelease

data = EnsemblRelease(93, species="human")
gene = data.gene_by_id("ENSG00000141510")
print(gene.id, gene.name, gene.contig, gene.start, gene.end, gene.strand)
ENSG00000141510 TP53 17 7661779 7687550 -

The result gives the gene ID, name, chromosome, start, end and strand. TP53 is on chromosome 17, on the minus strand. Coordinates are one-based and include both ends; start is the lower coordinate on either strand.

You can also search with data.genes_by_name("TP53"). It returns a list because a name can match several genes. Gene aliases let you search additional names such as p53.

Read protein and transcript sequences

A gene can have several transcripts. Continue in the same Python session with one selected transcript:

transcript = data.transcript_by_id("ENST00000269305")
print(transcript.protein_sequence[:30])
print(transcript.sequence[:30])
print(len(transcript.protein_sequence), len(transcript.sequence))
data.close()
MEEPQSDPSVEPPLSQETFSDLWKLLPENN
GTTTTCCCCTCCCATGTGCTCAAGACTGGC
393 2579

The first two lines show the first 30 amino acids of the protein and the first 30 bases of the transcript's cDNA, which begins with the 5′ UTR. The full sequences have 393 amino acids and 2,579 bases. The cDNA is spliced and already oriented 5′ to 3′, including for minus-strand transcripts. Noncoding or incomplete transcripts may have no protein sequence. Both come from the reference annotation, not from your samples.

data.close() releases the open index files. You can instead write with EnsemblRelease(93, species="human") as data: to close them automatically.

Next steps