Skip to content

Find genes and transcripts

These examples use the human GRCh38 / Ensembl release 93 data installed on the home page. They run in one Python session.

Find a gene by name or ID

A name lookup returns a list of every matching gene:

from pyensembl import EnsemblRelease

data = EnsemblRelease(93, species="human")
for gene in data.genes_by_name("TP53"):
    print(gene.id, gene.name, gene.contig, gene.biotype)
ENSG00000141510 TP53 17 protein_coding

A name can match several loci, such as copies on patch or haplotype contigs. When you know the gene ID, data.gene_by_id("ENSG00000141510") selects it directly. Alias lookup adds names from a separate source, such as HGNC.

Versioned IDs

Ensembl IDs can carry a version, which changes when a feature's structure or sequence changes. Every method that takes a gene, transcript, exon or protein ID accepts it with or without one:

gene = data.gene_by_id("ENSG00000141510.16")
print(gene.id, gene.versioned_id)
ENSG00000141510 ENSG00000141510.16

An ID without a version matches whatever version this release has. An ID with a version must match it: data.gene_by_id("ENSG00000141510.15") raises ValueError: ENSG00000141510.15 is not in this annotation, which has ENSG00000141510.16, so an ID copied from another release can't silently match a changed feature. Releases before Ensembl 77 record no versions, so they accept only IDs without one.

Find genes at a position

Give a position, or add end to search an interval:

print(data.gene_names_at_locus(contig="17", position=7668402))
for gene in data.genes_at_locus(contig="17", position=7660000, end=7690000):
    print(gene.name, gene.strand, gene.start, gene.end)
['TP53']
WRAP53 + 7686071 7703502
TP53 - 7661779 7687550
AC087388.1 - 7685260 7686371

Coordinates are one-based and inclusive, and must use the selected assembly. Overlapping genes on both strands are returned; pass strand="+" or "-" to restrict the search. transcripts_at_locus and exons_at_locus work the same way, and nearest_gene finds the closest gene when none overlaps.

Filter by contig or biotype

Whole-genome lists accept optional contig, strand and biotype filters:

coding_genes = data.genes(contig="17", biotype="protein_coding")
print(len(coding_genes))
1183

Biotype names come from the annotation and can differ between releases; list them with {gene.biotype for gene in data.genes()}. gene_ids, gene_names and transcript_ids return identifiers without building objects, and data.contigs() lists the chromosome and contig names.

Choose among a gene's transcripts

A gene usually has several transcripts. List order does not identify a preferred isoform, so choose by ID or by attributes that suit your analysis:

gene = data.gene_by_id("ENSG00000141510")
coding = [t for t in gene.transcripts
          if t.biotype == "protein_coding" and t.complete]
print(len(gene.transcripts), len(coding))
coding.sort(key=lambda t: len(t.protein_sequence), reverse=True)
for t in coding[:3]:
    print(t.id, t.name, t.support_level, len(t.protein_sequence))
28 19
ENST00000269305 TP53-201 1 393
ENST00000445888 TP53-205 1 393
ENST00000615910 TP53-221 5 382

TP53 has 28 transcripts; 19 are protein coding with annotated start and stop codons. support_level is Ensembl's transcript support level, from 1 (best supported by mRNA evidence) to 5, or None when the annotation omits it. Two transcripts can encode the same protein from different UTRs.

Sequences and exons

Get protein and transcript sequences continues with a transcript's protein, coding sequence, UTRs, exons and codon positions. The method overview lists every lookup. Call data.close() when you are finished; a with EnsemblRelease(93, species="human") as data: block closes it automatically.

Examples with other reference data

Fly annotation

This separate example selects fly release 100. Install it before running the Python fragment:

pyensembl install --release 100 --species drosophila_melanogaster
from pyensembl import EnsemblRelease

data = EnsemblRelease(100, species="drosophila_melanogaster")
gene = data.gene_by_id("FBgn0011747")
transcripts = gene.transcripts

Human release 77: HLA-A

This example uses human release 77. Install its data before querying HLA-A:

pyensembl install --release 77 --species human
from pyensembl import EnsemblRelease

data = EnsemblRelease(77, species="human")
print(data.gene_names_at_locus(contig=6, position=29945884))

The result is ["HLA-A"]. Release 77 also uses GRCh38, but its annotation differs from release 93, so pin the release you analyze.