Find genes and transcripts ¶
These examples use the human GRCh38 / Ensembl release 93 data installed on the home page. They run in one Python session.
Find a gene by name or ID¶
A name lookup returns a list of every matching gene:
from pyensembl import EnsemblRelease
data = EnsemblRelease(93, species="human")
for gene in data.genes_by_name("TP53"):
print(gene.id, gene.name, gene.contig, gene.biotype)
ENSG00000141510 TP53 17 protein_coding
A name can match several loci, such as copies on patch or haplotype contigs.
When you know the gene ID, data.gene_by_id("ENSG00000141510") selects it
directly. Alias lookup adds names from a separate source, such as
HGNC.
Versioned IDs ¶
Ensembl IDs can carry a version, which changes when a feature's structure or sequence changes. Every method that takes a gene, transcript, exon or protein ID accepts it with or without one:
gene = data.gene_by_id("ENSG00000141510.16")
print(gene.id, gene.versioned_id)
ENSG00000141510 ENSG00000141510.16
An ID without a version matches whatever version this release has. An ID with
a version must match it: data.gene_by_id("ENSG00000141510.15") raises
ValueError: ENSG00000141510.15 is not in this annotation, which has
ENSG00000141510.16, so an ID copied from another release can't silently match
a changed feature. Releases before Ensembl 77 record no versions, so they
accept only IDs without one.
Find genes at a position ¶
Give a position, or add end to search an interval:
print(data.gene_names_at_locus(contig="17", position=7668402))
for gene in data.genes_at_locus(contig="17", position=7660000, end=7690000):
print(gene.name, gene.strand, gene.start, gene.end)
['TP53']
WRAP53 + 7686071 7703502
TP53 - 7661779 7687550
AC087388.1 - 7685260 7686371
Coordinates are one-based and inclusive, and must use the selected assembly.
Overlapping genes on both strands are returned; pass strand="+" or "-" to
restrict the search. transcripts_at_locus and exons_at_locus work the same
way, and nearest_gene finds the closest gene when none overlaps.
Filter by contig or biotype¶
Whole-genome lists accept optional contig, strand and biotype filters:
coding_genes = data.genes(contig="17", biotype="protein_coding")
print(len(coding_genes))
1183
Biotype names come from the annotation and can differ between releases; list
them with {gene.biotype for gene in data.genes()}. gene_ids, gene_names
and transcript_ids return identifiers without building objects, and
data.contigs() lists the chromosome and contig names.
Choose among a gene's transcripts ¶
A gene usually has several transcripts. List order does not identify a preferred isoform, so choose by ID or by attributes that suit your analysis:
gene = data.gene_by_id("ENSG00000141510")
coding = [t for t in gene.transcripts
if t.biotype == "protein_coding" and t.complete]
print(len(gene.transcripts), len(coding))
coding.sort(key=lambda t: len(t.protein_sequence), reverse=True)
for t in coding[:3]:
print(t.id, t.name, t.support_level, len(t.protein_sequence))
28 19
ENST00000269305 TP53-201 1 393
ENST00000445888 TP53-205 1 393
ENST00000615910 TP53-221 5 382
TP53 has 28 transcripts; 19 are protein coding with annotated start and stop
codons. support_level is Ensembl's transcript support level, from 1 (best
supported by mRNA evidence) to 5, or None when the annotation omits it.
Two transcripts can encode the same protein from different UTRs.
Sequences and exons ¶
Get protein and transcript sequences continues with a
transcript's protein, coding sequence, UTRs, exons and codon positions. The
method overview lists every lookup.
Call data.close() when you are finished; a
with EnsemblRelease(93, species="human") as data: block closes it
automatically.
Examples with other reference data¶
Fly annotation ¶
This separate example selects fly release 100. Install it before running the Python fragment:
pyensembl install --release 100 --species drosophila_melanogaster
from pyensembl import EnsemblRelease
data = EnsemblRelease(100, species="drosophila_melanogaster")
gene = data.gene_by_id("FBgn0011747")
transcripts = gene.transcripts
Human release 77: HLA-A¶
This example uses human release 77. Install its data before querying HLA-A:
pyensembl install --release 77 --species human
from pyensembl import EnsemblRelease
data = EnsemblRelease(77, species="human")
print(data.gene_names_at_locus(contig=6, position=29945884))
The result is ["HLA-A"]. Release 77 also uses GRCh38, but its annotation
differs from release 93, so pin the release you analyze.