Use dated Ensembl annotations¶
Ensembl's new platform publishes annotations by
genome assembly and date instead of numbered releases. PyEnsembl treats each
annotation date as a release: EnsemblRelease("2026_04", species="human")
works like EnsemblRelease(116, species="human") and selects GRCh38.
Choose a reference compares both with custom
files. Installed datasets are not affected by the transition.
Install a dated release¶
This example downloads the GTF, cDNA and peptide files of the human GRCh38 annotation dated 2026_04 from its dataset directory.
from pyensembl import EnsemblRelease
data = EnsemblRelease("2026_04", species="human")
data.download()
data.index()
gene = data.gene_by_id("ENSG00000141510")
print(gene.name, gene.contig, gene.start, gene.end, gene.strand)
From the command line:
pyensembl install --release 2026_04 --species human
A date selects the annotation of the species' current assembly, the assembly
of its last numbered release, under PyEnsembl's name for that assembly. Dated
and numbered releases of an assembly share its cache directory, for example
GRCh38/ensembl2026_04/ beside GRCh38/ensembl116/.
For human GRCh38, 2026_04 has the same gene and transcript IDs and versions
as release 116; 2025_12 matches release 115 and 2023_03 matches release 110.
2024_11 is a GENCODE subset without non-canonical lncRNA transcripts; release
114 has the complete annotation.
Human GRCh38 uses genes-including_alt.gtf.gz, the counterpart of the
complete patch and haplotype GTF of numbered GRCh38 releases. Other species
use genes.gtf.gz; dated zebrafish GRCz11 therefore omits the alternative
loci of its numbered releases. The cDNA FASTA covers every transcript biotype.
For reference DNA, add genome_fasta=True; genome_fasta_mask="none",
"soft" or "hard" selects the corresponding combined genome FASTA. Only
toplevel DNA is published. Read genomic DNA
describes coordinates, strand and masking.
Find a date ¶
The new platform's FTP layout
groups data by GCA/GCF assembly accession, provider and annotation date. Browse
an assembly's directory for its dates, for example
human GRCh38.
An annotation date records when the annotation was built, not when it was
published: Arabidopsis's current annotation is dated 2010_09, and fly's
2022_07 annotation is older than release 116's. A website release label such
as 2026-07 is different from an annotation date and is rejected.
Other assemblies and providers¶
EnsemblAnnotation selects any dataset by assembly accession, provider and
date: an older assembly such as GRCh37 (2013_09), a species PyEnsembl does
not list, or another provider's annotation.
from pyensembl import EnsemblAnnotation
with EnsemblAnnotation(
assembly_accession="GCA_000001405.14",
annotation_date="2013_09",
provider="ensembl",
species="human",
reference_name="GRCh37",
) as data:
data.download()
data.index()
include_alt=True selects genes-including_alt.gtf.gz; the default selects
genes.gtf.gz. Alternative loci can add name matches. Confirm that the chosen
coverage file exists for your dataset, and use stable gene IDs when a name is
ambiguous. Accession, provider, date and coverage have separate default caches.
If overriding cache_directory_path, use a distinct directory for each dataset.
The optional species label guards species-specific alias sources; it does not
select or verify the assembly's species. Keep it consistent with the accession.
genome_fasta=True and genome_fasta_mask select reference DNA as above.
Community providers may have different available files. Check the directory
before installation. Use Genome with explicitly matched files when a dataset
does not provide the standard GTF, cDNA and peptide filenames. GFF3 and other
formats require conversion to GTF first. PyEnsembl does not select a rolling
latest annotation or integrate the new GraphQL/refget services.
Transition from numbered releases¶
As checked on 2026-10-05, Ensembl's transition announcement states that legacy FTP and API services remain available but stop receiving updates after Ensembl 116; new datasets are delivered through the new platform. PyEnsembl's numbered-release URLs and selection rules are preserved. Old species-name directories on the new platform were scheduled for retirement in August 2026.