Dataset constructors¶
EnsemblRelease selects a numbered Ensembl
release, such as 116, or a dated release of the new platform,
such as "2026_04". EnsemblAnnotation selects any assembly accession,
provider and dated geneset from the new platform. Both provide the shared
Genome methods.
EnsemblRelease ¶
EnsemblRelease(release=MAX_ENSEMBL_RELEASE, species=human, server=ENSEMBL_FTP_SERVER, *, genome_fasta=None, genome_fasta_type='toplevel', genome_fasta_mask='none', download_genome_fasta=None, genome_fasta_path=None)
Bundles together the genomic annotation and sequence data associated with a particular release of the Ensembl database.
A release is either numbered (up to 116, the last) or, on the new Ensembl
platform, the annotation date of the species' current assembly, e.g.
EnsemblRelease("2026_04", species="human") for GRCh38.
release : int or str Numbered Ensembl release, e.g. 116, or an annotation date on the new Ensembl platform in YYYY_MM form, e.g. "2026_04".
genome_fasta : True or path, optional
Reference DNA for sequence(): True for Ensembl's DNA for this
release (see genome_fasta_type and genome_fasta_mask), or a local
plain or gzip FASTA. Nothing is downloaded until
download_genome_fasta(), download(), or pyensembl install.
genome_fasta_urls
instance-attribute
¶
genome_fasta_urls = [genome_fasta_url] if genome_fasta is True else []
normalize_init_values
classmethod
¶
normalize_init_values(release, species, server)
Normalizes the arguments which uniquely specify an EnsemblRelease genome.
cached
classmethod
¶
cached(release=MAX_ENSEMBL_RELEASE, species=human, server=ENSEMBL_FTP_SERVER, *, genome_fasta=None, genome_fasta_type='toplevel', genome_fasta_mask='none', download_genome_fasta=None, genome_fasta_path=None)
Construct EnsemblRelease if it's never been made before, otherwise return an old instance.
from_dict
classmethod
¶
from_dict(state_dict)
Deserialize EnsemblRelease without creating duplicate instances.
EnsemblAnnotation ¶
EnsemblAnnotation(assembly_accession, annotation_date, *, provider='ensembl', include_alt=False, genome_fasta=False, genome_fasta_mask='none', species=None, reference_name=None, cache_directory_path=None, server=ENSEMBL_PLATFORM_FTP_SERVER)
A dated geneset for one assembly on the new Ensembl platform.
Select an existing assembly accession, provider and annotation date from
Ensembl's downloads page. These are not numbered EnsemblRelease values
or the website's YYYY-MM release label. Construction never downloads data;
call download() then index() to install the selected files.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
assembly_accession
|
str
|
Versioned INSDC accession, e.g. GCA_000001405.29. |
required |
annotation_date
|
str
|
Annotation directory in YYYY_MM form, e.g. 2023_03. |
required |
provider
|
str
|
Existing provider directory, usually "ensembl" or "community". |
'ensembl'
|
include_alt
|
bool
|
Select genes-including_alt.gtf.gz instead of genes.gtf.gz. Check that this file exists for the selected dataset; coverage is never inferred. |
False
|
genome_fasta
|
bool
|
Also install combined reference DNA. False keeps DNA optional. |
False
|
genome_fasta_mask
|
str
|
"none", "soft" or "hard", selecting the corresponding genome FASTA. |
'none'
|
species
|
str
|
Known species name, used to guard species-specific alias lookups. The assembly accession remains the authoritative dataset selection. |
None
|
reference_name
|
str
|
Display name for the assembly; defaults to the accession. |
None
|
cache_directory_path
|
str
|
Explicit cache directory. Use a distinct directory per dataset. |
None
|
server
|
str
|
Root of a mirror using the same accession/provider/date layout. |
ENSEMBL_PLATFORM_FTP_SERVER
|
species
instance-attribute
¶
species = find_species_by_name(species) if species is not None else None