Skip to content

Genome queries and setup

These methods apply to EnsemblRelease, EnsemblAnnotation and custom Genome objects. Find a method in the tables below; the complete reference follows them. Coordinates are one-based and inclusive. Gene, transcript, exon and protein IDs may include a version, which must match the annotation's.

Find a method

Genes

Method Returns
gene_by_id One Gene
genes_by_name Every Gene with a name, optionally including aliases
genes_at_locus Genes overlapping a position or interval
nearest_gene (distance, Gene) for the closest gene
genes All genes, filtered by contig, strand or biotype
gene_by_protein_id The gene encoding a protein
merged_gene_intervals Non-overlapping gene intervals on a contig

gene_ids, gene_names, gene_ids_at_locus, gene_names_at_locus, gene_ids_of_gene_name, gene_name_of_gene_id and similar methods return IDs or names instead of objects.

Transcripts

Method Returns
transcript_by_id One Transcript
transcripts_by_name Transcripts with a name such as TP53-201
transcripts_at_locus Transcripts overlapping a position or interval
nearest_transcript (distance, Transcript) for the closest transcript
transcripts All transcripts, filtered by contig, strand or biotype
transcript_by_protein_id The transcript encoding a protein

transcript_ids, transcript_ids_of_gene_id, transcript_ids_at_locus and similar methods return IDs or names. A gene's transcripts are also available as gene.transcripts.

Exons

Method Returns
exon_by_id One Exon
exons_at_locus Exons overlapping a position or interval
exons All exons, filtered by contig or strand

exon_ids_of_transcript_id, exon_ids_of_gene_id and similar methods return IDs. Use transcript.exons for exons in transcription order.

Proteins and sequences

Method Returns
protein_sequence Amino acids for a protein ID
protein_ids Protein IDs, filtered by contig or strand
transcript_sequence Spliced cDNA for a transcript ID

Transcript objects also provide protein_sequence, coding_sequence, sequence and UTR sequences; see protein and transcript sequences.

Reference DNA

These need reference DNA, configured with genome_fasta.

Method Returns
sequence Genomic DNA for an interval, optionally reverse-complemented
download_genome_fasta Downloads only the DNA
index_genome_fasta Builds the DNA index before the first query
fasta A zero-based pyfaidx reader, or None
genome_fasta_path The uncompressed FASTA path, or None

Setup and data

Method Returns
download, index Downloads and indexes files; existing files are kept
installed Whether queries are ready without setup
inspect_data Offline report of source files and indexes
contigs Contig names in the annotation
close Closes open index files; with does this automatically

See install and manage data for the command-line equivalents.

Genome

Genome(reference_name, annotation_name, annotation_version=None, gtf_path_or_url=None, transcript_fasta_paths_or_urls=None, protein_fasta_paths_or_urls=None, decompress_on_download=False, copy_local_files_to_cache=False, cache_directory_path=None, genome_fasta_path_or_url=None)

Bundles together the genomic annotation and sequence data associated with a particular genomic database source (e.g. a single Ensembl release) and provides a wide variety of helper methods for accessing this data.

Parameters:

Name Type Description Default
reference_name str

Name of genome assembly which annotations in GTF are aligned against (and from which sequence data is drawn)

required
annotation_name str

Name of annotation source (e.g. "Ensembl)

required
annotation_version int or str

Version of annotation database (e.g. 75)

None
gtf_path_or_url str

Path or URL of GTF file

None
transcript_fasta_paths_or_urls list

List of paths or URLs of FASTA files containing transcript sequences

None
protein_fasta_paths_or_urls list

List of paths or URLs of FASTA files containing protein sequences

None
decompress_on_download bool

If remote file is compressed, decompress the local copy?

False
copy_local_files_to_cache bool

If genome data file is local use it directly or copy to cache first?

False
cache_directory_path None

Where to place downloaded and cached files for this genome, by default inferred from reference name, annotation name, annotation version, and global cache dir for pyensembl.

None
genome_fasta_path_or_url str or Path

Combined reference DNA FASTA. Supports plain or gzip input; indexes and decompressed copies are stored in the cache.

None

reference_name instance-attribute

reference_name = reference_name

annotation_name instance-attribute

annotation_name = annotation_name

annotation_version instance-attribute

annotation_version = annotation_version

decompress_on_download instance-attribute

decompress_on_download = decompress_on_download

copy_local_files_to_cache instance-attribute

copy_local_files_to_cache = copy_local_files_to_cache

cache_directory_path instance-attribute

cache_directory_path = cache_directory_path

download_cache instance-attribute

download_cache = DownloadCache(reference_name=self.reference_name, annotation_name=self.annotation_name, annotation_version=self.annotation_version, decompress_on_download=self.decompress_on_download, copy_local_files_to_cache=self.copy_local_files_to_cache, install_string_function=self.install_string, cache_directory_path=cache_directory_path)

requires_gtf property

requires_gtf

requires_genome_fasta property

requires_genome_fasta

genome_fasta_path property

genome_fasta_path

Existing uncompressed DNA path, or None; never downloads data.

fasta property

fasta

Lazy pyfaidx reader, or None when reference DNA is unconfigured or not installed.

The reader uses zero-based, half-open slices. Prefer sequence for PyEnsembl's one-based, inclusive coordinates and for errors explaining missing DNA. This property never downloads missing remote data, but the first access may decompress a local gzip file and build its index.

requires_transcript_fasta property

requires_transcript_fasta

requires_protein_fasta property

requires_protein_fasta

db property

db

protein_sequences property

protein_sequences

transcript_sequences property

transcript_sequences

download_genome_fasta

download_genome_fasta(overwrite=False, show_progress=False)

Download only configured reference DNA, leaving annotation files alone.

show_progress displays progress bars for the download and decompression.

index_genome_fasta

index_genome_fasta(overwrite=False, show_progress=False)

Index configured local DNA without downloading other genome data.

sequence

sequence(contig, start, end, mask='upper', *, strand='+')

Return DNA using one-based, inclusive coordinates.

Contigs may be named as in the FASTA or as pyensembl reports them (e.g. gene.contig). strand="-" returns the reverse complement, so sequence(t.contig, t.start, t.end, strand=t.strand) reads a transcript's locus 5' to 3'. Invalid intervals and absent contigs raise ValueError; unconfigured or uninstalled DNA raises MissingGenomeFastaError. Use mask='raw' to preserve soft masking.

to_dict

to_dict()

Returns a dictionary of the essential fields of this Genome.

required_local_files

required_local_files()

required_local_files_exist

required_local_files_exist(empty_files_ok=False)

inspect_data

inspect_data(check_genome_fasta=False)

Offline inventory of configured sources and indexes.

Return readiness and role-keyed datacache FileInspection objects. Never downloads, imports, builds indexes, creates directories or loads pickle contents. Availability checks regular, readable, nonempty files; SQLite completeness and reference-DNA fingerprints are also checked. verified is not inferred from an acquisition provenance receipt.

installed

installed()

Whether every configured file is downloaded and indexed, so queries need no network access or setup.

Includes reference DNA when it is configured. Only reads the cache: never downloads, copies, indexes, or creates files.

download

download(overwrite=False, show_progress=False)

Download data files needed by this Genome instance.

Parameters:

Name Type Description Default
overwrite bool

Download files regardless whether local copy already exists.

False
show_progress bool

Display download progress bars.

False

index

index(overwrite=False, show_progress=False)

Assuming that all necessary data for this Genome has been downloaded, generate the GTF database and save efficient representation of FASTA sequence files.

show_progress displays progress bars while the GTF is read, the database is filled, and sequence files are read.

install_string

install_string()

Add every missing file to the install string shown to the user in an error message.

genome_fasta_install_string

genome_fasta_install_string()

Command that installs only this genome's reference DNA.

clear_cache

clear_cache()

Clear values cached in memory without deleting index files.

close

close()

Close resources opened by this genome.

delete_index_files

delete_index_files()

Delete SQLite and FASTA indexes, preserving source files.

Missing sources and indexes are allowed; nothing is downloaded or copied. Return (path, size_in_bytes) pairs for the removed files.

transcript_sequence

transcript_sequence(transcript_id)

Return cDNA nucleotide sequence of transcript, or None if transcript doesn't have cDNA sequence. A versioned ID such as ENST00000269305.8 must match this annotation's version.

protein_sequence

protein_sequence(protein_id)

Return amino-acid sequence of protein, or None if the protein is absent from the FASTA. A versioned ID such as ENSP00000269305.4 must match this annotation's version.

genes_at_locus

genes_at_locus(contig, position, end=None, strand=None)

Gene objects overlapping position, or the position..end interval, on contig. Coordinates are one-based and inclusive. Pass strand to restrict the search to one strand.

transcripts_at_locus

transcripts_at_locus(contig, position, end=None, strand=None)

Transcript objects overlapping position, or the position..end interval, on contig. Coordinates are one-based and inclusive. Pass strand to restrict the search to one strand.

exons_at_locus

exons_at_locus(contig, position, end=None, strand=None)

Exon objects overlapping position, or the position..end interval, on contig. Coordinates are one-based and inclusive. Pass strand to restrict the search to one strand.

nearest_gene

nearest_gene(contig, position, end=None, strand=None)

Find the gene on contig whose locus is nearest to the position (or position..end interval), even if no gene actually overlaps.

Returns (distance, Gene) where distance is the number of intervening bases (0 if the query interval falls inside the gene). Returns (inf, None) when the contig has no genes on the requested strand.

Pass strand to restrict the search to one strand of the contig.

merged_gene_intervals

merged_gene_intervals(contig, strand=None)

Return the union of all gene loci on contig as a sorted list of non-overlapping (start, end) tuples. Adjacent intervals (end + 1 == next start) are merged into one. Pass strand to restrict to one strand of the contig.

Useful for asking "is this position inside any gene?" or computing gene-density coverage without double-counting overlapping genes.

nearest_transcript

nearest_transcript(contig, position, end=None, strand=None)

Find the transcript on contig whose locus is nearest to the position (or position..end interval), even if no transcript overlaps. See :meth:nearest_gene for the return shape.

gene_ids_at_locus

gene_ids_at_locus(contig, position, end=None, strand=None)

gene_names_at_locus

gene_names_at_locus(contig, position, end=None, strand=None)

Names of genes overlapping position, or the position..end interval, on contig. Coordinates are one-based and inclusive. Pass strand to restrict the search to one strand.

exon_ids_at_locus

exon_ids_at_locus(contig, position, end=None, strand=None)

transcript_ids_at_locus

transcript_ids_at_locus(contig, position, end=None, strand=None)

transcript_names_at_locus

transcript_names_at_locus(contig, position, end=None, strand=None)

protein_ids_at_locus

protein_ids_at_locus(contig, position, end=None, strand=None)

locus_of_gene_id

locus_of_gene_id(gene_id)

Given a gene ID returns Locus with: chromosome, start, stop, strand

loci_of_gene_names

loci_of_gene_names(gene_name)

Given a gene name returns list of Locus objects with fields: chromosome, start, stop, strand You can get multiple results since a gene might have multiple copies in the genome.

locus_of_transcript_id

locus_of_transcript_id(transcript_id)

locus_of_exon_id

locus_of_exon_id(exon_id)

Given an exon ID returns Locus

contigs

contigs()

Returns all contig names for any gene in the genome (field called "seqname" in Ensembl GTF files)

genes

genes(contig=None, strand=None, biotype=None)

Returns all Gene objects in the database. Can be restricted to a particular contig/chromosome, strand, or biotype.

Parameters:

Name Type Description Default
contig str

Only return genes on the given contig.

None
strand str

Only return genes on this strand.

None
biotype str

Only return genes with this Ensembl gene_biotype (e.g. "protein_coding").

None

gene_by_id

gene_by_id(gene_id)

Construct a Gene object for the given gene ID, such as "ENSG00000141510". A versioned ID such as "ENSG00000141510.16" must match this annotation's version.

genes_by_name

genes_by_name(gene_name, aliases=None)

Get all the unique genes with the given name (there might be multiple due to copies in the genome), return a list containing a Gene object for each distinct ID.

aliases optionally supplies a name-to-gene-ID mapping, for example GeneNameAliases.from_hgnc(path). Include all exact and alias matches present in this annotation. No alias data is downloaded by a query.

gene_by_protein_id

gene_by_protein_id(protein_id)

Get the gene ID associated with the given protein ID, return its Gene object

gene_names

gene_names(contig=None, strand=None)

Return all genes in the database, optionally restrict to a chromosome and/or strand.

gene_name_of_gene_id

gene_name_of_gene_id(gene_id)

gene_name_of_transcript_id

gene_name_of_transcript_id(transcript_id)

gene_name_of_transcript_name

gene_name_of_transcript_name(transcript_name)

gene_name_of_exon_id

gene_name_of_exon_id(exon_id)

gene_ids

gene_ids(contig=None, strand=None, biotype=None)

What are all the gene IDs (optionally restrict to a given chromosome/contig, strand, or gene_biotype).

gene_ids_of_gene_name

gene_ids_of_gene_name(gene_name, aliases=None)

What are the gene IDs associated with a given gene name? (due to copy events, there might be multiple genes per name)

aliases is an optional mapping from names to one or more stable IDs. Unknown IDs are ignored; ambiguous names retain every matching gene.

gene_id_of_protein_id

gene_id_of_protein_id(protein_id)

What is the gene ID associated with a given protein ID?

transcripts

transcripts(contig=None, strand=None, biotype=None)

Construct Transcript object for every transcript entry in the database. Optionally restrict to a particular contig, strand, or transcript_biotype (e.g. "protein_coding").

transcript_by_id

transcript_by_id(transcript_id)

Construct a Transcript object for the given transcript ID, such as "ENST00000269305". A versioned ID such as "ENST00000269305.8" must match this annotation's version.

transcripts_by_name

transcripts_by_name(transcript_name)

Transcript objects with the given name, such as "TP53-201".

transcript_by_protein_id

transcript_by_protein_id(protein_id)

Transcript object that encodes the given protein ID.

transcript_names

transcript_names(contig=None, strand=None)

What are all the transcript names in the database (optionally, restrict to a given chromosome and/or strand)

transcript_names_of_gene_name

transcript_names_of_gene_name(gene_name)

transcript_name_of_transcript_id

transcript_name_of_transcript_id(transcript_id)

transcript_ids

transcript_ids(contig=None, strand=None, biotype=None)

transcript_ids_of_gene_id

transcript_ids_of_gene_id(gene_id)

transcript_ids_of_gene_name

transcript_ids_of_gene_name(gene_name)

transcript_ids_of_transcript_name

transcript_ids_of_transcript_name(transcript_name)

transcript_ids_of_exon_id

transcript_ids_of_exon_id(exon_id)

transcript_id_of_protein_id

transcript_id_of_protein_id(protein_id)

What is the transcript ID associated with a given protein ID?

exons

exons(contig=None, strand=None)

Create exon object for all exons in the database, optionally restrict to a particular chromosome using the contig argument.

exon_by_id

exon_by_id(exon_id)

Construct an Exon object for the given exon ID. A versioned ID must match this annotation's version.

exon_ids

exon_ids(contig=None, strand=None)

exon_ids_of_gene_id

exon_ids_of_gene_id(gene_id)

exon_ids_of_gene_name

exon_ids_of_gene_name(gene_name)

exon_ids_of_transcript_name

exon_ids_of_transcript_name(transcript_name)

exon_ids_of_transcript_id

exon_ids_of_transcript_id(transcript_id)

protein_ids

protein_ids(contig=None, strand=None)

What are all the protein IDs (optionally restrict to a given chromosome and/or strand)