Genome queries and setup¶
These methods apply to EnsemblRelease, EnsemblAnnotation and custom
Genome objects. Find a method in the tables
below; the complete reference follows them. Coordinates are
one-based and inclusive. Gene, transcript, exon and protein IDs may include a
version, which must match the annotation's.
Find a method¶
Genes¶
| Method | Returns |
|---|---|
gene_by_id |
One Gene |
genes_by_name |
Every Gene with a name, optionally including aliases |
genes_at_locus |
Genes overlapping a position or interval |
nearest_gene |
(distance, Gene) for the closest gene |
genes |
All genes, filtered by contig, strand or biotype |
gene_by_protein_id |
The gene encoding a protein |
merged_gene_intervals |
Non-overlapping gene intervals on a contig |
gene_ids, gene_names, gene_ids_at_locus, gene_names_at_locus,
gene_ids_of_gene_name, gene_name_of_gene_id and similar methods return IDs
or names instead of objects.
Transcripts¶
| Method | Returns |
|---|---|
transcript_by_id |
One Transcript |
transcripts_by_name |
Transcripts with a name such as TP53-201 |
transcripts_at_locus |
Transcripts overlapping a position or interval |
nearest_transcript |
(distance, Transcript) for the closest transcript |
transcripts |
All transcripts, filtered by contig, strand or biotype |
transcript_by_protein_id |
The transcript encoding a protein |
transcript_ids, transcript_ids_of_gene_id, transcript_ids_at_locus and
similar methods return IDs or names. A gene's transcripts are also available as
gene.transcripts.
Exons¶
| Method | Returns |
|---|---|
exon_by_id |
One Exon |
exons_at_locus |
Exons overlapping a position or interval |
exons |
All exons, filtered by contig or strand |
exon_ids_of_transcript_id, exon_ids_of_gene_id and similar methods return
IDs. Use transcript.exons for exons in transcription order.
Proteins and sequences ¶
| Method | Returns |
|---|---|
protein_sequence |
Amino acids for a protein ID |
protein_ids |
Protein IDs, filtered by contig or strand |
transcript_sequence |
Spliced cDNA for a transcript ID |
Transcript objects also provide protein_sequence, coding_sequence,
sequence and UTR sequences; see protein and transcript sequences.
Reference DNA¶
These need reference DNA, configured with
genome_fasta.
| Method | Returns |
|---|---|
sequence |
Genomic DNA for an interval, optionally reverse-complemented |
download_genome_fasta |
Downloads only the DNA |
index_genome_fasta |
Builds the DNA index before the first query |
fasta |
A zero-based pyfaidx reader, or None |
genome_fasta_path |
The uncompressed FASTA path, or None |
Setup and data¶
| Method | Returns |
|---|---|
download, index |
Downloads and indexes files; existing files are kept |
installed |
Whether queries are ready without setup |
inspect_data |
Offline report of source files and indexes |
contigs |
Contig names in the annotation |
close |
Closes open index files; with does this automatically |
See install and manage data for the command-line equivalents.
Genome ¶
Genome(reference_name, annotation_name, annotation_version=None, gtf_path_or_url=None, transcript_fasta_paths_or_urls=None, protein_fasta_paths_or_urls=None, decompress_on_download=False, copy_local_files_to_cache=False, cache_directory_path=None, genome_fasta_path_or_url=None)
Bundles together the genomic annotation and sequence data associated with a particular genomic database source (e.g. a single Ensembl release) and provides a wide variety of helper methods for accessing this data.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
reference_name
|
str
|
Name of genome assembly which annotations in GTF are aligned against (and from which sequence data is drawn) |
required |
annotation_name
|
str
|
Name of annotation source (e.g. "Ensembl) |
required |
annotation_version
|
int or str
|
Version of annotation database (e.g. 75) |
None
|
gtf_path_or_url
|
str
|
Path or URL of GTF file |
None
|
transcript_fasta_paths_or_urls
|
list
|
List of paths or URLs of FASTA files containing transcript sequences |
None
|
protein_fasta_paths_or_urls
|
list
|
List of paths or URLs of FASTA files containing protein sequences |
None
|
decompress_on_download
|
bool
|
If remote file is compressed, decompress the local copy? |
False
|
copy_local_files_to_cache
|
bool
|
If genome data file is local use it directly or copy to cache first? |
False
|
cache_directory_path
|
None
|
Where to place downloaded and cached files for this genome, by default inferred from reference name, annotation name, annotation version, and global cache dir for pyensembl. |
None
|
genome_fasta_path_or_url
|
str or Path
|
Combined reference DNA FASTA. Supports plain or gzip input; indexes and decompressed copies are stored in the cache. |
None
|
copy_local_files_to_cache
instance-attribute
¶
copy_local_files_to_cache = copy_local_files_to_cache
download_cache
instance-attribute
¶
download_cache = DownloadCache(reference_name=self.reference_name, annotation_name=self.annotation_name, annotation_version=self.annotation_version, decompress_on_download=self.decompress_on_download, copy_local_files_to_cache=self.copy_local_files_to_cache, install_string_function=self.install_string, cache_directory_path=cache_directory_path)
genome_fasta_path
property
¶
genome_fasta_path
Existing uncompressed DNA path, or None; never downloads data.
fasta
property
¶
fasta
Lazy pyfaidx reader, or None when reference DNA is unconfigured or not installed.
The reader uses zero-based, half-open slices. Prefer sequence for
PyEnsembl's one-based, inclusive coordinates and for errors explaining
missing DNA. This property never downloads missing remote data, but
the first access may decompress a local gzip file and build its index.
download_genome_fasta ¶
download_genome_fasta(overwrite=False, show_progress=False)
Download only configured reference DNA, leaving annotation files alone.
show_progress displays progress bars for the download and decompression.
index_genome_fasta ¶
index_genome_fasta(overwrite=False, show_progress=False)
Index configured local DNA without downloading other genome data.
sequence ¶
sequence(contig, start, end, mask='upper', *, strand='+')
Return DNA using one-based, inclusive coordinates.
Contigs may be named as in the FASTA or as pyensembl reports them
(e.g. gene.contig). strand="-" returns the reverse complement,
so sequence(t.contig, t.start, t.end, strand=t.strand) reads a
transcript's locus 5' to 3'. Invalid intervals and absent contigs raise
ValueError; unconfigured or uninstalled DNA raises
MissingGenomeFastaError. Use mask='raw' to preserve soft masking.
inspect_data ¶
inspect_data(check_genome_fasta=False)
Offline inventory of configured sources and indexes.
Return readiness and role-keyed datacache FileInspection objects.
Never downloads, imports, builds indexes, creates directories or loads
pickle contents. Availability checks regular, readable, nonempty files;
SQLite completeness and reference-DNA fingerprints are also checked.
verified is not inferred from an acquisition provenance receipt.
installed ¶
installed()
Whether every configured file is downloaded and indexed, so queries need no network access or setup.
Includes reference DNA when it is configured. Only reads the cache: never downloads, copies, indexes, or creates files.
download ¶
download(overwrite=False, show_progress=False)
Download data files needed by this Genome instance.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
overwrite
|
bool
|
Download files regardless whether local copy already exists. |
False
|
show_progress
|
bool
|
Display download progress bars. |
False
|
index ¶
index(overwrite=False, show_progress=False)
Assuming that all necessary data for this Genome has been downloaded, generate the GTF database and save efficient representation of FASTA sequence files.
show_progress displays progress bars while the GTF is read, the database is filled, and sequence files are read.
install_string ¶
install_string()
Add every missing file to the install string shown to the user in an error message.
genome_fasta_install_string ¶
genome_fasta_install_string()
Command that installs only this genome's reference DNA.
delete_index_files ¶
delete_index_files()
Delete SQLite and FASTA indexes, preserving source files.
Missing sources and indexes are allowed; nothing is downloaded or
copied. Return (path, size_in_bytes) pairs for the removed files.
transcript_sequence ¶
transcript_sequence(transcript_id)
Return cDNA nucleotide sequence of transcript, or None if
transcript doesn't have cDNA sequence. A versioned ID such as
ENST00000269305.8 must match this annotation's version.
protein_sequence ¶
protein_sequence(protein_id)
Return amino-acid sequence of protein, or None if the protein is
absent from the FASTA. A versioned ID such as ENSP00000269305.4
must match this annotation's version.
genes_at_locus ¶
genes_at_locus(contig, position, end=None, strand=None)
Gene objects overlapping position, or the position..end
interval, on contig. Coordinates are one-based and inclusive.
Pass strand to restrict the search to one strand.
transcripts_at_locus ¶
transcripts_at_locus(contig, position, end=None, strand=None)
Transcript objects overlapping position, or the position..end
interval, on contig. Coordinates are one-based and inclusive.
Pass strand to restrict the search to one strand.
exons_at_locus ¶
exons_at_locus(contig, position, end=None, strand=None)
Exon objects overlapping position, or the position..end
interval, on contig. Coordinates are one-based and inclusive.
Pass strand to restrict the search to one strand.
nearest_gene ¶
nearest_gene(contig, position, end=None, strand=None)
Find the gene on contig whose locus is nearest to the position
(or position..end interval), even if no gene actually overlaps.
Returns (distance, Gene) where distance is the number of
intervening bases (0 if the query interval falls inside the gene).
Returns (inf, None) when the contig has no genes on the
requested strand.
Pass strand to restrict the search to one strand of the contig.
merged_gene_intervals ¶
merged_gene_intervals(contig, strand=None)
Return the union of all gene loci on contig as a sorted list of
non-overlapping (start, end) tuples. Adjacent intervals
(end + 1 == next start) are merged into one. Pass strand to
restrict to one strand of the contig.
Useful for asking "is this position inside any gene?" or computing gene-density coverage without double-counting overlapping genes.
nearest_transcript ¶
nearest_transcript(contig, position, end=None, strand=None)
Find the transcript on contig whose locus is nearest to the
position (or position..end interval), even if no transcript
overlaps. See :meth:nearest_gene for the return shape.
gene_names_at_locus ¶
gene_names_at_locus(contig, position, end=None, strand=None)
Names of genes overlapping position, or the position..end
interval, on contig. Coordinates are one-based and inclusive.
Pass strand to restrict the search to one strand.
locus_of_gene_id ¶
locus_of_gene_id(gene_id)
Given a gene ID returns Locus with: chromosome, start, stop, strand
loci_of_gene_names ¶
loci_of_gene_names(gene_name)
Given a gene name returns list of Locus objects with fields: chromosome, start, stop, strand You can get multiple results since a gene might have multiple copies in the genome.
contigs ¶
contigs()
Returns all contig names for any gene in the genome (field called "seqname" in Ensembl GTF files)
genes ¶
genes(contig=None, strand=None, biotype=None)
Returns all Gene objects in the database. Can be restricted to a particular contig/chromosome, strand, or biotype.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
contig
|
str
|
Only return genes on the given contig. |
None
|
strand
|
str
|
Only return genes on this strand. |
None
|
biotype
|
str
|
Only return genes with this Ensembl |
None
|
gene_by_id ¶
gene_by_id(gene_id)
Construct a Gene object for the given gene ID, such as
"ENSG00000141510". A versioned ID such as
"ENSG00000141510.16" must match this annotation's version.
genes_by_name ¶
genes_by_name(gene_name, aliases=None)
Get all the unique genes with the given name (there might be multiple due to copies in the genome), return a list containing a Gene object for each distinct ID.
aliases optionally supplies a name-to-gene-ID mapping, for example GeneNameAliases.from_hgnc(path). Include all exact and alias matches present in this annotation. No alias data is downloaded by a query.
gene_by_protein_id ¶
gene_by_protein_id(protein_id)
Get the gene ID associated with the given protein ID, return its Gene object
gene_names ¶
gene_names(contig=None, strand=None)
Return all genes in the database, optionally restrict to a chromosome and/or strand.
gene_ids ¶
gene_ids(contig=None, strand=None, biotype=None)
What are all the gene IDs (optionally restrict to a given
chromosome/contig, strand, or gene_biotype).
gene_ids_of_gene_name ¶
gene_ids_of_gene_name(gene_name, aliases=None)
What are the gene IDs associated with a given gene name? (due to copy events, there might be multiple genes per name)
aliases is an optional mapping from names to one or more stable IDs. Unknown IDs are ignored; ambiguous names retain every matching gene.
gene_id_of_protein_id ¶
gene_id_of_protein_id(protein_id)
What is the gene ID associated with a given protein ID?
transcripts ¶
transcripts(contig=None, strand=None, biotype=None)
Construct Transcript object for every transcript entry in
the database. Optionally restrict to a particular contig, strand,
or transcript_biotype (e.g. "protein_coding").
transcript_by_id ¶
transcript_by_id(transcript_id)
Construct a Transcript object for the given transcript ID, such as
"ENST00000269305". A versioned ID such as
"ENST00000269305.8" must match this annotation's version.
transcripts_by_name ¶
transcripts_by_name(transcript_name)
Transcript objects with the given name, such as "TP53-201".
transcript_by_protein_id ¶
transcript_by_protein_id(protein_id)
Transcript object that encodes the given protein ID.
transcript_names ¶
transcript_names(contig=None, strand=None)
What are all the transcript names in the database (optionally, restrict to a given chromosome and/or strand)
transcript_id_of_protein_id ¶
transcript_id_of_protein_id(protein_id)
What is the transcript ID associated with a given protein ID?
exons ¶
exons(contig=None, strand=None)
Create exon object for all exons in the database, optionally
restrict to a particular chromosome using the contig argument.
exon_by_id ¶
exon_by_id(exon_id)
Construct an Exon object for the given exon ID. A versioned ID must match this annotation's version.
protein_ids ¶
protein_ids(contig=None, strand=None)
What are all the protein IDs (optionally restrict to a given chromosome and/or strand)