Data and caches¶
Use data inspection to check availability without downloading or indexing. Advisory provenance is not a trusted expected checksum and does not validate the biological contents of a file.
DownloadCache ¶
DownloadCache(reference_name, annotation_name, annotation_version=None, decompress_on_download=False, copy_local_files_to_cache=False, install_string_function=None, cache_directory_path=None)
Downloads remote files to cache, optionally copies local files into cache, raises custom message if data is missing.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
reference_name
|
str
|
Name of reference genome |
required |
annotation_name
|
str
|
Name of annotation database |
required |
annotation_version
|
str or int
|
Version or release of annotation database |
None
|
decompress_on_download
|
bool
|
If downloading a .fa.gz file, should we automatically expand it into a decompressed FASTA file? |
False
|
copy_local_files_to_cache
|
bool
|
If file is on the local file system, should we still copy it into the cache? |
False
|
install_string_function
|
fn
|
Function which returns an error message with install instructions. If not provided then the error tells the user what data is missing without install instructions. |
None
|
cache_directory_path
|
str
|
Where to place downloaded and temporary files, by default inferred from reference name, annotation name, annotation version, and the global cache directory determined by datacache. |
None
|
cache_subdirectory
instance-attribute
¶
cache_subdirectory = cache_subdirectory(reference_name=reference_name, annotation_name=annotation_name, annotation_version=annotation_version)
copy_local_files_to_cache
instance-attribute
¶
copy_local_files_to_cache = copy_local_files_to_cache
is_url_format ¶
is_url_format(path_or_url)
Is the given string a URL?
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
path_or_url
|
str
|
|
required |
Returns:
| Type | Description |
|---|---|
bool
|
|
cached_path ¶
cached_path(path_or_url)
When downloading remote files, the default behavior is to name local files the same as their remote counterparts.
local_path ¶
local_path(path_or_url)
Expected source location, without acquisition or directory creation.
inspect ¶
inspect(path_or_url, empty_files_ok=False)
Inspect a configured source without acquiring it.
download_or_copy_if_necessary ¶
download_or_copy_if_necessary(path_or_url, download_if_missing=False, overwrite=False, show_progress=False)
Download a remote file or copy Get the local path to a possibly remote file.
Download if file is missing from the cache directory and
download_if_missing is True. Download even if local file exists if
both download_if_missing and overwrite are True.
If the file is on the local file system then return its path, unless self.copy_local_to_cache is True, and then copy it to the cache first.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
path_or_url
|
str
|
|
required |
download_if_missing
|
bool
|
Download files if missing from local cache |
False
|
overwrite
|
bool
|
Overwrite existing copy if it exists |
False
|
local_path_or_install_error ¶
local_path_or_install_error(field_name, path_or_url, download_if_missing=False, overwrite=False, show_progress=False)
Database ¶
Database(gtf_path, install_string=None, cache_directory_path=None, restrict_gtf_columns=None, restrict_gtf_features=None)
Wrapper around sqlite3 database so that the rest of the library doesn't have to worry about constructing the .db file or writing SQL queries directly.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
gtf_path
|
str
|
Path to GTF annotation file |
required |
install_string
|
str
|
Message to tell user if database connection is requested before database is created. |
None
|
cache_directory_path
|
str
|
Path to directory where database should be written. If omitted then use path of GTF file. |
None
|
restrict_gtf_columns
|
list/set of str or None
|
If provided then extract only these columns before creating a database. |
None
|
restrict_gtf_features
|
list/set of str or None
|
If provided then only create tables for these features. |
None
|
PRIMARY_KEY_COLUMNS
class-attribute
instance-attribute
¶
PRIMARY_KEY_COLUMNS = {'gene': 'gene_id', 'transcript': 'transcript_id'}
ID_VERSION_COLUMNS
class-attribute
instance-attribute
¶
ID_VERSION_COLUMNS = {'gene_id': ('gene', 'gene_version'), 'transcript_id': ('transcript', 'transcript_version'), 'exon_id': ('exon', 'exon_version'), 'protein_id': ('CDS', 'protein_version')}
create ¶
create(overwrite=False, show_progress=False)
Create the local database (including indexing) if it's not
already set up. If overwrite is True, always re-create
the database from scratch. show_progress displays row insertion
progress.
Returns a connection to the database.
connect_or_create ¶
connect_or_create(overwrite=False, show_progress=False)
Return a connection to the database if it exists, otherwise create it.
With overwrite, rebuild it from the GTF even if it exists.
column_values_at_locus ¶
column_values_at_locus(column_name, feature, contig, position, end=None, strand=None, distinct=False, sorted=False)
Get the non-null values of a column from the database at a particular range of loci
distinct_column_values_at_locus ¶
distinct_column_values_at_locus(column, feature, contig, position, end=None, strand=None)
Gather all the distinct values for a property/column at some specified locus.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
column
|
str
|
Which property are we getting the values of. |
required |
feature
|
str
|
Which type of entry (e.g. transcript, exon, gene) is the property associated with? |
required |
contig
|
str
|
Chromosome or unplaced contig name |
required |
position
|
int
|
Chromosomal position |
required |
end
|
int
|
End position of a range, if unspecified assume we're only looking at the single given position. |
None
|
strand
|
str
|
Either the positive ('+') or negative strand ('-'). If unspecified then check for values on either strand. |
None
|
run_sql_query ¶
run_sql_query(sql, required=False, query_params=[])
Given an arbitrary SQL query, run it against the database and return the results.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
sql
|
str
|
SQL query |
required |
required
|
bool
|
Raise an error if no results found in the database |
False
|
query_params
|
list
|
For each '?' in the query there must be a corresponding value in this list. |
[]
|
query ¶
query(select_column_names, filter_column, filter_value, feature, distinct=False, required=False)
Construct a SQL query and run against the sqlite3 database, filtered both by the feature type and a user-provided column/value.
stored_id ¶
stored_id(id_column, identifier)
Form in which this annotation stores a gene, transcript, exon or protein ID, or the ID unchanged if it is absent or stored as given.
A bare Ensembl ID matches whatever version is installed. A versioned
ID must match the installed version: ValueError if the annotation
records a different version or none. Other IDs, such as TAIR
AT1G01010.1, are returned unchanged.
installed_versions ¶
installed_versions(id_column, identifier)
Stored forms of an Ensembl gene, transcript, exon or protein ID, with
or without a version, each mapped to its recorded version or None.
query_one ¶
query_one(select_column_names, filter_column, filter_value, feature, distinct=False, required=False)
query_feature_values ¶
query_feature_values(column, feature, distinct=True, contig=None, strand=None, biotype=None)
Run a SQL query against the sqlite3 database, filtered only on the feature type.
query_loci ¶
query_loci(filter_column, filter_value, feature)
Query for loci satisfying a given filter and feature type.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
filter_column
|
str
|
Name of column to filter results by. |
required |
filter_value
|
str
|
Only return loci which have this value in the their filter_column. |
required |
feature
|
str
|
Feature names such as 'transcript', 'gene', and 'exon' |
required |
Returns:
| Type | Description |
|---|---|
list of Locus
|
Matching genomic loci. |
query_locus ¶
query_locus(filter_column, filter_value, feature)
Query for unique locus, raises error if missing or more than one locus in the database.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
filter_column
|
str
|
Name of column to filter results by. |
required |
filter_value
|
str
|
Only return loci which have this value in the their filter_column. |
required |
feature
|
str
|
Feature names such as 'transcript', 'gene', and 'exon' |
required |
Returns:
| Type | Description |
|---|---|
Locus
|
The unique matching genomic locus. |
SequenceData ¶
SequenceData(fasta_paths, cache_directory_path=None)
Container for reference nucleotide and amino acid sequenes.
fasta_directory_paths
instance-attribute
¶
fasta_directory_paths = [split(path)[0] for path in self.fasta_paths]
fasta_filenames
instance-attribute
¶
fasta_filenames = [split(path)[1] for path in self.fasta_paths]
cache_directory_paths
instance-attribute
¶
cache_directory_paths = [cache_directory_path] * len(self.fasta_paths)
fasta_dictionary_filenames
instance-attribute
¶
fasta_dictionary_filenames = [filename + '.pickle' for filename in self.fasta_filenames]
fasta_dictionary_pickle_paths
instance-attribute
¶
fasta_dictionary_pickle_paths = [join(cache_path, filename) for cache_path, filename in zip(self.cache_directory_paths, self.fasta_dictionary_filenames)]
delete_index_files ¶
delete_index_files()
Delete cached FASTA dictionaries while preserving source files.
stored_id ¶
stored_id(sequence_id)
FASTA record ID for sequence_id, or None if it has no record.
A bare Ensembl ID matches the record whatever its version, and raises
ValueError if several versions are present. A versioned ID must
match the header's version: ValueError if the header carries a
different version or none.
recorded_versions ¶
recorded_versions(sequence_id)
Sorted versions that FASTA headers record for this stable ID.
fasta_version ¶
fasta_version(sequence_id)
Return the integer FASTA-header version for sequence_id, or
None if the header didn't carry a version (e.g. older Ensembl
releases) or the ID isn't in the FASTA.
Accepts either the versioned or bare form of an ENS ID — the
bare form is resolved through _stripped_index.
prune_genome_fastas ¶
prune_genome_fastas(dry_run=False, cache_root=None)
Remove owned shared objects with no release references.
Return (path, bytes) pairs. Malformed manifests abort the entire operation; dry_run performs the same reference checks without removing anything. Objects being downloaded or indexed are skipped. Attached local files, private downloads, and symlinks are never candidates.
MissingGenomeFastaError ¶
Reference DNA was not configured or has not been downloaded.