Skip to content

Install and manage data

PyEnsembl downloads annotation and sequence files once and indexes them in a local cache. After that, queries need no network access.

Install data

Install from the command line:

pyensembl install --release 93 --species human

or from Python, for example in a notebook or pipeline:

from pyensembl import EnsemblRelease

data = EnsemblRelease(93, species="human")
data.download()
data.index()

Both skip files that are already downloaded or indexed, so rerunning finishes an interrupted install. install prints one progress line per step on stderr, with progress bars in a terminal; add --verbose (-v) to see every download and database step. In Python, pass show_progress=True to download() and index(). Use --overwrite or overwrite=True to replace existing files.

Choose a reference explains release, species and assembly options. Whole-genome DNA is a separate download.

Check what is installed

pyensembl list
Species  Assembly  Release  Annotation   Reference DNA      Location
human    GRCh38    81       indexed      toplevel, indexed  ~/Library/Caches/pyensembl/GRCh38/ensembl81
human    GRCh38    82       not indexed  -                  ~/Library/Caches/pyensembl/GRCh38/ensembl82
custom   GRCm38    mine1    indexed      -                  ~/Library/Caches/pyensembl/GRCm38/mine1

Annotation is indexed when everything is downloaded and indexed, so queries need no network access or setup. not indexed means the files are downloaded but the first query would spend minutes indexing them, and incomplete means some downloads are missing. invalid means a source is empty or not a regular file; inaccessible means it cannot be read. Run pyensembl install for that release to finish (add --species for non-human genomes; custom genomes need their original install options).

In Python, installed() is True when everything a genome is configured with, reference DNA included, is downloaded and indexed. It only reads the cache. genome_for_reference_name picks the newest installed release of an assembly, else the newest downloaded one, else the newest supported Ensembl release:

from pyensembl import EnsemblRelease, genome_for_reference_name
EnsemblRelease(93).installed()
genome_for_reference_name("GRCh38")

# Ensembl releases with any files in the cache, ready or not:
from pyensembl.shell import collect_all_installed_ensembl_releases
collect_all_installed_ensembl_releases()

Cache location

PyEnsembl keeps all of its data under one directory, with a subdirectory per genome (<reference>/<annotation><version>, e.g. GRCh38/ensembl81). By default this is the platform cache directory that datacache chooses: ~/.cache/pyensembl on Linux, ~/Library/Caches/pyensembl on macOS, and %LOCALAPPDATA%\pyensembl\pyensembl\Cache on Windows. Releases that PyEnsembl 2.16 or earlier installed on Windows stay in their old per-genome directories and keep working. To use another location, set PYENSEMBL_CACHE_DIR; the data then goes in its pyensembl subdirectory:

export PYENSEMBL_CACHE_DIR=/custom/cache/dir

In Python, set os.environ["PYENSEMBL_CACHE_DIR"] before creating a genome.

Free disk space

pyensembl delete-index-files --release 93  # keep downloads, reindex on use
pyensembl delete-all-files --release 93    # all of release 93's files
pyensembl prune --dry-run                  # list unused shared DNA
pyensembl prune

Add --species for non-human releases. Compatible releases share one copy of Ensembl DNA, so deleting a release keeps DNA that other releases still use, and prune removes DNA that no release references. It skips DNA that is being downloaded or indexed, never touches local FASTA files, and deletes nothing if any release's DNA metadata is malformed (pyensembl list shows which one). delete-index-files keeps shared DNA indexes because other releases may use them; rebuild one with index_genome_fasta(overwrite=True). In Python, prune_genome_fastas(dry_run=True) returns (path, bytes) candidates, and pyensembl list --check-genome-fasta verifies each release's DNA index.

Inspect files without installing

Inspect a selected genome's source files and indexes, including paths, availability, sizes and any recorded download provenance:

pyensembl inspect --release 93
pyensembl inspect --release 93 --with-genome-fasta --json
# Custom genomes use the same --gtf / --transcript-fasta / --protein-fasta
# and --reference-name / --annotation-name options as install.

In Python, EnsemblRelease(93).inspect_data() returns annotation and optional reference-DNA readiness, installed, and a files dictionary keyed by role (for example gtf, gtf_index, transcript_fasta_1). Values are datacache FileInspection objects with path, status, error, size, mtime, source_url, fetched_at, recorded_sha256 and verified fields. CLI JSON is an array of reports, with errors rendered as strings. Add --check-genome-fasta (Python: check_genome_fasta=True) to validate existing DNA indexes more thoroughly.

Inspection only reads: no network, copying, directory creation, indexing or pickle deserialization. It checks that source files are readable, regular and nonempty and that SQLite indexes are complete; it does not validate biological contents. Recorded provenance is advisory, not a trusted checksum, and files from older versions have none. Downloads are published atomically, so a failed overwrite keeps the previous file; replace an invalid file with pyensembl install --overwrite.

Share a cache

Reads take no locks and write nothing, so a fully installed and indexed cache can be read-only for other users. A download or index build locks only the file it writes; registering, deleting and pruning releases briefly lock the whole cache.

New files follow your umask, as do lock files on Python 3.10+, so set umask 002 (or default ACLs) before installing into a group-shared cache. Files downloaded before PyEnsembl 2.13.1 were readable only by their owner; share them with chmod -R g+rX "$PYENSEMBL_CACHE_DIR/pyensembl" (or the platform cache directory). dna_cache may be a symlink, e.g. to a larger disk.

How reference DNA is stored

Ensembl DNA is stored once per upstream file under pyensembl/dna_cache/:

pyensembl/dna_cache/
  homo_sapiens/ftp.ensembl.org/GRCh38-GCA_000001405.18/
    toplevel/unmasked/fasta/<file key>/
      sequence.fa        uncompressed, even when downloaded as .fa.gz
      sequence.fa.fai
      object.json        full identity of the upstream file
      index.json

Before downloading, PyEnsembl reads Ensembl's small README and CHECKSUMS files to see whether another release already has the same file. The versioned assembly accession distinguishes assembly patches, and the 16-character file key (a SHA-256 prefix of the assembly, Ensembl's checksum and compressed size) distinguishes upstream revisions. These are metadata checks: Ensembl's Unix checksums are not cryptographic hashes. If the metadata is incomplete, the assembly directory ends in -unverified and each release keeps its own copy. Local FASTA files and custom mirrors are never shared.

Downloads retry transient HTTP failures and stalls of five minutes, and are checked against the upstream size. An interrupted DNA download resumes on the next install on POSIX systems, using only bytes the server confirms come from the same file; on Windows it starts over.