Read genomic DNA¶
Reference DNA lets you read any genomic interval, including introns, intergenic regions and flanking sequence. It is optional: a normal installation does not download a whole genome. Human DNA is about 1 GB to download and takes several GB of disk once decompressed.
Quick start¶
# Annotation and DNA, or just the DNA:
pyensembl install --release 93 --species human --with-genome-fasta
pyensembl install --release 93 --species human --only-genome-fasta
from pyensembl import EnsemblRelease
data = EnsemblRelease(93, species="human", genome_fasta=True)
data.download_genome_fasta() # does nothing if the DNA is already installed
with data:
bases = data.sequence("7", 117_480_000, 117_480_100)
tp53 = data.gene_by_id("ENSG00000141510") # TP53, on the minus strand
tp53_dna = data.sequence(
tp53.contig, tp53.start, tp53.end, strand=tp53.strand
)
genome_fasta=True only chooses the DNA. Nothing is downloaded until you call
download_genome_fasta() or download(), or run pyensembl install.
Python objects use reference DNA only when constructed with genome_fasta:
a plain EnsemblRelease(93) does not pick up DNA installed by the CLI, and its
error message names the call that does.
Reading sequences¶
sequence(contig, start, end, mask="upper", *, strand="+"):
- Coordinates are one-based and inclusive, like the rest of PyEnsembl,
with
1 <= start <= end <= contig length. - Bases are read from the plus strand;
strand="-"returns the reverse complement, so a gene or transcript reads 5' to 3'. - Contigs can be named as in the FASTA or as PyEnsembl reports them
(
gene.contig). If the names differ only by achrprefix, the error suggests the right one. - Results are uppercase;
mask="raw"keeps soft-masked repeats in lowercase. - Absent contigs and invalid ranges raise
ValueError. Missing DNA raisesMissingGenomeFastaError, aValueErrorwhose message explains how to install or enable it. Reads never download anything.
Related attributes and methods:
fastais a pyfaidx reader with zero-based, half-open slices (fasta[contig][start - 1:end].seq), as used by Varcode. It needs the FASTA's own contig names, and isNonewhen DNA is not configured or not installed.genome_fasta_pathis the uncompressed FASTA on disk, orNone.download()andindex()include DNA when it is configured.index_genome_fasta()builds the DNA index before the first query needs it.close()closes the reader. Readers already handed out stay usable afterclear_cache().- Attached DNA does not affect equality: genes and transcripts from the same release compare equal with or without it.
Choosing DNA¶
Python (EnsemblRelease) |
CLI (pyensembl install) |
|
|---|---|---|
| Ensembl's DNA | genome_fasta=True |
--with-genome-fasta or --only-genome-fasta |
| A local FASTA | genome_fasta="/data/ref.fa.gz" |
--genome-fasta-path /data/ref.fa.gz |
| Coverage | genome_fasta_type="primary_assembly" |
--genome-fasta-type primary_assembly |
| Masking | genome_fasta_mask="soft" |
--masked soft |
The default is unmasked toplevel DNA, which covers the patch and haplotype
contigs in Ensembl annotations. primary_assembly has the chromosomes and
unplaced/unlocalized sequences but no patches or haplotypes, and some older
releases and species don't provide it. Masking is none, soft (repeats in
lowercase), or hard (repeats replaced with N). See Ensembl's
DNA file definitions.
Local FASTA files¶
data = EnsemblRelease(
93, species="human", genome_fasta="/data/my_reference.fa.gz"
)
# Custom annotations can attach DNA too, from a path or URL:
from pyensembl import Genome
custom = Genome("custom", "my_annotations",
genome_fasta_path_or_url="/data/reference.fa")
Plain FASTA files are read in place; gzip and BGZF files are decompressed into
PyEnsembl's cache on first use. Indexes always live in the cache, so read-only
source directories work and your own files and indexes are never modified.
index() warns about annotation contigs that are missing from a local FASTA,
but matching contig names don't prove that the assembly matches.
Disk space and shared files ¶
Compatible releases share one copy of Ensembl DNA. To delete DNA, prune unused copies or share a cache with other users, see free disk space and how reference DNA is stored.
Upgrading from 2.11.0¶
| 2.11.0 | 2.12.0 and later |
|---|---|
EnsemblRelease(81, download_genome_fasta=True) |
EnsemblRelease(81, genome_fasta=True) |
EnsemblRelease(81, genome_fasta_path="/data/ref.fa") |
EnsemblRelease(81, genome_fasta="/data/ref.fa") |
The old keywords still work but emit a DeprecationWarning. Objects pickled or
serialized by 2.11.0 still load. CLI flags are unchanged.