twobitreader
This crate provides fast DNA sequence extraction from 2bit files, a standard format in bioinformatics.
The motivation for this crate is speed. Extracting sequences is consistently faster than the best alternative. The focus is raw reading from 2bit, but fast concatenation and reverse-complement methods are also provided to make higher-level use cases easier.
Examples
Extracting sequences is straightforward:
let tbr = open?; // Human genome, build 38
let seq = tbr.get; // -> String ("TAACC")
Concatenation works by iterating over (start, end) pairs. For example, assembling a spliced transcript:
// Exon ranges for human FKHL6 gene transcript (Gencode v43)
let exons = ; // Exon 2 (start, end)
let transcript = tbr.concat; // -> String
Parallelism is easy with crates like rayon.
For example, batch extraction of sequences:
use *;
let args = ;
let seqs = args.into_par_iter
.map
.; // -> Vec<String>
Or, assembling a batch of spliced transcripts in parallel:
use reverse_complement;
use *;
let transcripts = ;
// Concatenate exons and then reverse-complement if necessary.
// (Correct if exons listed in genome-coordinate order.)
let seqs = transcripts.into_par_iter
.map
.; // HashMap<&str, String>
let seq = &seqs; // -> &String to transcript sequence
Cold files are an order of magnitude slower to access than files already in memory ("hot"). Use prefetching to dramatically improve single-threaded speed:
let exons = ;
tbr.prefetch; // Ask the operating system to start paging this data from disk.
let seqs = tbr.get_batch; // Access the memory as it arrives.
Speed
Two tasks were benchmarked:
- exons: extract 133,388 distinct human exon sequences;
- transcripts: concatenate 319,468 exons into 29,180 human spliced transcript sequences.
Speed depends on parallelism and page cache (hot vs cold):
- hot runs represent repeated or interactive dna extraction scenarios;
- cold runs represent a first run of a genomics pipeline, bound by disk speed;
- prefetch runs are cold but with a
prefetchcall preceding extraction.
The table below shows running times in milliseconds. Experimental details are BENCH.md.
| EXONS | 1-thread / hot | 1-thread / cold | 1-thread / prefetch | 16-thread / hot | 16-thread / cold |
|---|---|---|---|---|---|
| twobitreader (rust) | 90 | 1,600 | 180 | 11 | 190 |
| py2bit (C, python) | 220 | 2,800 | n/a | n/a | n/a |
| GenomeKit (C++, python) | 340 | 2,400 | n/a | n/a | n/a |
| twobit (rust) | 390 | 3,500 | n/a | n/a | n/a |
| twobitToFa (C) | 1,000 | 4,500 | n/a | 340 | 850 |
| twobitreader (python) | 7,200 | 13,000 | n/a | 1,900 | 2,300 |
| Biopython (python) | 8,100 | 11,000 | n/a | n/a | n/a |
| TRANSCRIPTS | 1-thread / hot | 1-thread / cold | 1-thread / prefetch | 16-thread / hot | 16-thread / cold |
|---|---|---|---|---|---|
| twobitreader (rust) | 160 | 2,500 | 240 | 18 | 210 |
| py2bit (C, python) | 490 | 3,600 | n/a | n/a | n/a |
| GenomeKit (C++, python) | 720 | 2,900 | n/a | n/a | n/a |
| twobit (rust) | 930 | 1,900 | n/a | n/a | n/a |
| twobitToFa (C) | 2,100 | 6,900 | n/a | 840 | 1,200 |
| twobitreader (python) | 13,000 | 21,000 | n/a | 2,800 | 3,500 |
| Biopython (python) | 19,000 | 25,000 | n/a | n/a | n/a |
Dependencies
byteorderfor handling endian-nessmemmap2for memory mapping the 2bit fileseq-macrofor generating 2bit decoder lookup tablelibcfor prefetching file ranges on Apple targetswindows-sysfor prefetching file ranges on Windows targets
License: MIT OR Apache-2.0