twobitreader
This crate provides fast DNA sequence extraction from 2bit files, a standard format in bioinformatics.
The motivation for twobitreader is speed; see benchmarks below.
It is also available as a Python package named twobitreader-rs.
The focus is raw reading from 2bit, but fast concatenation and reverse-complement methods are also provided to make higher-level use cases easier.
Examples
Extracting sequences is straightforward:
let tbr = open?; // Human genome, build 38
let seq = tbr.get; // -> String ("TAACC")
Concatenation works by iterating over (start, end) pairs. For example, assembling a spliced transcript:
// Exon ranges for human FKHL6 gene transcript (Gencode v43)
let exons = ; // Exon 2 (start, end)
let transcript = tbr.concat; // -> String
Parallelism is easy with crates like rayon.
For example, batch extraction of sequences:
use *;
let args = ;
let seqs = args.into_par_iter
.map
.; // -> Vec<String>
Or, assembling a batch of spliced transcripts in parallel:
use reverse_complement;
use *;
let transcripts = ;
// Concatenate exons and reverse-complement if negative strand.
// (Correct when exons are listed in genome-coordinate order.)
let seqs = transcripts.into_par_iter
.map
.; // HashMap<&str, String>
let seq = &seqs; // -> &String to transcript sequence
Cold files are an order of magnitude slower to access than files already in memory ("hot"). Use prefetching to dramatically improve single-threaded speed:
let exons = ;
tbr.prefetch; // Ask the operating system to start paging this data from disk.
let seqs = tbr.get_batch; // Access the memory as it arrives.
Benchmarks
Two tasks were benchmarked:
- exons: extract 133,388 distinct human exon sequences;
- transcripts: concatenate 319,468 exons into 29,211 human spliced transcript sequences.
Speed depends on parallelism and page cache (hot vs cold):
- hot runs represent repeated or interactive dna extraction scenarios;
- cold runs represent a first run of a genomics pipeline, bound by disk speed;
- prefetch runs are cold but with a
prefetchcall preceding extraction.
The table below shows running times in milliseconds. Experimental details are BENCH.md.
This crate provides twobitreader (pure rust) and twobitreader_rs (python wrapper).
| EXONS | 1-thread / hot | 1-thread / cold | 1-thread / prefetch | 16-thread / hot | 16-thread / cold |
|---|---|---|---|---|---|
| twobitreader (rs) | 90 | 1,700 | 180 | 10 | 190 |
| twobitreader_rs (rs, py) | 100 | 1,900 | 190 | *27 | *210 |
| py2bit (c, py) | 220 | 2,800 | n/a | n/a | n/a |
| GenomeKit (cpp, py) | 340 | 2,400 | n/a | n/a | n/a |
| twobit (rs) | 390 | 3,500 | n/a | n/a | n/a |
| twobitToFa (c) | 1,000 | 4,500 | n/a | 340 | 850 |
| twobitreader (py) | 7,200 | 13,000 | n/a | 1,900 | 2,300 |
| Biopython (py) | 8,100 | 11,000 | n/a | n/a | n/a |
| TRANSCRIPTS | 1-thread / hot | 1-thread / cold | 1-thread / prefetch | 16-thread / hot | 16-thread / cold |
|---|---|---|---|---|---|
| twobitreader (rs) | 140 | 1,700 | 230 | 13 | 190 |
| twobitreader_rs (rs, py) | 180 | 2,100 | 310 | *27 | *220 |
| py2bit (c, py) | 490 | 3,600 | n/a | n/a | n/a |
| GenomeKit (cpp, py) | 720 | 2,900 | n/a | n/a | n/a |
| twobit (rs) | 930 | 1,900 | n/a | n/a | n/a |
| twobitToFa (c) | 2,100 | 6,900 | n/a | 840 | 1,200 |
| twobitreader (py) | 13,000 | 21,000 | n/a | 2,800 | 3,500 |
| Biopython (py) | 19,000 | 25,000 | n/a | n/a | n/a |
Entries marked * were run in free-threaded Python.
Dependencies
byteorderfor handling endian-nessmemmap2for memory mapping the 2bit fileseq-macrofor generating 2bit decoder lookup tablelibcfor prefetching file ranges on Apple targetswindows-sysfor prefetching file ranges on Windows targets
Use of AI
Claude Code: generated the OS-specific prefetch loops; improved handling of corrupt or malicious files; documented the Python bindings and mirrored their tests and benchmark; improved error checking and propagation to Python more broadly; and generated the CI configurations.
License: MIT OR Apache-2.0