Please check the build logs for more information.
See Builds for ideas on how to fix a failed build, or Metadata for how to configure docs.rs builds.
If you believe this is docs.rs' fault, open an issue.
Deacon
Deacon filters DNA sequences in FASTA/Q files and streams using SIMD-accelerated minimizer comparison with query sequence(s), emitting either matching sequences (search mode), or sequences without matches (deplete mode). Sequences match when they share enough distinct minimizers with the indexed query to exceed chosen absolute and relative thresholds. Query size has little impact on filtering speed, enabling ultrafast search and depletion with gene-, genome- and pangenome-scale queries using a laptop. Deacon filters uncompressed FASTA/Q at gigabases per second on recent AMD, Intel (x86_64), and Apple arm64 systems. Built with panhuman host depletion in mind—yet broadly useful for searching large sequence collections—Deacon delivers leading classification accuracy for host depletion and unrivalled speed using 5GB of RAM.
Default parameters are carefully chosen but easily changed. Classification sensitivity, specificity and memory requirements may be tuned by varying k-mer length (-k), window size (-w), absolute match threshold (-a) and relative match threshold (-r) . Minimizer k and w are chosen at query index time, while the match thresholds can be chosen at filter time. Matching sequences are those that share enough distinct minimizers with the indexed query to exceed both the absolute threshold (-a, default 2 shared minimizers) and the relative threshold (-r, default 0.01 [1%] shared minimizers). For paired sequences, hits in either mate counts towards a single match threshold for the pair. Deacon reports filtering performance during execution and optionally writes a JSON --summary upon completion. Sequences can optionally be renamed using --rename for privacy and smaller file sizes. Deacon fully supports stdin, stdout and natively handles gz, zst and xz compression formats, detected by file extension.
Benchmarks for panhuman host depletion of complex microbial metagenomes are described in a preprint. Deacon with the panhuman-1 (k=31, w=15) index exhibited the highest balanced accuracy for both long and short simulated reads. Deacon was less specific only than Hostile for short reads.
Use cases
- Depletion of human or other host genome sequences in FASTQ reads or streams.
- Ultrafast binary classification of genes, genomes or pangenomes in terabase genome catalogues like AllTheBacteria without tedious pre-indexing.
Install
Conda/mamba/pixi 
Cargo 
RUSTFLAGS="-C target-cpu=native"
[!IMPORTANT] Cargo installation requires Rust 1.88 or newer. Update using
rustup update.
Docker 
Containers are available from the BioContainers registry.
Quickstart
Ultrafast panhuman host depletion
# Download validated 3GB human pangenome index (version 0.13.0 or later)
# Deplete long reads
# Deplete short paired reads
Ultrafast gene/genome/pangenome search
N.B. Indexing a 3Gbp human genome takes ~30s using 18GB of RAM with default parameters. Filtering uses 5GB.
Prebuilt indexes
Prebuilt pangenome indexes are provided for human and mouse host classification and depletion. These can be downloaded using the links below, or with deacon index fetch <name>.
| Name & URL | Composition | Minimizers | Subtracted minimizers | Size | Date |
|---|---|---|---|---|---|
panhuman-1 (k=31, w=15) Cloud, Zenodo |
HPRC Year 1 ∪ CHM13v2.0 ∪ GRCh38.p14 - bacteria (FDA-ARGOS) - viruses (RefSeq) |
409,907,949 | 20,671 (0.0050%) | 3.3GB | 2025-04 |
panmouse-1 (k=31, w=15) Cloud, Zenodo |
GRCm39 ∪ PRJEB47108 - bacteria (FDA-ARGOS) - viruses (RefSeq) |
551,041,865 | 9,866 (0.0018%) | 4.4GB | 2025-11 |
[!NOTE]
Index compatibility. Deacon
0.11.0and above uses index format version 3. Using version 3 indexes with older Deacon versions and vice versa triggers an error. Prebuilt indexes in legacy formats are archived in object storage and Zenodo to ensure reproducibility. To download indexes in legacy formats, replace the/3/in any prebuilt index download URL with either/2/or/1/accordingly.
- Deacon
0.11.0and above uses index format version3- Deacon
0.7.0through to0.10.0used index format version2- Deacon
0.1.0through to0.6.0used index format version1
Usage
Filtering
The main command deacon filter accepts an index path followed by up to two FASTA/FASTQ file paths, depending on whether input sequences originate from stdin, a single file, or paired input files. Indexes are built with deacon index build. Paired queries are supported as either separate files or interleaved stdin, and written interleaved to either stdout or file, or else to separate paired output files. For paired reads, distinct minimizer hits originating from either mate are counted. By default, input sequences must meet both an absolute threshold of 2 minimizer hits (-a 2) and a relative threshold of 1% of minimizers (-r 0.01) to pass the filter. Filtering can be inverted for e.g. host depletion using the --deplete (-d) flag. Gzip, Zstandard, and xz compression formats are detected automatically by file extension.
Examples
# Keep only sequences matching a collection of genes
# Host depletion using the panhuman-1 index
# High sensitivity host depletion with absolute threshold of 1 and no relative threshold
# High specificity 10% relative match threshold
# Stdin and stdout
|
# True multithreaded gzip decompression with rapidgzip
|
# Zstandard compression
# Paired reads
|
# Save summary JSON
# Replace read headers with incrementing integers
# Only look for minimizer hits inside the first 1000bp per record
# Output FASTA regardless of input format (discards quality scores)
# Debug mode: see sequences with minimizer hits in stderr
[!NOTE]
deacon filteruses 8 threads by default. Using more threads (e.g.--threads 16) can accelerate filtering given sufficient resources, especially with uncompressed sequences whose processing is not rate limited by decompression. Since version0.13.0, Deacon writes gzipped output files (e.g-o out.fastq.gz) in parallel, providing particular practical benefit for gzipped paired reads. If output file(s) ending in.gzare detected, total--threadsare allocated 1:1 to compression and filtering tasks respectively. Gzip compression thread allocation can be overriden with--compression-threads.
Indexing
# Index one FASTA/FASTQ file
# Index many FASTA/FASTQ files using stdin
|
deacon index build accepts either a FASTA/FASTQ file or a stdin stream (-), enabling convenient indexing of compressed sequences in one or many files with a single step. Indexing a human genome takes a few seconds. Indexing uses 2-4x as much RAM as filtering. For indexing large collections approaching terabase scale—such as mammalian pangenomes—it may be practical to index genomes individually in parallel and later combine them using the deacon index union set operation, described below.
Set operations
A differentiating feature of Deacon is the ease of combining, subtracting and intersecting minimizer indexes. For example, deacon index diffcan be used to subtract shared minimizers between target and host genomes when building custom indexes for host depletion.
-
Use
deacon index union 1.idx 2.idx 3.idx… > 1+2+3.idxto succinctly combine two or more indexes. -
Use
deacon index diff 1.idx 2.idx > 1-2.idxto subtract minimizers in 2.idx from 1.idx. Useful for masking out shared minimizer content between e.g. target and host genomes.deacon index diffalso supports subtracting minimizers from an index using a fastx file or stream directly, e.g.deacon index diff 1.idx 2.fa.gz > 1-2.idxorzcat *.fa.gz | deacon index diff 1.idx - > 1-2.idx. This enables diffing with larger-than-memory sequence collections if desired.
-
Use
deacon index intersect 1.idx 2.idx… > 1∩2.idxto find the intersection of minimizers in two or more indexes.
Inspecting indexes
- Use
deacon index info 1.idxto display index information including minimizer k and w parameters, number of minimizers, and index format version. - Use
deacon index dump 1.idx > 1.fato dump a minimizer index to FASTA.
Command line reference
Filtering
<INDEX> Path
)
)
)
)
)
; )
)
)
)
& ; )
Indexing
)
)
)
)
<INPUT> Path )
)
)
)
)
Filtering summary statistics
Use -s summary.json to save detailed filtering statistics:
Server mode
From version 0.11.0, it is possible to eliminate index loading overhead at the start of each filter operation by preloading the index in the memory of a local server process. Subsequent filtering commands with --use-server are executed by the server process using a UNIX socket. Having started a server process, the index of the first filtering command it receives persists in memory for the life of that server process, enabling subsequent filter commands to be served rapidly without hash set construction overhead.
# Start the server
# The first filter command loads the index as usual
# Subsequent filter commands use the existing index stored in memory
# Stop the server
Workflow manager integration
Nextflow (nf-core)
- Modules
deacon_indexanddeacon_filter - Subworkflow
fastq_index_filter_deacon
Galaxy
Work in progress: https://github.com/galaxyproject/tools-iuc/pull/7473
Citation
Bede Constantinides, John Lees, Derrick W Crook. "Deacon: fast sequence filtering and contaminant depletion" bioRxiv 2025.06.09.658732, https://doi.org/10.1101/2025.06.09.658732
Please also consider citing the SimdMinimizers paper:
Ragnar Groot Koerkamp, Igor Martayan. "SimdMinimizers: Computing random minimizers, fast" bioRxiv 2025.01.27.634998, https://doi.org/10.1101/2025.01.27.634998