A fast, accurate, multi-threaded toolkit for trimming and filtering short-read FASTQ data, written in Rust. Its name is the plural of chela, the pincer-like claws of crustaceans: a nod to what it does to FASTQ reads.
Contents
Every option, the input and output rules, and more examples are in docs/usage.md.
Overview
chelae trims and filters short-read FASTQ in one multi-threaded pass, and identifies the adapters in a library you know nothing about.
- Accurate paired-end adapter trimming.
chelaefinds each pair's insert from where R1 and R2 overlap, then checks the implied adapter against known adapter sequences. That catches adapters too short for sequence matching alone to find, without mistaking adapter-like sequence inside the insert for adapter. - Fast. SIMD kernels and a pipeline built for many cores trimmed about 1.7 M read pairs per second on 8 cores in our benchmark.
- Benchmarked. Against six other trimmers,
chelaewas the fastest, 1.25× faster than the runner-up and 2.6–5.7× faster than cutadapt, trim-galore-rs and fastp, and the most accurate on 8 of 11 simulated datasets. See Performance. - One pass does it all: poly-G, adapter, read-structure hard-trimming with UMI extraction (which also trims the mate's UMI from reads that run through a short insert), poly-X and quality trimming, then length, N-base and quality filters.
- Adapter detection.
chelae detectreports the adapters in a library and writes them as FASTA forchelae trim --adapter-fasta. - Fits into pipelines. Split or interleaved paired-end files, stdin and stdout, gzip or BGZF input detected automatically, and a fastp-compatible JSON report for MultiQC.
Every option, the input and output rules, and more examples are in docs/usage.md.
Examples
Paired-end reads can be trimmed without naming any adapters. chelae finds adapters from where R1 and R2 overlap, and checks each one against every adapter in its built-in database (TruSeq, Nextera, small RNA, AVITI and MGI/DNBSEQ):
When you know the kit, name it - specify the right kit will make chelae a little faster and a little more accurate (since it can't match to the wrong kits) The following example also moves an 8 bp UMI from the start of R1 into the read name, skips the next 4 bases, and quality-trims 3' ends:
When a pair's insert is short enough that R2 reads through to the far end of the molecule, R2 ends in the reverse complement of R1's UMI and skip bases, so chelae trims those 12 bases from R2 as well. See Read-structures on paired-end reads.
-i and -o default to stdin and stdout, and interleaved paired-end input is recognized from its read names, so chelae can sit in a pipeline with no intermediate files:
| | |
For a library of unknown provenance, chelae detect finds the adapters, and its FASTA feeds straight back into chelae trim:
For single-end reads, chelae detect reports which built-in kit matches best:
Performance
chelae was benchmarked against six other FASTQ trimmers on simulated short-read libraries, for runtime and for adapter-trim accuracy against the simulator's ground truth. RESULTS.md has the method, the dataset definitions and every tool's numbers.
Runtime
Wall seconds for 50 M read pairs at 8 threads, median of 3 runs (WGS, cfDNA):
| Tool | WGS, insert 350 ± 60 | cfDNA, insert 170 ± 30 |
|---|---|---|
| chelae | 29.7 | 29.0 |
| adapterremoval | 37.4 (1.26×) | 36.1 (1.25×) |
| cutadapt | 77.6 (2.61×) | 81.9 (2.83×) |
| trim-galore-rs | 80.0 (2.69×) | 93.9 (3.24×) |
| fastp | 169.6 (5.71×) | 132.5 (4.58×) |
- Both libraries are simulated 2×150 paired-end reads with TruSeq adapters, trimmed with a realistic configuration: adapter, poly-G, sliding-window quality and length trimming, and an N filter.
- The host was an EC2
r8a.2xlarge(AMD EPYC 9R45, Zen 5, 8 cores). Inputs and outputs were on a RAM disk so the timings measure the tools rather than storage, and every tool with a compression option wrote gzip at level 1. - The three runs of each tool agreed to within 6%.
- bbduk and trimmomatic weren't timed at this scale: on the paired-end accuracy datasets below, they took 8–16× and 4–27× as long as
chelae(see broad toolset timings).
| Tool | User CPU s, WGS | User CPU s, cfDNA | Max RSS |
|---|---|---|---|
| chelae | 207 | 202 | 117 MB |
| adapterremoval | 257 | 237 | 55 MB |
| cutadapt | 358 | 418 | 76 MB |
| trim-galore-rs | 618 | 731 | 140 MB |
| fastp | 1,251 | 974 | 1,274 MB |
Accuracy
Eight paired-end and three single-end datasets of 1–4 M reads each simulate common library types: WGS at several insert sizes, cfDNA, exome with Nextera adapters, miRNA, and a high error rate (definitions). Accuracy is the RMSE of each read's trim point against the truth, in bases (lower is better).
| Tool | Most accurate on | Median RMSE, paired-end | Median RMSE, single-end |
|---|---|---|---|
| chelae | 8 of 11 | 0.013 | 0.265 |
| adapterremoval | 2 of 11 | 0.100 | 0.412 |
| cutadapt | 0 of 11 | 0.237 | 0.270 |
| trim-galore-rs | 0 of 11 | 0.237 | 0.270 |
| fastp | 1 of 11 | 1.687 | 0.280 |
| trimmomatic | 0 of 11 | 1.063 | 1.172 |
| bbduk | 0 of 11 | 1.599 | 0.902 |
ar = adapterremoval, tg-rs = trim-galore-rs, tmatic = trimmomatic. Each ID links to that dataset's full results.
| ID | Layout | Insert | Err rate | #1 | #2 | #3 |
|---|---|---|---|---|---|---|
| 1 | 2×150 | 150 ± 30 | 0.1–1% | chelae 0.039 | ar 0.050 | fastp 0.169 |
| 2 | 2×150 | 250 ± 40 | 0.1–1% | chelae 0.013 | ar 0.113 | cutadapt 0.193 |
| 3 | 2×150 | 350 ± 60 | 0.1–1% | chelae 0.008 | cutadapt 0.099 | tg-rs 0.099 |
| 4 | 2×150 | 450 ± 80 | 0.1–1% | chelae 0.011 | cutadapt 0.089 | tg-rs 0.089 |
| 5 | 2×250 | 450 ± 80 | 0.1–1% | chelae 0.010 | ar 0.130 | cutadapt 0.165 |
| 6 | 2×150 | 250 ± 60 | 1–5% | ar 0.095 | chelae 0.417 | cutadapt 1.314 |
| 7 | 2×150 | 170 ± 30 | 0.1–1% | chelae 0.041 | ar 0.077 | fastp 0.220 |
| 8 | 2×76 | 140 ± 25 | 0.1–1% | chelae 0.013 | ar 0.074 | cutadapt 0.280 |
| 9 | 1×150 | 300 ± 80 | 0.1–1% | chelae 0.265 | cutadapt 0.270 | tg-rs 0.270 |
| 10 | 1×150 | 120 ± 30 | 0.1–1% | fastp 0.650 | ar 0.742 | chelae 0.801 |
| 11 | 1×76 | 30 ± 2 | 0.1–1% | ar 0.000 | tmatic 0.040 | chelae 0.065 |
Dataset 8 is an exome library with Nextera adapters and dataset 11 is miRNA; the rest are WGS or cfDNA with TruSeq adapters. On dataset 11 the top six tools differ by at most 5 of 4.2 M reads.
Versions: chelae 0.2.0, adapterremoval 3.0.2, bbduk 40.02, cutadapt 5.2, fastp 1.3.7, trim-galore-rs 2.3.0 and trimmomatic 0.41, the latest on bioconda as of 2026-09-25.
Installing
From bioconda
Using pixi, after adding the bioconda channel:
pixi add chelae
Or using your favorite conda client (conda, mamba, micromamba, …):
conda install -c bioconda chelae
With cargo
With Rust 1.89 or newer installed:
cargo install chelae
To build from source, see CONTRIBUTING.md.
About Fulcrum Genomics
Visit us at Fulcrum Genomics to learn more about how we can power your Bioinformatics with chelae and beyond.