ruwex 0.1.0

Fast Rust rewrite of wikiextractor: extract and clean text from Wikimedia XML dumps
Documentation

ruwex

Rust Wiki Extractor (ruwex) is a fast Rust rewrite of wikiextractor: extracts and cleans plain text from Wikimedia XML dumps. Usable as a Rust library and as drop-in replacement binarieswikiextractor, extractPage, and cirrus-extract — accepting the same command-line parameters as the Python originals.

Why

  • Fast. Full English Wikipedia (25.7M pages, 26.5 GB bz2) in ~24 minutes with template expansion, or ~4.6 minutes without, on an Apple M4 Pro — where Python wikiextractor extrapolates to roughly 13 hours. On identical workloads it measures 8× faster; see BENCHMARKS.md.
  • Parallel where it counts. Wikimedia multistream dumps are read with seek-and-decompress parallelism across all cores (the index file is detected automatically), 6.5× faster than serial bz2 decoding of the same content.
  • Compatible. Output is byte-for-byte identical to Python wikiextractor on golden fixtures (default, --json, --links, and template-expansion modes), and --templates cache files are interchangeable between the two tools.
  • Less mangling. Several upstream bugs are deliberately fixed rather than replicated: text corruption from stale comment offsets, raw XML leaking into article text, articles silently dropped by a multiprocessing race, and empty documents emitted for redirects. See Divergences below.

Install

From crates.io:

cargo install ruwex            # installs wikiextractor, extractPage, cirrus-extract

Or from a clone of this repository:

cargo install --path .

As a library, add it to your Cargo.toml:

cargo add ruwex

Usage

Same CLI as Python wikiextractor:

# extract to out/AA/wiki_00, ... with template expansion (two passes)
wikiextractor enwiki-latest-pages-articles-multistream.xml.bz2 -o out

# keep the multistream index file next to the dump for parallel reading:
#   enwiki-latest-pages-articles-multistream.xml.bz2
#   enwiki-latest-pages-articles-multistream-index.txt.bz2

# faster, without template expansion; JSON lines; to stdout
wikiextractor dump.xml.bz2 --no-templates --json -o -

# reuse a template cache (created on first run; also readable by the Python tool)
wikiextractor dump.xml.bz2 --templates templates.cache -o out

# extract one page by exact title, as HTML, to stdout
wikiextractor enwiki-latest-pages-articles-multistream.xml.bz2 --title "Anarchism" --html

# single page by id (raw XML), CirrusSearch dumps
extractPage --id 12 dump.xml.bz2
cirrus-extract enwiki-cirrussearch-content.json.gz -o out

Fast by-title lookup (--title)

--title "Some Page" extracts a single page by title to stdout, honoring --html / --json / --links / --html-safe. It requires a multistream dump (with its *-index.txt.bz2 alongside). The title is normalized — underscores become spaces and surrounding whitespace is trimmed, so --title Richard_Dawkins finds Richard Dawkins — but internal spacing and first-letter capitalization are left as typed (both can be significant). On first use it builds a compact FST title index next to the dump — e.g. enwiki-…-multistream.xml.title.fst, ~435 MB for full enwiki (25.7M titles), ~40 s one-time — and reuses it automatically afterwards (rebuilding only if the dump is newer). A lookup then seeks straight to the one ~100-page bz2 block that holds the page: ~0.1 s, versus tens of seconds to scan the raw index.

Templates are expanded by default, lazily through the same index: each template the page uses is fetched on demand (the title index locates it; the containing block is decompressed at most once), so no separate template database is built. Expansion adds well under a second — full enwiki lookups land around 0.8–1.5 s even for template-heavy pages:

wikiextractor enwiki-…-multistream.xml.bz2 --title "Anarchism" --html   # expanded, ~1 s
wikiextractor enwiki-…-multistream.xml.bz2 --title "Anarchism" --no-templates   # ~0.1 s

--no-templates skips expansion for the fastest lookup. --templates FILE forces the bulk template database instead (built once from a full scan and cached, ~3.5 min / ~1.6 GB) — only needed when you want expansion identical to a full-dump run rather than the lazy equivalent.

Note: as with Python wikiextractor, {{#invoke:}} Lua modules ({{convert}}, many {{lang}} variants, most infobox internals) are not executed, so on modern enwiki the visible effect of expansion on the extracted plain text is often small — the templates that would add inline prose are mostly Lua.

As a library:

use std::path::Path;
use std::sync::Arc;
use ruwex::{ExtractorConfig, PageSource, ShardedWriter, expand::templates};

let config = ExtractorConfig::default();
let db = Arc::new(templates::load(Some(Path::new("dump.xml.bz2")), None, config.workers)?);
let source = PageSource::open(Path::new("dump.xml.bz2"))?;
let mut sink = ShardedWriter::create(Path::new("out"), 1 << 20, false)?;
let stats = ruwex::run_with_templates(source, &db, &config, &mut sink)?;

Compatibility and divergences

The cleaner, template expander, parser functions, and output formats are ports of the Python implementation, validated against it byte-for-byte on fixtures and on real dumps. Divergences are deliberate: ruwex fixes upstream bugs (comment-offset text corruption, XML leaking from empty <text/> elements, randomly dropped articles, empty redirect documents, -ns being ignored under --no-templates) and otherwise matches Python's behavior, quirks included.

{{#invoke:}} (Scribunto/Lua modules) is not executed — invocations expand to nothing, exactly as in the Python tool. Running Lua modules behind an optional feature is possible future work; a faithful minimal Scribunto environment is a substantial project on its own.

Development

cargo test                                  # unit + golden-fixture + e2e tests
cargo clippy --all-targets -- -D warnings
benchmarks/run.sh                           # benchmark suite → BENCHMARKS.md
benchmarks/compare.sh                       # side-by-side diff vs Python wikiextractor

Golden fixtures under tests/fixtures/ were generated by running the Python implementation in-process (tests/fixtures/gen_golden.py); the tests themselves never require Python.

License

AGPL-3.0, the license of the Python wikiextractor this is a port of.