1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
//! Test corpus utilities for validation and benchmarking.
//!
//! This module provides parsers and generators for standard spelling correction
//! test corpora, including:
//!
//! - **Norvig's big.txt**: Natural language word frequency distribution
//! - **Birkbeck corpus**: Real-world spelling errors from native speakers
//! - **Holbrook corpus**: Secondary school spelling errors
//! - **Aspell corpus**: Technical and domain-specific term errors
//! - **Wikipedia corpus**: Common web-scale misspellings
//!
//! ## Usage
//!
//! This module is only available in test and benchmark builds:
//!
//! ```rust,ignore
//! use liblevenshtein::corpus::{BigTxtCorpus, MittonCorpus};
//!
//! // Load Norvig's big.txt for dictionary construction
//! let big_txt = BigTxtCorpus::load("data/corpora/big.txt")?;
//! let unique_words = big_txt.unique_words(); // ~32K words
//!
//! // Load Holbrook corpus for error validation
//! let holbrook = MittonCorpus::load("data/corpora/holbrook.dat")?;
//! for (correct, misspellings) in &holbrook.errors {
//! for (misspelling, frequency) in misspellings {
//! // Validate spell correction...
//! }
//! }
//! ```
//!
//! ## Corpus Formats
//!
//! ### Norvig's big.txt
//!
//! Plain text, one token per line (preserving frequency):
//!
//! ```text
//! the
//! the
//! the
//! ...
//! ```
//!
//! ### Mitton `.dat` Format
//!
//! Used by Holbrook, Aspell, and Wikipedia corpora:
//!
//! ```text
//! $correct_word
//! misspelling1 frequency1
//! misspelling2 frequency2
//! ...
//! ```
//!
//! Each correct word is preceded by `$` and followed by its misspellings
//! with optional frequency counts (default: 1).
pub use ;
pub use ;