coremlit 0.1.1

Safe, synchronous CoreML runtime for macOS (CPU/GPU/Neural Engine) with opt-in on-device multimodal pipelines: speech (Whisper STT, forced alignment, speaker diarization, Silero VAD), AudioSet sound-event tagging, and audio/text/image embeddings (CLAP, granite, SigLIP)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
//! Bridge between the chordai base960h CTC vocabulary and the
//! `tokenizers`-crate schema asry's seam builder needs (design spec §3.1/§6,
//! `docs/superpowers/specs/2026-07-11-alignkit-forced-alignment-design.md`).
//!
//! `chordai/wav2vec2-base960h-aligner-coreml` ships a raw `{token: id}` CTC
//! dict (`Models/alignkit/base960h_dict.json`), not a HuggingFace
//! `tokenizer.json` — both asry's own `Aligner::from_paths` and alignkit's
//! [`crate::audio::align::aligner::Aligner::from_paths`] (via
//! [`asry::emissions::EmissionsAligner::builder`]) need the latter. This
//! module owns the derived, committed asset that fills that gap
//! (`assets/chordai_base960h_tokenizer.json`) plus the vocabulary constants
//! the seam validates a loaded tokenizer against — and [`Vocabulary`], which
//! applies the same rule set at run time to the table ANY model ships beside
//! it, so an aligner spells with its model's own alphabet.
//!
//! # Generator note (reproducibility record)
//!
//! `assets/chordai_base960h_tokenizer.json` is mechanically derived from
//! `Models/alignkit/base960h_dict.json` (SHA-256
//! `ef41495ab958d4416ad2f81ea51a77d4a3c79cace96e92e978c443c7bfbdd2e5`, the
//! same file `tests/model_io.rs` pins) and the staged model's contract,
//! [`AcousticContract::BASE960H`], by this rule set — re-running it
//! reproduces the asset field for field, and the aligner runs it for any
//! table under any contract:
//!
//! 1. Parse the dict file as a flat JSON object `{token: id}` (29 entries).
//! 2. Copy every `(token, id)` pair unmodified into `model.vocab`. Key
//!    order is not semantically meaningful — the `tokenizers` crate's
//!    `WordLevel` model deserializes `vocab` into a hash map — so the
//!    asset simply preserves the dict's own id-ascending order for
//!    reviewability.
//! 3. Set `model.type = "WordLevel"` and `model.unk_token = "<unk>"`.
//!    `unk_token` is a REQUIRED key for `tokenizers` 0.23's `WordLevel`
//!    deserializer (its visitor's `missing_fields` check covers `vocab`
//!    and `unk_token`), but it is never validated against `vocab` at parse
//!    time — `WordLevelBuilder::build` stores whatever string it's given
//!    unchecked, and `Model::get_vocab_size` counts only `vocab`'s own
//!    entries. `"<unk>"` is deliberately NOT one of the 29 vocab entries
//!    (this CTC alphabet has no unknown-token concept), and doesn't need
//!    to be for the file to parse or for `VOCAB_SIZE` to stay exactly 29.
//! 4. Set every other top-level field but `added_tokens` to its schema
//!    default: `version = "1.0"` (the only value `tokenizers` 0.23 accepts),
//!    `truncation` / `padding` / `normalizer` / `pre_tokenizer` /
//!    `post_processor` / `decoder` = `null`. Neither this crate's vocab
//!    bridge nor asry's own runtime tokenization
//!    (`asry/src/runner/aligner/algorithm/tokenize.rs`) ever calls
//!    `Tokenizer::encode` — both go through `token_to_id` /
//!    `get_vocab_size` directly — so these pipeline fields are inert for
//!    this asset's purpose.
//! 5. Declare the contract's non-lexical tokens the table spells as special
//!    added tokens, in id order: its blank (the token at the contract's
//!    blank id), its word delimiter (`|` or `" "`, unless it has none), and every
//!    special it names. Each is `{"id", "content", "single_word": false,
//!    "lstrip": false, "rstrip": false, "normalized": false, "special":
//!    true}` at its own vocabulary id, so the table's ids and its size are
//!    unchanged. A model with special tokens declares them in its tokenizer
//!    document as special added tokens, and asry reads its reserved ids off
//!    `added_tokens[].special`, never off a spelling: no transcript character
//!    is spelled onto a reserved column. Under the staged contract that is
//!    `-` (id 0, the blank) and `|` (id 1, the delimiter). An added token
//!    whose content is empty is dropped when the `tokenizers` crate parses the
//!    document (`AddedVocabulary::add_tokens` skips it), so a table's empty
//!    entry is declared here and never reserved by the declaration: the
//!    aligner refuses a declared special spelled that way unless the blank's
//!    id or the stated delimiter reserves it
//!    ([`TokenizationError::EmptySpecial`](crate::audio::align::error::TokenizationError::EmptySpecial)),
//!    and after building the seam it reads back the set the seam reserves and
//!    refuses one that is not this step's
//!    ([`AlignerError::ReservedSetMismatch`](crate::audio::align::error::AlignerError::ReservedSetMismatch)).
//!
//! Step 4's claim about asry holds since asry 0.2 (asry#21); this crate
//! requires 0.3. asry 0.1 classified each character by running it alone
//! through `Tokenizer::encode`, and a `WordLevel` model whose declared
//! `unk_token` is absent from its vocabulary — step 3's shape, on purpose —
//! answers that with `MissingUnkToken` for every character outside the 29: the
//! whole chunk failed before any OOV policy could decide the character. asry
//! 0.2 looks the (ASCII-uppercased) character up with `Tokenizer::token_to_id`
//! instead, so a character with no entry is an `OovKind::Symbol` event, and the
//! absent unknown token costs nothing.
//!
//! # Where the asset is consumed
//!
//! Parsing a vocabulary into a live tokenizer, and reporting a parse or
//! delimiter failure, both happen inside asry's seam builder when an
//! [`crate::audio::align::aligner::Aligner`] hands it the tokenizer document a
//! [`Vocabulary`] writes for the model's contract — for the bundled table
//! under [`AcousticContract::BASE960H`], the document
//! [`tokenizer_json_bytes`] records; that failure surfaces as
//! [`crate::audio::align::error::AlignerError::Seam`]. This module constructs
//! no `Tokenizer` itself: [`Vocabulary::from_json`] only validates a table's
//! tokens and ids, and the document is written by the rule set above. The
//! aligner parses the document it handed asry once more, with the same
//! `tokenizers` crate, to read back the columns the seam reserves.
//!
//! # A table does not say which class is the blank
//!
//! A flat `{token: id}` table names columns; it does not say which one the head
//! scores as "no token here". Names are no answer: HuggingFace calls its blank
//! `<pad>`, chordai and torchaudio call theirs `-` at id 0, and a table can hold
//! a `<pad>` or a `-` that is an ordinary class beside a blank of another name.
//! So a [`Vocabulary`] carries no blank at all, and no special either: the
//! blank, the delimiter and the specials are the model's [`AcousticContract`]
//! statement, which the aligner checks against the table's ids at load and
//! declares in the tokenizer document it writes for that contract.

use core::num::NonZeroUsize;
use std::{
  borrow::Cow,
  collections::{BTreeMap, BTreeSet, btree_map::Entry},
  path::Path,
};

use crate::audio::align::{
  acoustic::{AcousticContract, Tokenization, WordDelimiter},
  error::{MissingId, VocabularyError, VocabularyRead},
};

/// Number of entries in the chordai base960h CTC vocabulary, including the
/// blank and word-delimiter tokens.
///
/// Derived from `Models/alignkit/base960h_dict.json` (see this module's
/// `# Generator note`). asry's `validate_vocab_dim` requires the CTC head's
/// output width `V` to equal the tokenizer's vocab size EXACTLY;
/// `base960h_aligner.mlmodelc` declares `emissions` `[1, 2999, 29]`
/// (`tests/model_io.rs`), so this is the width of the model the bundled table
/// belongs to. The encoder reads a model's width rather than assuming this
/// one; [`Vocabulary::size`] is what an aligner pairs it with.
pub const VOCAB_SIZE: usize = 29;

/// CTC blank-token id in the chordai base960h vocabulary.
///
/// The dict maps the literal token `"-"` to id `0` — chordai's own CTC
/// blank convention. This is distinct from the `<pad>` / `[PAD]` /
/// `<blank>` special-token probe asry's `detect_blank_token_id` performs by
/// default: this vocabulary has no `<pad>`-style entry at all, only the bare
/// `"-"` at id `0`. It is the blank of
/// [`AcousticContract::BASE960H`](crate::audio::align::acoustic::AcousticContract::BASE960H),
/// which every aligner passes to the seam builder's `.blank_token_id(..)`
/// explicitly (the default auto-detect would fail construction here, and
/// guess by name elsewhere).
pub const BLANK_ID: u32 = 0;

/// wav2vec2 inter-word delimiter token.
///
/// asry resolves the delimiter dynamically via `tokenizer.token_to_id("|")`
/// (`asry/src/runner/aligner/aligner.rs:1132`, in
/// `validate_word_delimiter_present`) rather than assuming a fixed id; this
/// constant is the TOKEN STRING that lookup uses, not its id — id `1` in
/// this vocabulary (`tests::word_delimiter_resolves_via_token_to_id`).
pub const WORD_DELIMITER: &str = "|";

/// Bytes of the committed tokenizer asset
/// (`assets/chordai_base960h_tokenizer.json`): the document written for the
/// bundled table under [`AcousticContract::BASE960H`], its blank and its
/// delimiter declared special, in the `tokenizers`-crate schema asry's loader
/// accepts on its fast path. Unlike the model
/// artifacts under the gitignored `Models/` store, this asset is
/// deliberately committed: it is a small authored text file this crate
/// owns, not a downloaded artifact. Its schema is an explicit
/// `"model": {"type": "WordLevel", ...}` object never needs the
/// `load_tokenizer_with_compat` compat-patch shim
/// (`asry/src/runner/aligner/aligner.rs:1198`) that exists only for
/// upstream exports missing that discriminator.
///
/// # Why bytes, not a path
///
/// `include_bytes!` embeds the asset in the compiled artifact at build
/// time. A path helper built on `env!("CARGO_MANIFEST_DIR")` would only
/// resolve on the machine and source tree that built the crate — it reads
/// back correctly today only by accident of running in-tree, and breaks
/// the moment the crate is used as an installed/packaged dependency
/// elsewhere. Bytes also match asry's loader one step further downstream
/// than a path would: `load_tokenizer_with_compat` immediately turns
/// whatever path it's given into bytes (`std::fs::read`) before ever
/// calling `Tokenizer::from_bytes` — never `Tokenizer::from_file`, despite
/// that function's own error-message text saying so. The document
/// [`crate::audio::align::aligner::Aligner::from_paths`] hands to
/// [`asry::emissions::EmissionsAligner::builder`], with no filesystem
/// round-trip, is written by the same rule set and holds exactly these fields
/// (`tests::the_bundled_document_declares_exactly_the_staged_contracts_specials`).
/// A vocabulary that ships beside a model is the caller's file, like the model
/// itself, and is read through [`Vocabulary::from_file`].
pub const fn tokenizer_json_bytes() -> &'static [u8] {
  include_bytes!("../assets/chordai_base960h_tokenizer.json")
}

/// [`VOCAB_SIZE`] as the [`NonZeroUsize`] a [`Vocabulary`] carries. The
/// conversion is infallible: `VOCAB_SIZE` is the nonzero constant `29`.
const BUNDLED_SIZE: NonZeroUsize = match NonZeroUsize::new(VOCAB_SIZE) {
  Some(size) => size,
  None => unreachable!(),
};

/// A `{token: id}` JSON object read entry by entry, so a token the object
/// names twice reaches [`Vocabulary::from_json`] twice. A map's insert would
/// keep one of the two ids and say nothing.
struct Entries(Vec<(String, u32)>);

impl<'de> serde::Deserialize<'de> for Entries {
  fn deserialize<D: serde::Deserializer<'de>>(deserializer: D) -> Result<Self, D::Error> {
    struct Visit;
    impl<'de> serde::de::Visitor<'de> for Visit {
      type Value = Entries;
      fn expecting(&self, f: &mut core::fmt::Formatter<'_>) -> core::fmt::Result {
        f.write_str("a JSON object mapping each token to its id")
      }
      fn visit_map<A: serde::de::MapAccess<'de>>(self, mut map: A) -> Result<Entries, A::Error> {
        let mut entries = Vec::new();
        while let Some(entry) = map.next_entry::<String, u32>()? {
          entries.push(entry);
        }
        Ok(Entries(entries))
      }
    }
    deserializer.deserialize_map(Visit)
  }
}

/// The unknown token a written tokenizer document declares (the generator
/// note's step 3).
const UNKNOWN_TOKEN: &str = "<unk>";

/// The bundled table's tokens, in id order: `base960h_dict.json`'s 29 entries,
/// the ones [`tokenizer_json_bytes`] spells
/// (`tests::bundled_tokens_are_the_committed_tables`).
const BUNDLED_TOKENS: [Cow<'static, str>; VOCAB_SIZE] = [
  Cow::Borrowed("-"),
  Cow::Borrowed("|"),
  Cow::Borrowed("E"),
  Cow::Borrowed("T"),
  Cow::Borrowed("A"),
  Cow::Borrowed("O"),
  Cow::Borrowed("N"),
  Cow::Borrowed("I"),
  Cow::Borrowed("H"),
  Cow::Borrowed("S"),
  Cow::Borrowed("R"),
  Cow::Borrowed("D"),
  Cow::Borrowed("L"),
  Cow::Borrowed("U"),
  Cow::Borrowed("M"),
  Cow::Borrowed("W"),
  Cow::Borrowed("C"),
  Cow::Borrowed("F"),
  Cow::Borrowed("G"),
  Cow::Borrowed("Y"),
  Cow::Borrowed("P"),
  Cow::Borrowed("B"),
  Cow::Borrowed("V"),
  Cow::Borrowed("K"),
  Cow::Borrowed("'"),
  Cow::Borrowed("X"),
  Cow::Borrowed("J"),
  Cow::Borrowed("Q"),
  Cow::Borrowed("Z"),
];

/// A CTC vocabulary: the table an aligner spells with, one entry per class of
/// its model's CTC head.
///
/// [`Self::bundled`] is the 29-class English table every
/// [`Aligner::from_paths`](crate::audio::align::aligner::Aligner::from_paths)
/// binds. A model that ships its own table beside it — a flat `{token: id}`
/// JSON object, as chordai's `base960h_dict.json` and HuggingFace's
/// `vocab.json` are — is read with [`Self::from_file`] (or [`Self::from_json`])
/// and paired with that model by
/// [`Aligner::from_paths_with_vocabulary`](crate::audio::align::aligner::Aligner::from_paths_with_vocabulary),
/// which refuses the pair at load unless the table has exactly one entry per
/// class of the model's head. This is how an aligner comes to spell a language
/// other than English: the model supplies the alphabet, not this crate.
///
/// A vocabulary names columns and carries no blank and no special: which
/// column is the blank, which token delimits words and which tokens are never
/// letters is the model's [`AcousticContract`] statement (see the module doc's
/// "A table does not say which class is the blank").
#[derive(Clone)]
pub struct Vocabulary {
  /// The table's tokens, in id order: the id of each is its index.
  tokens: Cow<'static, [Cow<'static, str>]>,
  /// Number of entries: the CTC head width this table names.
  size: NonZeroUsize,
}

impl Vocabulary {
  /// The bundled 29-class English table (chordai base960h), whose document
  /// under [`AcousticContract::BASE960H`] is [`tokenizer_json_bytes`]. Its
  /// blank is `-`, id [`BLANK_ID`], which that contract names.
  #[must_use]
  pub const fn bundled() -> Self {
    Self {
      tokens: Cow::Borrowed(&BUNDLED_TOKENS),
      size: BUNDLED_SIZE,
    }
  }

  /// Read a flat `{token: id}` JSON table — the shape of the vocabulary a CTC
  /// model ships beside it.
  ///
  /// The object must name each token once, and every id in `0..n` exactly once
  /// (`n` its entry count): a CTC head has one column per class, and each id
  /// is the column its token is scored in. No entry is taken for the blank or
  /// a special, whatever its name: those are the model's contract's to state.
  /// The tokenizer document asry parses is written for a contract by this
  /// module's generator rule set, the one the bundled asset was derived by:
  /// read through here, the staged `base960h_dict.json` yields the bundled
  /// table.
  ///
  /// # Errors
  /// [`VocabularyError::Parse`] if `json` is not a JSON object mapping each
  /// token to a non-negative integer id that fits a `u32`;
  /// [`VocabularyError::DuplicateToken`] if the object names a token twice;
  /// [`VocabularyError::Empty`] if it names no token;
  /// [`VocabularyError::MissingId`] if an id in `0..n` names no token.
  pub fn from_json(json: &[u8]) -> Result<Self, VocabularyError> {
    let Entries(entries) =
      serde_json::from_slice(json).map_err(|error| VocabularyError::Parse(error.to_string()))?;
    let mut table = BTreeMap::new();
    for (token, id) in entries {
      match table.entry(token) {
        Entry::Occupied(repeated) => {
          return Err(VocabularyError::DuplicateToken(repeated.key().clone()));
        }
        Entry::Vacant(slot) => {
          slot.insert(id);
        }
      }
    }
    let size = NonZeroUsize::new(table.len()).ok_or(VocabularyError::Empty)?;

    // `n` ids, each in `0..n` at most once, is exactly "each id in `0..n`
    // once": a duplicate or an id past the end leaves one below it unnamed.
    let mut named = vec![false; size.get()];
    for &id in table.values() {
      if let Some(slot) = usize::try_from(id).ok().and_then(|id| named.get_mut(id)) {
        *slot = true;
      }
    }
    if let Some(id) = named.iter().position(|named| !named) {
      return Err(VocabularyError::MissingId(MissingId::new(id, size.get())));
    }
    // Every id in `0..n` is named once (just checked), so each slot is
    // written exactly once.
    let mut tokens = vec![Cow::Borrowed(""); size.get()];
    for (token, id) in table {
      if let Some(slot) = usize::try_from(id).ok().and_then(|id| tokens.get_mut(id)) {
        *slot = Cow::Owned(token);
      }
    }
    Ok(Self {
      tokens: Cow::Owned(tokens),
      size,
    })
  }

  /// Read the `{token: id}` table in the file at `path` — the vocabulary that
  /// ships beside a model, such as `base960h_dict.json` beside
  /// `base960h_aligner.mlmodelc`.
  ///
  /// # Errors
  /// [`VocabularyError::Read`] if the file cannot be read; otherwise as
  /// [`Self::from_json`].
  pub fn from_file(path: impl AsRef<Path>) -> Result<Self, VocabularyError> {
    let path = path.as_ref();
    let json = std::fs::read(path)
      .map_err(|source| VocabularyError::Read(VocabularyRead::new(path.to_path_buf(), source)))?;
    Self::from_json(&json)
  }

  /// Number of entries: the width of the CTC head a model paired with this
  /// table must have.
  #[must_use]
  pub const fn size(&self) -> NonZeroUsize {
    self.size
  }

  /// The tokenizer document asry's seam builder parses for a model of this
  /// table under `contract`, written by the module doc's generator rule set:
  /// the table as a `WordLevel` vocabulary, and `contract`'s non-lexical tokens
  /// declared as special added tokens at their own ids
  /// ([`Self::non_lexical`]), so asry reserves their columns.
  pub(crate) fn tokenizer_json(&self, contract: &AcousticContract) -> Vec<u8> {
    let vocab: BTreeMap<&str, usize> = self
      .tokens()
      .enumerate()
      .map(|(id, token)| (token, id))
      .collect();
    let added_tokens: Vec<serde_json::Value> = self
      .non_lexical(contract.blank(), contract.tokenization())
      .into_iter()
      .map(|id| {
        serde_json::json!({
          "id": id,
          "content": self.tokens[id],
          "single_word": false,
          "lstrip": false,
          "rstrip": false,
          "normalized": false,
          "special": true,
        })
      })
      .collect();
    serde_json::json!({
      "version": "1.0",
      "truncation": null,
      "padding": null,
      "added_tokens": added_tokens,
      "normalizer": null,
      "pre_tokenizer": null,
      "post_processor": null,
      "decoder": null,
      "model": {
        "type": "WordLevel",
        "vocab": vocab,
        "unk_token": UNKNOWN_TOKEN,
      },
    })
    .to_string()
    .into_bytes()
  }

  /// The ids of a contract's non-lexical tokens this table spells, in id
  /// order: the token at its `blank` id, its word delimiter, and every special
  /// `tokenization` names. Each is the contract's statement, never inferred
  /// from a spelling; a statement naming no entry of the table names nothing
  /// here.
  ///
  /// The one definition of what is not a letter: the tokenizer document
  /// declares exactly these special, the aligner refuses at load a seam that
  /// does not reserve exactly these, and every lexical check at load —
  /// whitespace, letter case, granularity — reads every other entry
  /// ([`Self::lexical`]).
  pub(crate) fn non_lexical(&self, blank: u32, tokenization: Tokenization) -> BTreeSet<usize> {
    let blank = usize::try_from(blank)
      .ok()
      .filter(|&id| id < self.size.get());
    let delimiter = match tokenization.delimiter() {
      stated @ (WordDelimiter::Pipe | WordDelimiter::Space) => self.id_of(stated.seam_token()),
      WordDelimiter::Absent => None,
    };
    let specials = tokenization
      .specials()
      .iter()
      .filter_map(|special| self.id_of(special));
    blank.into_iter().chain(delimiter).chain(specials).collect()
  }

  /// The table's LEXICAL tokens under a contract stating `blank` and
  /// `tokenization`: every entry but [`Self::non_lexical`]'s, in id order.
  pub(crate) fn lexical(
    &self,
    blank: u32,
    tokenization: Tokenization,
  ) -> impl Iterator<Item = &str> {
    let reserved = self.non_lexical(blank, tokenization);
    self
      .tokens()
      .enumerate()
      .filter(move |(id, _)| !reserved.contains(id))
      .map(|(_, token)| token)
  }

  /// The id of `token`, when the table spells it.
  pub(crate) fn id_of(&self, token: &str) -> Option<usize> {
    self.tokens().position(|spelled| spelled == token)
  }

  /// The table's tokens, in id order.
  pub(crate) fn tokens(&self) -> impl Iterator<Item = &str> {
    self.tokens.iter().map(|token| token.as_ref())
  }

  /// Whether the table spells `token`.
  pub(crate) fn contains(&self, token: &str) -> bool {
    self.tokens().any(|spelled| spelled == token)
  }
}

/// The table's size rather than the tokenizer document's bytes.
impl core::fmt::Debug for Vocabulary {
  fn fmt(&self, f: &mut core::fmt::Formatter<'_>) -> core::fmt::Result {
    f.debug_struct("Vocabulary")
      .field("size", &self.size)
      .finish_non_exhaustive()
  }
}

#[cfg(test)]
mod tests;