1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
// The workspace-root anchor every `models_dir()` below resolves against, and
// the sibling-checkout anchor the oracle gates read. FOUND by searching upward
// for the `[workspace]` manifest, never counted in `../` hops — see its module
// doc for why a count is the wrong shape here. Re-exported so the binaries
// that pull this `common` in share the one resolver.
pub use ;
use ;
/// Directory containing the downloaded alignkit model artifacts.
///
/// Overridable via `ALIGNKIT_TEST_MODELS`; otherwise falls back to
/// `<workspace>/Models/alignkit` — gitignored, fetched dev-time (mirrors
/// whisperkit's `WHISPERKIT_TEST_MODELS`/`Models/` and dia-coreml's
/// `DIA_COREML_TEST_MODELS`/`Models/dia-coreml` conventions, one directory
/// level down for this crate's own model set).
/// Path to the compiled forced-aligner artifact.
///
/// Compiled from the downloaded `base960h_aligner.mlpackage` via `xcrun
/// coremlcompiler compile` at model-acquisition time (`coremlit::Model::load`
/// only accepts a compiled `.mlmodelc`; see `tests/model_io.rs`'s module doc
/// for the full acquisition record: source, revision, licence, per-file
/// SHA-256).
/// Path to the `{token: id}` CTC vocabulary the model ships beside it
/// (`base960h_dict.json`, the table the bundled tokenizer asset was derived
/// from).
///
/// `#[allow(dead_code)]`: only `tests/align/align_chunk.rs` uses it.
/// Path to the 60 s @ 16 kHz mono fixture used by the graph-truth test and
/// by `tests/parity_words.rs`'s **unpadded** half.
///
/// alignkit has no committed audio fixtures of its own. `ted_60.wav` in the
/// whisperkit crate's `tests/fixtures/audio/` is already exactly 960,000
/// samples (60.000000 s @ 16 kHz mono int16, `afinfo`-verified at write
/// time) — precisely the `[1, 960000]` window `base960h_aligner.mlmodelc`
/// requires, with no padding needed — so this crate borrows it by relative
/// path instead of committing a second copy of a ~1.9 MB binary fixture that
/// would then need to stay byte-identical to the original forever. Both
/// crates live in this workspace and move together.
///
/// That exact-window property is the whole reason the parity gate wants this
/// clip: see [`TED_60_TRANSCRIPT`].
///
/// `#[allow(dead_code)]`: only `tests/model_io.rs` and `tests/parity_words.rs`
/// use it; the per-binary `common` copy in `tests/align_chunk.rs` does not.
/// The verified transcript for [`ted_60_wav_path`]'s audio — the opening 60 s
/// of Tim Urban's TED talk *Inside the mind of a master procrastinator*.
///
/// # Why this clip has a transcript at all
///
/// `jfk.wav` is 176,000 samples; the encoder window is 960,000, so alignkit
/// zero-pads it by 81.7% and the **unpadded path has never been gated**.
/// `ted_60.wav` is exactly 960,000 samples, so
/// `Encoder::emissions_raw` takes its `Cow::Borrowed` branch — *no zeros are
/// ever appended* — and `truncated_frame_count(960_000, 2999)` keeps all 2,999
/// frames instead of jfk's 549. Forced alignment needs a transcript, and this
/// clip shipped without one, which is exactly why the gap survived B5.
///
/// # Provenance: ASR, because ASR is what feeds forced alignment
///
/// Produced by **this workspace's own whisperkit** — the real production
/// pipeline, ASR → forced alignment — via
/// `cargo run -p coremlit --features whisper --example whisper_transcribe_wav`, on
/// `openai_whisper-large-v3` (`argmaxinc/whisperkit-coreml`). It is therefore
/// the *kind* of text a caller actually aligns: readable ASR output, not a
/// hand-made verbatim transcription.
///
/// # Verification — three sources, and the ASR lost twice
///
/// A transcript is an **input** to forced alignment: a wrong word does not
/// fail, it silently becomes a wrong alignment target. So this text was
/// cross-checked against two independent readings of the same audio —
/// whisper-small (the same pipeline, a weaker checkpoint) and a **greedy CTC
/// decode of alignkit's own wav2vec2 emissions**, which is the acoustic model
/// that will actually consume it — and every disagreement was settled against
/// the emission posteriors and the RMS envelope, never by vote:
///
/// - **`ninety-page`, not `90-page`** (large-v3's spelling). The 29-class CTC
/// vocabulary has **no digits**, so `90` can only align as out-of-vocabulary
/// wildcards. The greedy decode reads the audio as `NINETY | PAGE` — two
/// words — and [`asry::EnglishNormalizer`] splits on the hyphen, so the
/// spoken form lands as exactly those two words. Spoken form is the correct
/// register for a forced-alignment target.
/// - **`happen to every single paper`**, not whisper-small's `happen in`. The
/// greedy decode emits a bare `T` at 37,520 ms between `HAPPEN` and `EVERY`
/// — the /t/ that `in` does not have. large-v3 agrees.
/// - **`everything gets done and things stay civil`** — the `and` is real,
/// though the greedy decode drops it. Frames 1055–1057 (21,100–21,140 ms)
/// carry `A`/`I` → `N` → `D` posterior mass in sequence beneath a dominant
/// word-delimiter, over a non-zero RMS of 0.081 → 0.062 → 0.021: a reduced,
/// unstressed /ənd/ in the 160 ms gap. Greedy CTC routinely swallows those
/// (it also lost the `may` of `maybe` and the `st` of `stay` right here).
/// - **`I knew for a paper like that`** — large-v3's leading **`And` is a
/// hallucination and is NOT in this transcript.** Frames 2285–2296
/// (45,700–45,920 ms) are digital silence: RMS 0.001–0.003, `logP(blank)`
/// fp16-saturated at exactly `0.0`, and no letter posterior above −8. Speech
/// resumes at 45,920 ms and the model fires `I` at 45,940 ms with a −0.06
/// log-prob — there is neither room nor evidence for an `And`. It is a
/// textbook Whisper segment-initial discourse-marker insertion: large-v3
/// opened a new segment at exactly 45.84 s, right after that pause, and
/// whisper-small — which did not break a segment there — never wrote it.
///
/// The clip's last word is complete: the greedy decode closes `IT` at
/// 59,980 ms, 20 ms inside the 60,000 ms edge.
///
/// # What is deliberately NOT here
///
/// Whisper elides disfluencies, and at least three survive in the audio: a
/// false start (`I would`… ~27,520–28,220 ms) before `I would have it all
/// ready to go`, a stammered `then` (~30,300–30,700 ms) before `actually`, and
/// a `would would` repetition near 31,800 ms. They are **left out on purpose**
/// — this is ASR output, which is what production feeds an aligner, and *both*
/// aligners receive the identical text, so the omission cannot bias the
/// comparison. It does leave real speech with no transcript word under it,
/// which makes those spots the natural places for the trellis to diverge;
/// `tests/parity_words.rs`'s divergence ledger is where that shows up, and it
/// is pinned rather than tolerated.
///
/// `#[allow(dead_code)]`: only `tests/parity_words.rs` uses it.
pub const TED_60_TRANSCRIPT: &str = concat!;
/// SHA-256 of [`ted_60_wav_path`]'s **decoded** buffer — the 960,000 f32
/// samples [`load_wav_mono_f32`] returns, hashed as little-endian bytes.
/// Exactly [`JFK_SAMPLES_SHA256`]'s role, for the second clip: it pins the
/// audio the gate's ted_60 bounds were measured on, and it is *also* the pin
/// that the clip still fills the window exactly — a re-encode that changed the
/// length by one sample would silently move alignkit onto its zero-padding
/// branch and quietly retire the very path this fixture exists to cover.
///
/// `#[allow(dead_code)]`: only `tests/parity_words.rs` uses it.
pub const TED_60_SAMPLES_SHA256: &str =
"b14ed488eb68545e49893bd424d78a0849941b97c3f042c3e4461e3bfb513dd5";
/// Path to the 11 s @ 16 kHz mono `jfk.wav` fixture (176,000 samples, well
/// inside the encoder's 960,000 window), borrowed from the whisperkit crate
/// exactly as [`ted_60_wav_path`] is. Its known transcript is
/// [`JFK_TRANSCRIPT`]; together they drive `tests/align_chunk.rs`'s
/// end-to-end alignment.
///
/// `#[allow(dead_code)]`: only `tests/align_chunk.rs` uses it.
/// The known transcript for [`jfk_wav_path`]'s audio (whisperkit's
/// `tests/fixtures/golden/jfk_tiny_golden.json`).
///
/// `#[allow(dead_code)]`: only `tests/align_chunk.rs` uses it.
pub const JFK_TRANSCRIPT: &str = "And so my fellow Americans ask not what your country can do for \
you, ask what you can do for your country.";
/// SHA-256 of [`jfk_wav_path`]'s **decoded** buffer — the 176,000 f32
/// samples [`load_wav_mono_f32`] returns, hashed as little-endian bytes.
///
/// This is the input-identity pin for `tests/parity_words.rs`. That gate
/// compares alignkit's word timings against asry's ONNX aligner, and such a
/// comparison is worth exactly nothing if the two sides are not looking at
/// the same audio: the FIRST attempt at an alignkit-vs-asry comparison
/// (`.superpowers/sdd/alignkit-gate1-diagnostic.md`) reported an alarming
/// "86.6% divergence" that turned out to be a harness bug — one side got a
/// padded buffer, the other an unpadded one. The number was measuring the
/// harness, not the models.
///
/// The gate feeds one `Vec<f32>`, by reference, to both aligners, so
/// buffer identity holds by construction; this digest additionally pins the
/// FIXTURE, so a `jfk.wav` that is silently re-encoded, resampled, or
/// swapped out from under the cross-crate relative path fails loudly instead
/// of re-measuring parity on different audio.
///
/// `#[allow(dead_code)]`: only `tests/parity_words.rs` uses it.
pub const JFK_SAMPLES_SHA256: &str =
"ebd52851100536db02d12c49fddd010372dcdc70243562e057553d476b706ae0";
/// Lowercase-hex SHA-256 of a decoded sample buffer, over its little-endian
/// `f32` bytes. Backs [`JFK_SAMPLES_SHA256`].
///
/// `#[allow(dead_code)]`: only `tests/parity_words.rs` uses it.
/// Directory holding asry's ONNX wav2vec2 oracle — the `models/` directory of a
/// co-located `asry` checkout (a sibling of this repo). This is TEST DATA, not
/// the code dependency: alignkit depends on asry as a crates.io version
/// (`coremlit/Cargo.toml`), so building the crate does NOT put asry's
/// `models/` on disk. The default path below assumes the dev-worktree layout (a
/// sibling `asry`); set `ALIGNKIT_ASRY_MODELS` when it lives elsewhere.
///
/// `#[allow(dead_code)]`: only `tests/parity_words.rs` uses it.
/// asry's ONNX wav2vec2-base-960h export (`onnx-community/
/// wav2vec2-base-960h-ONNX`, fetched by asry's own `build.rs`). Raw
/// **logits**, 32-class head — the oracle's encoder.
///
/// `#[allow(dead_code)]`: only `tests/parity_words.rs` uses it.
/// The 32-class HuggingFace tokenizer matching [`asry_onnx_model_path`].
/// **Not** alignkit's bundled 29-class chordai asset: each tokenizer belongs
/// to its own CTC head, and asry's `Aligner::from_paths` validates the width.
///
/// `#[allow(dead_code)]`: only `tests/parity_words.rs` uses it.
/// Reads a 16 kHz mono 16-bit PCM WAV into normalized f32 samples.
///
/// Mirrors whisperkit's `tests/common::load_wav_mono_f32`.
/// Lowercase-hex SHA-256 digest of a file's contents.
///
/// Backs `tests/model_io.rs`'s provenance/integrity pin over the downloaded
/// model artifacts. `common` is a `mod`, not a separate crate, so each
/// `tests/*.rs` integration-test binary compiles its own copy; binaries
/// that don't happen to call this one (e.g. `tests/align_chunk.rs`)
/// would otherwise warn `dead_code` on it.
// ── Model-gate visibility (#61) ─────────────────────────────────────────────
//
// NOT `#[ignore]`d, deliberately. This is the ordinary-run half of the gate
// accounting: an ignored-ONLY run (`-- --ignored`, what every CI gate uses)
// never selects it, and it never appears in an ignored-only `--list`, so the
// anti-vacuum counts those gates take are unchanged. What it adds is the case
// no gate covers — a plain, modelless run — where the skipped gates otherwise
// say nothing but `ignored`. Mechanism, and what it does and does not refuse,
// in the shared module.
/// Reports how many of this binary's tests are `#[ignore]`d alignkit model gates
/// that did not run, and whether the models root they read is on disk.