1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
//! Ground-truth introspection of the FluidInference unified Silero VAD
//! artifact (design spec §4/§5). Every claim below comes from loading the real
//! `.mlmodelc` via `coremlit::Model::load` + `.description()`, or from actually
//! running it — the artifact's own `metadata.json` is a HYPOTHESIS re-verified
//! here, not trusted blind, and it wins over the plan wherever they differ.
//!
//! # Artifact (`Models/vadkit/`, COMMITTED — not a dev-time download)
//!
//! Source: <https://huggingface.co/FluidInference/silero-vad-coreml>, revision
//! (commit SHA) `b419383c55c110e2c9271fa6ee0ea83d03c70d96` — pinned at download
//! time (`hf api`/the HF API `sha` field). Artifact
//! `silero-vad-unified-256ms-v6.2.1.mlmodelc`.
//!
//! Unlike every other model this crate loads, these bytes are VENDORED into the
//! repository (1.1 MiB total, MIT; `.gitignore` un-ignores exactly this path,
//! and NOTICE sections 1-2 plus a LICENSE inside the artifact record the
//! redistribution). The gates below therefore need no fetch step and run in CI
//! on a fresh checkout. The `hf download` in `tests/vad/common/mod.rs`
//! re-fetches the same bytes from the Hub, which is how the committed copy is
//! verified against its source.
//!
//! | File | Role |
//! |---|---|
//! | `silero-vad-unified-256ms-v6.2.1.mlmodelc/metadata.json` | I/O contract |
//! | `silero-vad-unified-256ms-v6.2.1.mlmodelc/model.mil` | model graph |
//! | `silero-vad-unified-256ms-v6.2.1.mlmodelc/weights/weight.bin` | weights |
//! | `silero-vad-unified-256ms-v6.2.1.mlmodelc/coremldata.bin` | compiled model data |
//! | `silero-vad-unified-256ms-v6.2.1.mlmodelc/analytics/coremldata.bin` | analytics blob |
//!
//! Unlike alignkit, the targeted `.mlmodelc` came off the Hub pre-compiled
//! (v6.2.1 ships no `.mlpackage`), so every one of its files is byte-pinned
//! below by SHA-256 — there is no local `coremlcompiler` output whose bytes
//! could legitimately drift. Those pins now guard the COMMITTED copy too: they
//! are what would catch a checkout that rewrote the two TEXT files
//! (`model.mil`, `metadata.json`), which is why `.gitattributes` marks the whole
//! artifact `-text`. The `LICENSE` file inside the directory is coremlit's own
//! addition and is deliberately NOT pinned — it did not come from the Hub.
//!
//! # License
//!
//! HuggingFace `cardData.license` = `mit`. MIT end to end: upstream Silero VAD
//! is MIT, and FluidInference's CoreML conversion is MIT. MIT requires
//! preserving the notice, not a specific attribution string. Because the
//! artifact is now REDISTRIBUTED here rather than only loaded, that obligation
//! is live in this repository: NOTICE sections 1-2 carry the full upstream
//! notice and record both where FluidInference asserts MIT (card front matter,
//! HF tag, README) and that they ship no license file and assert no copyright
//! line; a second copy of the notice travels inside the artifact directory as
//! `LICENSE`. This module's record (repo id, revision, per-file SHA-256) is the
//! byte-level half of the same provenance.
//!
//! # Per-file SHA-256 (the artifacts as published on the Hub)
//!
//! | File | SHA-256 |
//! |---|---|
//! | `metadata.json` | `2740be542c611e1ba358e1849b4e265c65cdf0b17192767e1e5de86a31ac94d6` |
//! | `model.mil` | `c6a9d1bf22d413265da0a07a1d14151c3ea2fad296b3aa5859275b33ef1c3270` |
//! | `weights/weight.bin` | `53ecc8b5081146140ab654c89109cf001f2183abddd7a2411c5081feeffff063` |
//! | `coremldata.bin` | `7db35a4fd995222a7fb0129713473b15d1462572ab4a2e5e4d56bcaad9e40f41` |
//! | `analytics/coremldata.bin` | `8067594eb3126ab8318af507f0c00cabfed40d5fedb8a0ee5075dd02e903d909` |
//!
//! # DECISION
//!
//! - **Target: `silero-vad-unified-256ms-v6.2.1.mlmodelc`** — the exact version
//! FluidAudio pins (spec §5). The repo also ships `-v6.0.0`, an un-suffixed
//! `-256ms` sibling, `silero_vad*` and a 4-bit variant; only v6.2.1 is
//! targeted, downloaded, and pinned.
//!
//! # Spec-vs-reality
//!
//! 1. The plan expected "4160 in, one probability out" (spec §4, from
//! `VadManager.swift:21-26`). Introspection CONFIRMS `audio_input [1, 4160]`
//! f32 and a probability output, and reveals the artifact ALSO declares the
//! explicit LSTM state I/O `VadManager` drives:
//! `hidden_state`/`cell_state [1, 128]` f32 in → `new_hidden_state`/
//! `new_cell_state [1, 128]` f32 out. `stateSchema` is EMPTY — this is NOT a
//! CoreML `MLState` model; the recurrent state is ordinary feature I/O.
//! 2. `vad_output` is a rank-3 `[1, 1, 1]` f32 tensor, not a bare scalar `[1]`.
//! 3. v6.2.1 ships ONLY as a compiled `.mlmodelc` — there is no `.mlpackage`
//! for this version (the plan brief said "mlpackage + mlmodelc"), so no
//! `coremlcompiler` step is needed or possible.
//! 4. Spec §5 / plan T2 called for the revision to be "revision-pinned in
//! `MODELS_LOCK`". Reality: `MODELS_LOCK` is held by a whisperkit hermetic
//! gate (`coremlit/tests/whisper/models_lock.rs`) to EXACTLY the
//! tables CI actually downloads (vad is not among them), and the convention for
//! an adopted, CI-untested model (alignkit, speakerkit) is to pin its
//! revision + per-file SHA-256 in the crate's own `model_io.rs` — which is
//! where this record lives. Following that gated convention, not the plan's
//! letter, keeps the workspace green and consistent with the sibling kits.
//! Vendoring settled it for good: at 1.1 MiB the artifact is COMMITTED
//! instead, so CI never downloads it and a `MODELS_LOCK` table would buy
//! nothing — no lock entry, no cache key, no ci.yml parser change. The model
//! gates below are no longer CI-untested either; they run in the `check`
//! job against the committed bytes.
use ;
/// The model-layer contract, EXACT in both directions (design spec §4). Every
/// input AND output feature's name, shape and dtype is pinned against the live
/// model's introspected `description()` — the ground truth `VadModel::load_with`
/// validates at construction (mutating any pin here, or the matching check in
/// `crate::model::check_feature`, turns one side red).
/// Byte-pins every downloaded file of the artifact. A drift, corruption, or
/// silent re-download of DIFFERENT bytes fails here (the alignkit
/// `source_artifacts_match_pinned_sha256` precedent).
/// **Compute-placement characterization** (design spec §4 "ANE honesty" / §6).
/// Records what is actually measurable through this runtime, and asserts ONLY
/// that — never "runs on the ANE". `coremlit` exposes no `MLComputePlan`
/// placement introspection (only `ComputeUnits` SELECTION and I/O
/// `description()`), so the honest characterization is cross-placement
/// NUMERICAL agreement: run the same first chunk under every `ComputeUnits`
/// and record the resulting probability.
///
/// Measured (this test's own run, `02_pyannote_sample` chunk 0 from the
/// initial state, on the machine this branch was cut on):
/// - `cpu_only` : 0.083007812
/// - `cpu_and_gpu` : 0.083007812
/// - `cpu_and_neural_engine`: 0.083007812
/// - `all` : 0.083007812
/// - worst cross-placement |Δ| vs `cpu_only`: 0.000e0 (bit-identical)
///
/// All four placements agreed to the bit here — CoreML happened to schedule
/// this small graph identically. That is a MEASUREMENT, not a guarantee: the
/// ANE/GPU paths compute in fp16 (the graph is `Mixed(Float16, Float32)`) and
/// the fp32-capable `cpu_only` reference could diverge by up to one fp16 step
/// (~1e-3 near these values) on other hardware/OS, which is what the pinned
/// bound below allows for. Every placement must still return a finite
/// probability in `[0, 1]` (the noisy-OR output's range). `cpu_only` is the
/// deterministic reference the Swift trace gate (`parity_swift.rs`) and the
/// state gates pin against. This test asserts what it measures — agreement and
/// range — and never that any op runs on the ANE (`coremlit` cannot observe
/// that).
/// Worst tolerated cross-placement probability delta vs the `cpu_only`
/// reference — the fp16 headroom the ANE/GPU paths incur. **Measured worst:
/// `0.000e0`** (bit-identical across all four placements,
/// `compute_placement_is_characterized_not_asserted_ane`). Pinned at `1e-2` —
/// ~10x the fp16 output granularity (~1e-3 near these values) so genuine
/// cross-hardware fp16 drift does not flake, yet a full order of magnitude
/// below the O(1e-1) probability swing a real fp16 misplacement/corruption
/// produces (cf. the alignkit/speakerkit fp16 collapses), so it cannot mask one.
const PLACEMENT_AGREEMENT_TOL: f64 = 1e-2;