1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
658
659
660
661
662
663
664
665
666
667
668
669
670
671
672
673
674
675
676
677
678
679
680
681
682
683
684
685
686
687
688
689
690
691
692
693
694
695
696
697
698
699
700
701
702
703
704
705
706
707
708
709
710
711
712
713
714
715
716
717
718
719
720
721
722
723
724
725
726
727
728
729
730
731
732
733
734
735
736
737
738
739
740
741
742
743
744
745
746
747
748
749
750
751
752
753
754
755
756
757
758
759
760
761
762
763
764
765
766
767
768
769
770
771
772
773
774
775
776
777
778
779
780
781
782
783
784
785
786
787
788
789
790
791
792
793
794
795
796
797
798
799
800
801
802
803
804
805
806
807
808
809
810
811
812
813
814
815
816
817
818
819
820
821
822
823
824
825
826
827
828
829
830
831
832
833
834
835
836
837
838
839
840
841
842
843
844
845
846
847
848
849
850
851
852
853
854
855
856
857
858
859
860
861
862
863
864
865
866
867
868
869
870
871
872
873
874
875
876
877
878
879
880
881
882
883
884
885
886
887
888
889
890
891
892
893
894
895
896
897
898
899
900
901
902
903
904
905
906
907
908
909
910
911
//! Ground-truth introspection of every candidate segmentation and embedding
//! artifact named in the design spec
//! (`docs/superpowers/specs/2026-07-11-dia-coreml-backends-design.md` §4, §9
//! open item). Every value below comes from loading the real `.mlmodelc` via
//! `coremlit::Model::load` + `.description()` — the spec's table is a
//! HYPOTHESIS; reality wins, and every place it differs is marked `SPEC
//! DELTA`. Feeds Task 2 (`SegmentModel`) and later tasks (`EmbedModel`,
//! `Extractor`).
//!
//! # Artifacts (`Models/speakerkit/`, gitignored, fetched dev-time)
//!
//! | File | Role | Targeted? |
//! |---|---|---|
//! | `pyannote_segmentation.mlmodelc` | segmentation | **yes** — see DECISION |
//! | `Segmentation.mlmodelc` | segmentation, alt conversion | no |
//! | `wespeaker.mlmodelc` | embedding, fp32 | **yes** — see DECISION (issue #15) |
//! | `wespeaker_v2.mlmodelc` | embedding, int8 — RETIRED from shipping | no (tested sibling) |
//! | `wespeaker_int8.mlmodelc` | embedding, byte-identical to `wespeaker_v2` | no (same file) |
//! | `FBank.mlmodelc` | embedding frontend, split-pipeline alt | no |
//! | `Embedding.mlmodelc` | embedding backend, split-pipeline alt | no |
//!
//! The repo also ships `PLDA.mlmodelc` / `PldaRho.mlmodelc`. Neither is a
//! candidate and neither is introspected here: clustering — and with it the
//! PLDA projection — stays in `diaric` (spec §3 non-goal), which projects in
//! f64 on the host. `coremlit::audio::speaker::extract`'s `into_offline_input`
//! records why the CoreML graphs are unusable even if that ever changed.
//!
//! ## Retired coverage: `plda_io_recorded_out_of_scope` (a deliberate loss)
//!
//! `PLDA.mlmodelc` once had a row in the table above and an `#[ignore]`d
//! introspection test, `plda_io_recorded_out_of_scope`. Both were DELETED, not
//! relocated, and deleting the test did cost coverage that nothing else
//! provides. "No coverage was lost" would be false, so here is the ledger:
//!
//! - **What it uniquely asserted:** that `PLDA.mlmodelc` loads at all
//! (`Model::load`, `CpuOnly`), and that its description carries the input
//! `embeddings [1, 256]` F32 and the output `plda_features [1, 128]` F32.
//! - **What still touches the artifact, and how far that reaches:** only
//! `tests/fp16_guards.rs`. Its sweep requires the bundle to be present with a
//! parseable `model.mil` and hard-fails if the pinned guard sites change or
//! vanish — so deletion of the artifact and drift in its numerical guards are
//! still caught. But it reads the MIL as TEXT: it never calls `Model::load`,
//! so it says nothing about CoreML loadability, and it never reads the model
//! description, so it says nothing about feature names, shapes or dtypes.
//! - **So the loadability and schema assertions are simply gone**, and nothing
//! replaces them. A re-conversion that reshaped the I/O while leaving the
//! guard sites alone would now pass unnoticed.
//! - **Why that is the intended trade rather than an oversight:** the artifact
//! is on no runtime path — nothing in this crate calls `Model::load` on it,
//! and nothing is planned to, because the projection is
//! `diaric::plda::PldaTransform` in f64 on the host. The test pinned the I/O
//! contract of a graph this crate will never consume, and being `#[ignore]`d
//! it ran only when someone staged the models and asked for it by name.
//! Retiring it with this record beats keeping a check whose presence implies
//! the graph is still a live candidate; the one scenario in which the lost
//! assertions would matter — the graph becoming a candidate again — is also
//! the one scenario in which it would be re-introspected from scratch anyway.
//! - **If it ever does become one**, restore exactly the assertions listed
//! above; they are spelled out here so the restoration needs no `git log`
//! archaeology.
//!
//! # Artifact provenance (issue #15)
//!
//! Two artifacts — the two the shipping pipeline loads — come from
//! <https://huggingface.co/FinDIT-Studio/speakerkit-coreml>, revision (commit
//! SHA) `3db69988bf2de12bab250614d6ac2b03d35132a2`, and every file of each is
//! byte-pinned in this suite:
//!
//! - `pyannote_segmentation.mlmodelc`
//! ([`fp16_safe_segmentation_matches_pinned_sha256`]);
//! - `wespeaker.mlmodelc` ([`fp16_safe_wespeaker_fp32_matches_pinned_sha256`]).
//!
//! ```text
//! hf download FinDIT-Studio/speakerkit-coreml \
//! --revision 3db69988bf2de12bab250614d6ac2b03d35132a2 \
//! --include 'pyannote_segmentation.mlmodelc/*' 'wespeaker.mlmodelc/*' \
//! --local-dir Models/speakerkit
//! ```
//!
//! Every other artifact in `Models/speakerkit/` is FluidInference's
//! `speaker-diarization-coreml` conversion at revision
//! `1ed7a662fdc7109e36d822db793ee6eebdaf8594`, verified 2026-07-27 against the
//! HuggingFace API (LFS sha256 for `weights/weight.bin`, git blob SHA-1 for
//! `model.mil`) and byte-pinned for the tested int8 sibling by
//! [`int8_wespeaker_matches_fluidinference_pinned_sha256`]:
//!
//! ```text
//! hf download FluidInference/speaker-diarization-coreml \
//! --revision 1ed7a662fdc7109e36d822db793ee6eebdaf8594 \
//! --local-dir Models/speakerkit
//! ```
//!
//! (Fetch FluidInference first, then overlay the two FinDIT-Studio artifacts —
//! the second download replaces `pyannote_segmentation.mlmodelc` and
//! `wespeaker.mlmodelc` in place. A tree that skips the overlay fails the two
//! byte-pin tests with re-download instructions rather than silently running
//! the pre-repair artifacts.)
//!
//! Both FinDIT-Studio artifacts are **re-conversions of the same upstream
//! weights**, not different models:
//!
//! - `pyannote_segmentation.mlmodelc`: FluidInference's fp16 conversion ended
//! `softmax` → `log(epsilon = 0x0p+0)`; `0` is below fp16's smallest
//! subnormal (`2^-24`), so wherever the graph computes in fp16 the guard is
//! inert and an underflowed softmax reaches an unguarded `log(0)`. Measured
//! on `09_mrbeast_dollar_date`, 1033 chunks, the shipping
//! `ComputeUnits::All` placement: minimum `segments` value **−45440.0**
//! against **−32.31** on `CpuOnly`. The re-conversion emits the fused
//! `reduce_log_sum_exp` → `sub` form — no `log` op at all, nothing to
//! saturate — and the same measurement gives **−31.80** on `All`.
//! - `wespeaker.mlmodelc`: byte-identical weights (`weights/weight.bin`
//! sha256 `680837ec…` on both sides — the Phase-A repair was MIL-only) and a
//! `model.mil` differing ONLY in the two attentive-stat pooling guard
//! constants (`1e-8` → `0x1p-24`, the fp16 floor) plus coremlc buildInfo
//! strings. On this host the two MILs measured EQUAL to the last DER error
//! unit on every clip-09 remedy-matrix arm — the guard sites add their
//! epsilon to mask sums that are never near zero — so the repair is a
//! static fp16-floor guarantee (`tests/fp16_guards.rs`'s defect class), not
//! a measured quality change. Adopting it removes the shipping embedder
//! from the sub-fp16-guard defect roster.
//!
//! **What the SEGMENTATION swap did NOT do — an OBSERVATION from the int8
//! era, kept for attribution.** Re-running the then-current four gated clips
//! (06 / 14 / 10 / 09, int8 shipping arms) with only the segmentation
//! artifact changed reproduced every gated number, clip 09's 5-of-8-speaker
//! collapse included. Environment: Apple M1 Max, macOS 26.5 (build 25F71),
//! arm64; the reference side dia's ONNX on `ort`'s CPU EP. That measurement
//! is why the collapse was never the segmentation tail's to fix: it removed a
//! real, silent, four-orders-of-magnitude corruption on the default placement
//! and moved no gated DER. The collapse belonged to the embedder — see
//! "Clip 09" below — and its repair is the EMBEDDING decision, not this
//! swap.
//!
//! The published `wespeaker_int8` re-conversion stays NOT adopted: it is also
//! a RE-PALETTIZATION (different LUTs), and it moves clip 14's int8 ANE arm
//! from 0.8178 % to 1.4860 % DER (isolated by swapping one artifact at a
//! time; see issue #15). Its `fp16_guards` pin therefore stands, and the
//! retired-but-tested int8 sibling keeps FluidInference's original bytes
//! ([`int8_wespeaker_matches_fluidinference_pinned_sha256`]). The published
//! fp32 `wespeaker` re-conversion IS adopted — the clip-14 regression
//! belongs to the re-palettized LUTs, which the fp32 artifact does not have,
//! and its remedy-matrix arms are measured DER-equal to FluidInference's
//! fp32 on this host (see "Artifact provenance" above).
//!
//! # Clip 09: what the cross-products establish, and at WHICH configuration
//!
//! Two model cross-products exist. They were run at different configurations,
//! **they disagree**, and neither speaks for the other. Both hold one audio
//! buffer, one chunk grid, one decode and one clustering constant, varying
//! only which conversion computes a stage.
//!
//! **(a) fp32 embedder (`wespeaker.mlmodelc`) on `CpuOnly`** — the original
//! run (issue #15). In that configuration the CoreML path does not undercount;
//! it produces no answer at all:
//!
//! ```text
//! ONNX-seg + ONNX-emb 8 spk, 0.0000 % the dia-ort reference
//! COREML-seg + ONNX-emb Err(AmbiguousAliveCluster sp[13] = 1.706e-7)
//! ONNX-seg + COREML-emb 8 spk, 0.0000 %
//! COREML-seg + COREML-emb Err(AmbiguousAliveCluster sp[13] = 1.700e-7)
//! ```
//!
//! - *Established*: at fp32/`CpuOnly`, the segmentation conversion alone is
//! SUFFICIENT to produce that `Err`, and the fp32 embedder alone is not.
//! - *Not established by (a)*: anything about the int8 embedder or about
//! `ComputeUnits::All` — a different artifact, different kernels, and not
//! even the same failure (a refusal to cluster, versus the silent 5-of-8
//! undercount this crate shipped). The sentence this doc carried until
//! 2026-07-26, "the embedder is exonerated", generalized (a) to a
//! configuration it never touched.
//!
//! **(b) int8 embedder (`wespeaker_v2.mlmodelc`) on `ComputeUnits::All` — the
//! configuration that SHIPPED until issue #15 retired it.** `coremlit-parity`'s `tests/speaker/backend_factorial.rs` runs the
//! identical design where the defect actually lives, same clip, same host:
//!
//! ```text
//! segmentation | embedding | spk | DER | conf
//! -------------+-----------+-----+----------+---------
//! ONNX | ONNX | 8 | 0.0000% | 0.0000%
//! ONNX | COREML | 5 | 16.5904% | 16.5904%
//! COREML | ONNX | 9 | 1.3011% | 1.3011%
//! COREML | COREML | 5 | 16.5904% | 16.5904%
//! ```
//!
//! - *Established at the shipping configuration*: swapping ONLY the
//! **embedding** conversion, over dia's own reference segmentation,
//! reproduces the shipping collapse exactly — 5 of 8 speakers, 16.5904 %
//! DER, 11 999 confusion units, the same numbers as the all-CoreML corner.
//! Swapping ONLY the **segmentation** conversion does not: it OVERcounts by
//! one (9 speakers, 1.3011 %), a real defect an order of magnitude smaller
//! that the shipping arm masks. Both corners reproduce their independently
//! pinned numbers (dia-ort's 8 / 0.0000 %; `parity_shipping_der`'s
//! 5 / 16.5904 %), which is what makes the hybrid cells readable at all.
//! - *Not established by (b)*: WHICH property of the CoreML embedding path
//! carries it. The factor varied is the BACKEND, so the implicated object is
//! the then-shipping bundle — int8 palettization **plus** `All` placement **plus**
//! that conversion — as one unit. That is what (c) separates.
//!
//! **(c) The embedding path's three properties, separated.**
//! `backend_factorial.rs`'s `embedding_precision_x_placement` holds dia's
//! reference segmentation fixed for every arm and runs the embedding arm across
//! precision x placement, same clip and host:
//!
//! ```text
//! cell | embedding arm | spk | DER | conf | err units
//! -----+--------------------------+-----+----------+----------+----------
//! A | ONNX fp32 / ort CPU EP | 8 | 0.0000% | 0.0000% | 0
//! B | CoreML int8 / All | 5 | 16.5904% | 16.5904% | 11999
//! C | CoreML fp32 / All | 7 | 2.5427% | 2.5427% | 1839
//! D | CoreML int8 / CpuOnly | 6 | 16.3636% | 16.3636% | 11835
//! E | CoreML fp32 / CpuOnly | 8 | 0.0000% | 0.0000% | 0
//! ```
//!
//! - *Established*: **the embedding CONVERSION is exonerated.** Cell E — the
//! same CoreML graph with both other factors removed — reproduces dia-ort
//! frame-perfectly: 8 of 8 speakers and `err_units == 0`, not one
//! collar-scored speaker-frame different, at mean AND minimum cosine
//! 1.000000 against dia's fp32 ONNX over all 2 114 `(chunk, slot)` rows.
//! *(b)*'s "the CoreML embedding path" must not be read as "the CoreML
//! embedding conversion".
//! - *Established*: the two remaining factors are separable, additive on
//! speaker count, and very unequal on error mass. **int8 palettization costs
//! 2 speakers at BOTH placements** (E 8 -> D 6; C 7 -> B 5), moving 11 835
//! of the shipping arm's 11 999 error units — 98.6 % — in the measured run.
//! **The `All` placement costs 1 speaker at BOTH precisions** (E 8 -> C 7;
//! D 6 -> B 5), 1 839 / 164 error units respectively in that run. The
//! speaker counts and per-cell units are guarded
//! (`backend_factorial`'s verdicts, cells to ±10 units); the derived
//! splits (164, 98.6 %) are REPORTED measurements of that run. Neither
//! factor alone reproduces the shipping 5-of-8 collapse.
//! - *Not established by (c) alone*: the arms run over ONNX segmentation,
//! which is not what this crate ships, and nothing here extends beyond clip
//! 09 or this host. (d) below is the real-pipeline pricing that settled the
//! remedy.
//!
//! **(d) The mechanism, and the remedy priced in the REAL pipeline.**
//! `backend_factorial.rs`'s `quantization_error_structure` decomposes each
//! arm's embedding perturbation against cell E and projects it through
//! diaric's own frozen community-1 transform: the int8 delta is a COHERENT
//! shared displacement (coherence 0.50 vs the 0.022 isotropic null, 98.7 % of
//! rows aligned) that compresses between-speaker centroid margins by up to
//! +0.05 cosine, concentrated on pairs involving the clusters that then lose
//! their identity (the probe's printed contingency names them), with
//! within-cluster tightness unchanged; the `All` placement's delta is 1.5x
//! larger per row
//! but near-isotropic (coherence 0.06). Perturbation SIZE anti-correlates
//! with damage; the coherent component predicts it. The full DECISION section
//! carries the artifact-level cause (38 per-tensor 256-entry LUTs) and the
//! remedy-matrix table (real pipeline: `seg@All + fp32@All` = 8/8 at
//! 2.9810 %; the CPU-embedder arms sit on the segmentation knife edge — 9
//! speakers / `Err(AmbiguousAliveCluster)`); `parity_shipping_der.rs` gates
//! the adopted configuration per clip and pins the clip-09 record.
//!
//! The mechanism INSIDE the segmentation graph is a further step again, and it
//! is not established either: `segments` is the only tensor either graph
//! exposes, so a divergence in it cannot be attributed to the log-softmax tail
//! rather than to the trunk feeding it. Measured on this clip at `All`, the
//! CoreML segmentation differs from dia's ONNX by at most 0.6574 in
//! log-probability and flips 565 of 608 437 powerset argmax frames (0.0929 %);
//! of those flips only 50 carry an exact tie at the CoreML row maximum, and a
//! tie is the ONLY argmax change a monotone per-row shift can produce — so at
//! least 515 of them originate upstream of the tail. `backend_factorial`'s
//! `seg_divergence` carries the full argument and its caveats. On clip 09
//! that trunk divergence surfaces as the shipping suite's three
//! separately-pinned placement outcomes (9 spk / `Err` / 8-with-confusion);
//! reading them as one near-threshold cluster whose state flips with
//! placement is the interpretation the pattern supports — the pinned record
//! (`assert_clip09_record`) states the distinction.
//!
//! # Licenses (`Models/speakerkit/README.md`)
//!
//! The repo's HuggingFace frontmatter declares `license: cc-by-4.0` for the
//! model repo as a whole; the body clarifies "the SDK itself is Apache 2.0,
//! but the parent model from Pyannote is `cc-by-4.0`" ("SDK" = FluidAudio's
//! conversion tooling, not the weights this crate loads). The newer
//! "community-1" conversion set (`Segmentation`/`FBank`/`Embedding`/`PLDA`,
//! see DECISION below) additionally self-declares `"license": "CC-BY-4.0"`
//! inside its own `metadata.json`, confirming the same terms independently.
//! CC-BY-4.0 requires attribution; the README's Citations section gives the
//! required BibTeX: segmentation model (Plaquet & Bredin, "Powerset
//! multi-class cross entropy loss for neural speaker diarization",
//! INTERSPEECH 2023), speaker embedding model (Wang et al., "Wespeaker: A
//! research and production oriented speaker embedding learning toolkit",
//! ICASSP 2023), and speaker clustering / VBx (Landini et al., "Bayesian
//! HMM clustering of x-vector sequences (VBx) in speaker diarization",
//! Computer Speech & Language 2022) — the last is irrelevant to speakerkit
//! (clustering stays in `diaric`, spec §3) but ships in the same README and is
//! reproduced here for completeness.
//!
//! # DECISION
//!
//! - **Segmentation: `pyannote_segmentation.mlmodelc`.** The two candidates
//! are NOT contract-equal — the plan brief's stated tiebreaker condition
//! ("pick pyannote_segmentation if contract-equal") does not actually
//! hold, see `segmentation_alt_io_recorded_not_targeted` below. It is
//! chosen anyway because its single-chunk, fixed-shape `segments` output
//! (per-frame powerset **log-probabilities** — the graph's tail is
//! `reduce_log_sum_exp` → `sub`, see `crate::segment`'s module doc) matches both the
//! spec's pinned contract (§4 table) and the `SegmentModel::infer`
//! single-chunk API (§5) exactly, and it is FluidAudio's shipping name —
//! the brief's fallback tiebreaker.
//! - **Embedding: `wespeaker.mlmodelc` (fp32, the FinDIT-Studio fp16-safe
//! MIL) — issue #15.** The original DECISION here was the int8-palettized
//! `wespeaker_v2.mlmodelc` (byte-identical to `wespeaker_int8.mlmodelc`;
//! see `wespeaker_v2_and_wespeaker_int8_are_byte_identical` below). It was
//! retired on measurement:
//!
//! *The collapse.* On 8-speaker audio the int8 embedder silently loses
//! speakers: `09_mrbeast_dollar_date`, 5 of 8 speakers at 16.5904 % DER
//! (100 % confusion) at the then-shipping `int8/All`, where dia-ort is
//! frame-perfect. `backend_factorial.rs` isolated the factors over dia's
//! reference segmentation: the CoreML embedding CONVERSION is exonerated
//! (fp32/`CpuOnly` reproduces dia-ort frame-perfectly, mean AND minimum
//! cosine 1.000000 over all 2 114 rows), the int8 palettization costs 2
//! speakers at either placement (98.6 % of the shipping arm's error mass —
//! a reported figure of the measured run, cells guarded to ±10 units),
//! and the `All` placement costs 1 more at either precision.
//!
//! *The mechanism* (`backend_factorial.rs`'s
//! `quantization_error_structure`, same clip/host): the palettization
//! error is NOT isotropic noise — it is a COHERENT shared displacement.
//! Against the fp32/`CpuOnly` base, the int8 arm's per-row delta carries
//! half its mass in ONE shared direction (coherence 0.50 vs the 0.022
//! independent-scatter null; 98.7 % of rows aligned), producing a
//! near-constant ~2.4 %-of-norm shift on every embedding. Through diaric's
//! frozen community-1 projection (center on a FROZEN mean → L2-normalize →
//! LDA → re-center → re-normalize → PLDA-whiten) that shared shift
//! compresses BETWEEN-speaker centroid margins by up to +0.05 cosine,
//! concentrated on pairs involving the clusters that then lose their
//! identity (the probe's printed contingency names them), while
//! within-cluster tightness is unchanged to three decimals. The `All` placement's fp16 scatter is the
//! opposite shape — 1.5× LARGER per row but near-isotropic (coherence
//! 0.06) — which is why the perturbation SIZE anti-correlates with the
//! damage. The earlier "quantization is roughly isotropic and survives the
//! frozen basis" rationale in `parity_shipping_der` was refuted by this
//! measurement: palettization is the coherent, basis-hostile perturbation;
//! the placement is the noisy benign one. The structural reason is in the
//! artifact: 38 `constexpr_lut_to_dense` sites, each ONE flat 256-entry
//! fp32 codebook for the WHOLE tensor (~0.8-1.0 % rel-RMS weight error per
//! ResNet conv, compounding over 34 layers), covering even the
//! deterministic DSP constants (STFT cos/sin bases, mel filterbank) and
//! the 5120→256 embedding head — the same ΔW applied to every input's
//! shared activation statistics is a constant output-space bias. A
//! repaired quantization must break that coherence (per-channel/grouped
//! LUTs, higher bits, exempting the head and DSP constants) and re-enter
//! the full DER validation; the one published re-palettization
//! (`wespeaker_int8` at the FinDIT revision) is measured to REGRESS clip
//! 14's ANE arm from 0.8178 % to 1.4860 % and stays rejected.
//!
//! *The remedy pricing* (issue #15 remedy matrix, real pipeline — CoreML
//! fp16-safe segmentation + each embedder arm, Apple M1 Max, macOS 26.5
//! build 25F71, arm64, release harness): on clip 09, `seg@All +
//! fp32@All` = **8 of 8 speakers at 2.9810 %** (the only composition with
//! the right count); `seg@All + fp32@CpuOnly` = 9 speakers at 1.3011 %
//! (a spurious cluster survives against a bit-near-ONNX embedder — the
//! segmentation conversion's overcount class); `seg@CpuOnly +
//! fp32@CpuOnly` = `Err(AmbiguousAliveCluster)` (an alive-band refusal;
//! one shared near-threshold cluster is the supported interpretation, not
//! an asserted identity — the pinned record states the distinction); the
//! retired `seg@All + int8@All` = 5 at 16.5904 %. Speed does not
//! adjudicate the artifact choice: two warm runs of THIS bench
//! (`shipping_embedder_cost_int8_vs_fp32`, 120 s of clip 10, one config at
//! a time, this host, post-swap artifacts) put the int8-vs-fp32 extraction
//! difference ≤ ~15 % on every placement WITH THE SIGN FLIPPING BETWEEN
//! RUNS (`All`: fp32 4.33 s vs int8 4.95 s in the first run, then int8
//! 4.13 s vs fp32 4.65 s in the second; the CPU rows flipped likewise) —
//! inside warm-run scheduler variability, so neither artifact holds a
//! stable edge and the bench prints rather than asserts a winner. The
//! stable cost axis is PLACEMENT (CPU-embedder configs run ~2x slower than
//! `All`, both artifacts, both runs). Palettization's remaining edge is
//! ~21 MB of footprint (8.0 MB vs 29.4 MB) — retired as the price of not
//! silently losing 3 speakers.
//!
//! *The artifact choice within fp32*: the FinDIT-Studio fp16-safe MIL
//! (pooling eps `1e-8` → `0x1p-24`) over FluidInference's original —
//! measured EQUAL to the last DER error unit on every clip-09 arm on this
//! host, adopted for the static fp16-floor guarantee (see "Artifact
//! provenance" above and `tests/fp16_guards.rs`).
//!
//! The remaining gate evidence lives in `parity_shipping_der.rs` (per-clip
//! placement matrix + clip-09 record); a parity gate (spec §6.2)
//! separately confirms the fp32 conversion carries no NaN/Inf corruption
//! (spec §1).
//!
//! `FBank.mlmodelc` + `Embedding.mlmodelc` (the split fbank-then-embed
//! pipeline) are NOT targeted per spec §2.4: the wespeaker artifacts
//! compute fbank in-graph from raw waveform, so the split frontend is
//! unnecessary.
//!
//! # Spec-vs-reality deltas
//!
//! 1. The segmentation candidates are NOT contract-equal (see DECISION).
//! `Segmentation.mlmodelc` is part of a distinct, newer "community-1"
//! conversion set (`coremltools` 9.0b1/`torch` 2.8.0, converted
//! 2025-10-13, minimum macOS 14) vs `pyannote_segmentation.mlmodelc`
//! (`coremltools` 8.3.0/`torch` 2.6.0, minimum macOS 12): it batches
//! 1..=32 chunks per call (default shape `[32, 1, 160000]`) and its sole
//! output is named `log_probs` (log-softmaxed, per its `metadata.json`)
//! with a shape CoreML leaves unpinned — not `segments`, raw powerset
//! logits, fixed `[1, 589, 7]`. Only `pyannote_segmentation`'s contract
//! matches the spec's table, which introspection confirms exactly:
//! `audio [1, 1, 160000]` f32 -> `segments [1, 589, 7]` f32.
//! 2. `wespeaker_v2.mlmodelc` (and its `wespeaker`/`wespeaker_int8`
//! siblings) carry an undocumented second output, `constant`:
//! fixed-shape (rank-0/scalar, NOT a symptom of input flexibility —
//! `hasShapeFlexibility` is false in `metadata.json`) `Some(F32)`. Not
//! in the spec; Task 2 ignores it and reads `embedding` only.
//! 3. `wespeaker_v2.mlmodelc` and `wespeaker_int8.mlmodelc` are the same
//! file (see DECISION) — the spec's table names only `wespeaker_v2` and
//! doesn't mention this duplication.
//! 4. Every output whose shape depends on a flexible input
//! (`Segmentation`'s `log_probs`, `FBank`'s `fbank_features`,
//! `Embedding`'s `embedding`) introspects to an EMPTY shape (`[]`) with
//! the dtype still populated — `coremlit`'s `FeatureInfo` reports a real
//! `multiArrayConstraint` (so `data_type()` resolves) but CoreML
//! declares no static shape for it. None of these are targeted
//! artifacts, so this doesn't block Task 2, but it is the shape a future
//! flexible-batch design would need to handle explicitly (a predict-time
//! concern, not a load-time one).
//! 5. Flexible-shape INPUTS (`Segmentation`'s `audio`, `FBank`'s `audio`,
//! `Embedding`'s `fbank_features`/`weights`) introspect to their
//! declared DEFAULT shape, not an empty/unconstrained one:
//! `Segmentation`'s default is `[32, 1, 160000]` (the max of its 1..=32
//! enumerated range), `FBank`'s default is `[1, 1, 160000]` (batch 1),
//! `Embedding`'s are the low end of its range constraints (`[1, 1, 80,
//! 998]`, `[1, 589]`). None of the four targeted-artifact inputs are
//! flexible, so this doesn't affect Task 2 either.
use ;
use ;
/// Recursively collects every FILE under `dir` as a `/`-separated path relative
/// to `root`. Used by `wespeaker_v2_and_wespeaker_int8_are_byte_identical` to
/// compare two `.mlmodelc` bundle trees file-for-file.
///
/// OS-generated sidecars are skipped: AppleDouble `._*` files and `.DS_Store`.
/// macOS materializes these inside bundles on non-native filesystems
/// (exFAT/FAT/SMB); CoreML's loader never reads them, so excluding them from
/// discovery cannot mask a functional artifact change — whereas NOT excluding
/// them would false-fail the byte-identity comparison below as a phantom
/// "unpinned extra" even though every real byte is untouched.
/// Hermetic non-vacuity proof for [`collect_files_rel`]'s sidecar filter (no
/// staged model needed). On exFAT/FAT/SMB volumes macOS materializes
/// AppleDouble `._*` and `.DS_Store` sidecars inside `.mlmodelc` bundles;
/// discovery must drop EXACTLY those, while every real file — crucially
/// including an unpinned real extra — still reaches the byte-identity gate in
/// [`wespeaker_v2_and_wespeaker_int8_are_byte_identical`]. This proves the
/// filter fixes the false-failure WITHOUT blanket-suppressing genuine extras.
/// Byte-pins every file of the fp16-safe segmentation artifact against the
/// published `FinDIT-Studio/speakerkit-coreml` revision (module doc,
/// "Segmentation provenance").
///
/// This is the gate that makes the issue-#15 swap non-reversible by accident.
/// The whole difference between the two conversions lives inside `model.mil`:
/// the pre-swap FluidInference artifact has the identical filename, the
/// identical `audio [1,1,160000]` → `segments [1,589,7]` contract, and passes
/// every other test in this file — while restoring a `segments` minimum of
/// −45440 on the shipping `ComputeUnits::All` placement. Only these bytes carry
/// the repair, so only a byte pin can defend it.
/// Byte-pins every file of the fp16-safe fp32 embedder against the published
/// `FinDIT-Studio/speakerkit-coreml` revision (module doc, "Artifact
/// provenance") — the shipping-embedder twin of
/// [`fp16_safe_segmentation_matches_pinned_sha256`], and the issue-#15 gate
/// that makes the int8 retirement non-reversible by accident.
///
/// The whole difference between this artifact and FluidInference's original
/// fp32 conversion lives inside `model.mil` (two pooling guard constants and
/// buildInfo strings; the weights are byte-identical), and the two are
/// measured DER-equal on this host — so nothing but a byte pin can tell them
/// apart, and only these bytes carry the static fp16-floor repair that keeps
/// `tests/fp16_guards.rs`'s roster clean for the shipping embedder.
/// Byte-pins the retired int8 sibling (`wespeaker_v2.mlmodelc`) against
/// FluidInference's `speaker-diarization-coreml` at the pinned revision
/// (module doc, "Artifact provenance").
///
/// The artifact no longer ships, but the issue-#15 record RUNS on it:
/// `backend_factorial.rs`'s B/D cells and the `quantization_error_structure`
/// mechanism probe reproduce the collapse from exactly these bytes. A silent
/// swap (for instance to the published re-palettization, whose different LUTs
/// regress clip 14) would quietly re-baseline that whole record, so the bytes
/// are pinned like the shipping artifacts'.
/// The shipped segmentation graph reaches its powerset log-probabilities through
/// the FUSED `reduce_log_sum_exp` → `sub` tail, and contains no `log` op that a
/// vanishing epsilon could leave unguarded. Reading the MIL is what the whole
/// issue-#15 investigation turned on — the docs asserted "raw powerset logits"
/// and nobody went looking for the `log` — so the structural claim is asserted,
/// not narrated. `tests/fp16_guards.rs` checks epsilon FLOORS across every
/// vendor; this checks the one structural property that makes this graph
/// epsilon-free in the first place.