coremlit 0.1.2

Safe, synchronous CoreML runtime for macOS (CPU/GPU/Neural Engine) with opt-in on-device multimodal pipelines: speech (Whisper STT, forced alignment, speaker diarization, Silero VAD), AudioSet sound-event tagging, and audio/text/image embeddings (CLAP, granite, SigLIP)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
658
659
660
661
662
663
664
665
666
667
668
669
670
671
672
673
674
675
676
677
678
679
680
681
682
683
684
685
686
687
688
689
690
691
692
693
694
695
696
697
698
699
700
701
702
703
704
705
706
707
708
709
710
711
712
713
714
715
716
717
718
719
720
721
722
723
724
725
726
727
728
729
730
731
732
733
734
735
736
737
738
739
740
741
742
743
744
745
746
747
748
749
750
751
752
753
754
755
756
757
758
759
760
761
762
763
764
765
766
767
768
769
770
771
772
773
774
775
776
777
778
779
780
781
782
783
784
785
786
787
788
789
790
791
792
793
794
795
796
797
798
799
800
801
802
803
804
805
806
807
808
809
810
811
812
813
814
815
816
817
818
819
820
821
822
823
824
825
826
827
828
829
830
831
832
833
834
835
836
837
838
839
//! Native CoreML **CED** (tiny/mini/small/base) AudioSet sound-event tagging —
//! coremlit's first multi-label classifier: 16 kHz mono waveform in, ranked
//! AudioSet predictions out (527 rated classes: name + permanent `SoundEventId`
//! + `/m/…` mid + class index + sigmoid confidence), long clips via windowed
//! chunking + Mean/Max aggregation.
//!
//! CED (Consistent Ensemble Distillation, arXiv 2308.11957; upstream
//! RicherMans/CED, `mispeech/ced-{tiny,mini,small,base}`) is a distilled
//! AudioSet transformer. The four sizes are contract-identical here — one
//! size-invariant mel→logits I/O; they differ only in internal transformer
//! width (see [`CedModel`]). The mel front-end runs in Rust (the private `mel`
//! submodule) and the mel→logits transformer runs natively on Apple silicon as
//! one fp16 `.mlmodelc` — an in-graph STFT/mel is the exact fragility class
//! behind the ORT CoreML EP zeroed-logits bug this feature closes. NO `ort`
//! anywhere.
//!
//! Design spec: `docs/superpowers/specs/2026-07-23-ced-native-ane-design.md`.
//!
//! # Model artifacts
//!
//! No model is bundled (a `.mlmodelc` is a directory artifact). Each size's
//! fp16 CED graph is converted owner-side (Wave B), distributed via Hugging
//! Face, and staged as a gitignored dev-time download under
//! `Models/ced/ced-<size>/` (env override `CED_TEST_MODELS` points at the
//! `Models/ced` family root); per-file SHA-256 and I/O contract are pinned per
//! size by `tests/ced/model_io.rs` once staged. [`CedModel`] owns the repo ids
//! and the `<dir>/<bundle>` path spelling ([`CedModel::mlmodelc_path`]). The
//! four graphs are I/O-identical, so model identity is caller-supplied:
//! coremlit cannot — and does not — detect which size a `.mlmodelc` is.
//!
//! # Rust front-end around an fp16 CoreML graph
//!
//! The graph takes the believed `[1, 64, 1001]` log-mel (`mel`, f32) computed
//! by this module's Rust front-end and emits `[1, 527]` **pre-sigmoid** logits
//! (`logits`, f32); sigmoid, ranking, and long-clip aggregation run in Rust.
//! The believed mel numerics are probe-pinned in Wave B (see the `mel`
//! submodule docs). This `[1, 64, 1001]` shape is shared by all four sizes.
//!
//! Upstream's `target_length = 1012` is NOT this input width: it is the
//! transformer's time positional-embedding capacity and its long-form mel
//! chunk size. A canonical 10 s window is 160 000 samples → 1001 mel frames
//! (hop 160, `center=True`), consumed unpadded with the pos embed sliced to 62
//! of its 63 patch columns; padding to 1012 would compute a different
//! function. So `1001 <= 1012` is on-distribution, not a truncation (verified
//! against RicherMans/CED `audiotransformer.py` and the mispeech feature
//! extractor); the `mel` submodule carries the full derivation.
//!
//! # From window scores to events
//!
//! [`Classifier::classify_windows`] is the seam between this crate and event
//! detection. It returns `Vec<`[`WindowConfidences`]`>`, and
//! [`WindowConfidences`] *is* [`windit::windowed::Windowed`]`<`[`Confidences`]`>`
//! — just as [`Span`] *is* [`windit::plan::Span`]. coremlit's long-clip output
//! is already `windit`'s own value type, so the post-processing stack composes
//! with no adapter and no repacking:
//!
//! ```text
//! Classifier::classify_windows -> Vec<Windowed<Confidences>>  coremlit (this module)
//!   index one class            -> Vec<Windowed<f32>>          one slice read per window
//!   windit::smooth             -> Vec<Windowed<f32>>          Ema / CadenceEma
//!   zuoer::RunSegmenter        -> Run { span, mean, peak }    hysteresis + durations
//! ```
//!
//! ## coremlit ships no CED convenience layer, deliberately
//!
//! `audio::vad` offers a one-call `detect_speech` because Silero has
//! *upstream-authored* defaults — 0.5 threshold, 250 ms minimum speech, 100 ms
//! minimum silence — that `zuoer`'s hysteresis derives from, so "the default"
//! there is a real, attributable thing. CED's 527 classes have no equivalent. A
//! threshold and minimum duration right for `Glass` (class 441, a sub-second
//! transient) are wrong for `Music` (class 137, continuous for minutes); a hop
//! fine enough to localize a single bark wastes an order of magnitude of
//! inference on ambient scene tagging. Any defaults coremlit shipped would be
//! silently wrong in some scenario, and silently wrong is the failure mode a
//! classifier can least afford.
//!
//! So the glue below stays in your application. It is about a dozen lines, and
//! **every parameter in it is scenario-dependent** — which is precisely why it
//! is not a coremlit function. This module chooses no threshold, no smoothing
//! constant, no minimum event duration, and no set of classes to watch.
//!
//! ## Per-class orchestration is the consumer's job
//!
//! CED emits 527 *independent* sigmoids per window: not a softmax, and they do
//! not sum to one. An event detector picks the handful of classes it cares
//! about and runs one independent smoother + segmenter per class, each with its
//! own parameters. Running all 527 is possible but rarely wanted — a real
//! recording has a sparse active set at any moment, and 527 segmenters is 527
//! parameter sets nobody has tuned.
//!
//! ## Timestamps are window-resolution, not event-resolution
//!
//! Each score summarizes a whole [`WINDOW_SAMPLES`] (10 s) window, and a
//! segmenter treats it as one point sample placed at that window's start. Run
//! boundaries are therefore quantized to [`WindowPlan::hop_samples`], and the
//! audio a run actually observed is its reported interval extended forward by
//! one whole window — so an event reported at `3 s..8 s` happened somewhere in
//! `3 s..18 s`. A shorter hop buys finer quantization, never a narrower smear.
//!
//! ## Dependencies
//!
//! `windit` and `zuoer` are coremlit's own dependencies. Only [`Span`] and
//! [`WindowConfidences`] cross into *this* module's API, and this module
//! re-exports no smoothing tier, so depend on `windit` directly. (Under the
//! `clap` feature, `embeddings::clap::smooth` does re-export windit's smoothing
//! seam — but only the parts a 512-wide *embedding* can use: the scalar `Ema`
//! and `CadenceEma` this table names are not among them.)
//!
//! ```toml
//! windit = "0.4"   # smoothing; already in your graph via `ced`
//! ```
//!
//! `zuoer` is the other case. With the `vad` feature on, `audio::vad`
//! re-exports the whole set needed to drive a segmenter — `Run`,
//! `RunSegmenter`, `RunOptions`, `SampleRate` and its `Error` / `Result` — so
//! the segmenting block below names them through coremlit and needs no direct
//! dependency. Under `ced` alone, `zuoer` is not in your graph at all, and you
//! add it yourself:
//!
//! ```toml
//! zuoer = "0.2"    # only for `ced` WITHOUT `vad`
//! ```
//!
//! `windit` also ships its own gate/segment tier
//! (`windit::segment::{Hysteresis, Segmenter, SegmentOptions}`, composed by
//! `windit::decode`) which needs no extra dependency at all. It returns element
//! `Range`s and *no* probability aggregates, so prefer it when a plain interval
//! is enough, and `zuoer::RunSegmenter` when the event needs a confidence
//! attached — which is what the rest of this section shows.
//!
//! ## Scoring the clip
//!
//! Loading a model and running it is the ONE step that needs a staged
//! `.mlmodelc`, so this block — and only this block — is `no_run`:
//! `cargo test --doc` **compiles it and never executes it**. Nothing in it is
//! verified behavior. Everything downstream of it is, because everything
//! downstream of it is arithmetic on the returned numbers:
//!
//! ```no_run
//! use coremlit::audio::ced::{CedModel, Classifier, WindowConfidences, WindowPlan};
//!
//! # let samples_16k: Vec<f32> = Vec::new();
//! let classifier = Classifier::from_file(CedModel::Small.mlmodelc_path("Models/ced"))?;
//! // A 1 s hop across the fixed 10 s window: 90% overlap, one score per second.
//! let plan = WindowPlan::new().with_hop_samples(16_000);
//! let windows: Vec<WindowConfidences> = classifier.classify_windows(&samples_16k, &plan)?;
//! # Ok::<(), coremlit::audio::ced::Error>(())
//! ```
//!
//! ## Projecting one class, and smoothing it
//!
//! [`Confidences::try_from_slice`] builds that `windows` vector by hand, which
//! is what lets the rest of the pipeline **run** here with no model staged —
//! and what lets a consumer unit-test their own event logic the same way:
//!
//! ```
//! use coremlit::audio::ced::{
//!   Confidences, Error, NUM_CLASSES, RatedSoundEvent, Span, WINDOW_SAMPLES, WindowConfidences,
//! };
//! use windit::{
//!   smooth::{Ema, SmoothPolicy},
//!   windowed::Windowed,
//! };
//!
//! let hop = 16_000; // the `WindowPlan` hop the scores were produced at
//! let dog = RatedSoundEvent::from_key("Dog")[0].index();
//! let music = RatedSoundEvent::from_key("Music")[0].index();
//! assert_eq!((dog, music), (74, 137));
//!
//! // Twelve windows of 527 scores, standing in for `classify_windows` output.
//! // `Music` outscores `Dog` in every one of them: the 527 sigmoids are
//! // independent, so two classes can both be loud and they never sum to one.
//! let barks = [0.02, 0.04, 0.71, 0.86, 0.31, 0.90, 0.88, 0.09, 0.03, 0.01, 0.01, 0.02];
//! let windows = barks
//!   .iter()
//!   .enumerate()
//!   .map(|(i, &p)| {
//!     let mut scores = vec![0.0; NUM_CLASSES];
//!     scores[dog] = p;
//!     scores[music] = 0.93;
//!     Ok(WindowConfidences::new(
//!       Confidences::try_from_slice(&scores)?,
//!       Span::new(i * hop, WINDOW_SAMPLES, WINDOW_SAMPLES),
//!     ))
//!   })
//!   .collect::<Result<Vec<WindowConfidences>, Error>>()?;
//!
//! // Stored exactly as handed over: `try_from_slice` takes confidences, not
//! // logits, so it applies no sigmoid (which would read 0.7027 here) and no
//! // renormalization (0.4804 here, and a sum that could never pass one).
//! assert_eq!(windows[3].value().as_slice()[dog], 0.86);
//! assert!(windows[3].value().as_slice().iter().sum::<f32>() > 1.0);
//!
//! // One column out of 527. The span rides along untouched.
//! let track: Vec<Windowed<f32>> = windows
//!   .iter()
//!   .map(|w| Windowed::new(w.value().as_slice()[dog], w.span()))
//!   .collect();
//!
//! // The asked-for class, never the loudest one — projecting an argmax would
//! // have followed `Music` and returned a flat 0.93 track.
//! assert_eq!(*track[3].value(), 0.86);
//! assert_eq!(track[3].span(), windows[3].span());
//!
//! // `Ema`'s alpha is per push; `CadenceEma` denominates its time constant in
//! // input samples instead, so one setting survives an irregular hop.
//! let smoothed = Ema::new(0.6).smooth(&track)?;
//!
//! // Spans are preserved and values rewritten: the lone 0.31 dip at window 4
//! // lifts to ~0.46, so a 0.5/0.35 hysteresis will not tear the event in two.
//! // Unsmoothed that window still reads 0.31; read alpha as the decay weight
//! // rather than the innovation weight and it reads 0.4387.
//! assert_eq!(smoothed[4].span(), track[4].span());
//! assert!((smoothed[4].value() - 0.46).abs() < 5e-3);
//! # Ok::<(), Error>(())
//! ```
//!
#![cfg_attr(
  feature = "vad",
  doc = "## Segmenting the track into events",
  doc = "",
  doc = "`zuoer` reaches this crate through the `vad` feature, so this block is shown",
  doc = "and run under `ced` + `vad`. Like the projection above, it **runs** — the",
  doc = "segmenter takes probabilities, not audio — and it continues that same",
  doc = "example: the hidden preamble is the previous block's code verbatim, so the",
  doc = "whole chain from hand-built window scores to this event's confidence",
  doc = "executes. Note where the segmenter comes from: `audio::vad` re-exports it,",
  doc = "so nothing here names `zuoer`.",
  doc = "",
  doc = "```",
  doc = "use core::time::Duration;",
  doc = "",
  doc = "use coremlit::audio::vad::{RunOptions, RunSegmenter, SampleRate};",
  doc = "# use coremlit::audio::ced::{",
  doc = "#   Confidences, Error, NUM_CLASSES, RatedSoundEvent, Span, WINDOW_SAMPLES, WindowConfidences,",
  doc = "# };",
  doc = "# use windit::{",
  doc = "#   smooth::{Ema, SmoothPolicy},",
  doc = "#   windowed::Windowed,",
  doc = "# };",
  doc = "",
  doc = "let hop = 16_000;",
  doc = "// `smoothed` is the previous block's result verbatim: twelve hand-built",
  doc = "// `WindowConfidences`, the `Dog` column projected out, `Ema::new(0.6)`.",
  doc = "# let dog = RatedSoundEvent::from_key(\"Dog\")[0].index();",
  doc = "# let barks = [0.02, 0.04, 0.71, 0.86, 0.31, 0.90, 0.88, 0.09, 0.03, 0.01, 0.01, 0.02];",
  doc = "# let windows = barks",
  doc = "#   .iter()",
  doc = "#   .enumerate()",
  doc = "#   .map(|(i, &p)| {",
  doc = "#     let mut scores = vec![0.0; NUM_CLASSES];",
  doc = "#     scores[dog] = p;",
  doc = "#     Ok(WindowConfidences::new(",
  doc = "#       Confidences::try_from_slice(&scores)?,",
  doc = "#       Span::new(i * hop, WINDOW_SAMPLES, WINDOW_SAMPLES),",
  doc = "#     ))",
  doc = "#   })",
  doc = "#   .collect::<Result<Vec<WindowConfidences>, Error>>()?;",
  doc = "# let track: Vec<Windowed<f32>> = windows",
  doc = "#   .iter()",
  doc = "#   .map(|w| Windowed::new(w.value().as_slice()[dog], w.span()))",
  doc = "#   .collect();",
  doc = "# let smoothed = Ema::new(0.6).smooth(&track)?;",
  doc = "",
  doc = "// Every number here is a scenario choice for this one class. coremlit picks",
  doc = "// none of them, and there is no CED default set to fall back on.",
  doc = "let options = RunOptions::default()",
  doc = "  .with_sample_rate(SampleRate::Rate16k)",
  doc = "  .with_start_threshold(0.5)",
  doc = "  .with_end_threshold(0.35)",
  doc = "  .with_min_run_duration(Duration::from_secs(2))",
  doc = "  .with_min_gap_duration(Duration::from_secs(2))",
  doc = "  .with_pad(Duration::ZERO);",
  doc = "let mut segmenter = RunSegmenter::new(options);",
  doc = "// One score per planned window, so the segmenter's frame hop IS the plan's.",
  doc = "segmenter.set_frame_hop(hop);",
  doc = "",
  doc = "let mut events = Vec::new();",
  doc = "for window in &smoothed {",
  doc = "  if let Some(run) = segmenter.push_probability(*window.value()) {",
  doc = "    events.push(run);",
  doc = "  }",
  doc = "}",
  doc = "if let Some(run) = segmenter.finish() {",
  doc = "  events.push(run);",
  doc = "}",
  doc = "",
  doc = "assert_eq!(events.len(), 1);",
  doc = "let event = events[0];",
  doc = "// Both thresholds are load-bearing, and the track straddles them: windows 2",
  doc = "// (0.4388) and 7 (0.3812) clear the 0.35 end threshold but not the 0.5 start",
  doc = "// one. Collapse the pair to a single 0.35 gate and the run opens a window",
  doc = "// early, at 2.0; raise the end threshold to 0.40 and it closes a window",
  doc = "// early, at 7.0. (An end threshold at or above the start one is not a third",
  doc = "// option: `RunOptions` normalizes that back to its derived value.)",
  doc = "assert_eq!((event.start_seconds(), event.end_seconds()), (3.0, 8.0));",
  doc = "",
  doc = "// The event's confidence, on zuoer's terms: mean and peak over the run's own",
  doc = "// frames — padding excluded, bridged frames included. See `audio::vad`'s",
  doc = "// \"Segment confidence\" section for the full statement of those rules. So",
  doc = "// this is the mean of the five SMOOTHED in-run windows — over the whole clip",
  doc = "// it would be 0.3230, and smoothing left out entirely (`Ema::new(1.0)`) the",
  doc = "// same options report a 2.0..7.0 event with mean 0.7320.",
  doc = "assert!((event.mean_probability() - 0.6157).abs() < 1e-3);",
  doc = "assert!((event.peak_probability() - 0.8180).abs() < 1e-3);",
  doc = "# Ok::<(), Error>(())",
  doc = "```"
)]
//!
//! # Compute placement (measured, never marketed)
//!
//! [`DEFAULT_COMPUTE`] ships as [`crate::ComputeUnits::All`], MEASURED: the
//! Wave-C pass (`tests/ced/placement.rs`) characterized per-unit parity and
//! latency across all four sizes and this default is what it pinned. See
//! [`DEFAULT_COMPUTE`] for the numbers.
//!
//! # Performance: construct once, reuse, prewarm
//!
//! Construction pays model load/specialization; [`Classifier::prewarm`] runs
//! one throwaway inference to absorb first-prediction specialization before
//! serving. Fan-out is one [`Classifier`] per worker ([`crate::Model`] is
//! `Send` but deliberately not `Sync`).
//!
//! macOS only (built on [`crate`]).

use std::path::Path;

use crate::{
  ComputeUnits, DataType, Model, MultiArray,
  model::contract::{
    Checked, ContractViolation, Dim, FeatureContract, LoadContract, Rendered, StateContract,
  },
};

pub mod aggregate;
pub mod error;
pub mod model;
pub mod prediction;
pub mod window;

mod mel;

pub use aggregate::{ChunkAggregation, aggregate_windows};
pub use error::{
  AudioTooLong, ClassCountMismatch, ContractMismatch, Error, InvalidConfidence, OutputShape,
};
pub use model::{CedModel, ParseCedModelError};
pub use prediction::{
  Confidences, EventPrediction, RatedSoundEvent, SoundEventId, WindowConfidences,
};
pub use window::{DropBelowMin, Span, TailPolicy, WindowPlan};

use crate::audio::ced::{
  error::{Result, WinditError},
  mel::{MelExtractor, N_FRAMES, N_MELS},
};

#[cfg(test)]
mod tests;

/// The sample rate this module's contract is defined at: callers decode and
/// resample to **16 kHz mono f32** before calling (sans-I/O — the workspace
/// convention; CED natively matches it).
pub const SAMPLE_RATE_HZ: u32 = 16_000;

/// The fixed inference-window length in samples: 160 000 = 10 s at 16 kHz,
/// CED's training window. The CoreML export is fixed-shape, so this is model
/// geometry, not a knob (soundevents exposes `window_samples` only because its
/// ONNX graph is dynamic-length — recorded non-goal).
pub const WINDOW_SAMPLES: usize = 160_000;

/// Number of AudioSet classes the model scores: the 527 released rated classes.
/// Compile-time-pinned to `RatedSoundEvent::events().len()` below, so the
/// dataset crate and this module can never drift apart silently.
pub const NUM_CLASSES: usize = 527;

const _: () = assert!(
  soundevents_dataset::RatedSoundEvent::events().len() == NUM_CLASSES,
  "soundevents-dataset's rated label set must have exactly NUM_CLASSES entries"
);

/// Default compute placement: [`ComputeUnits::All`].
///
/// MEASURED, not provisional: the Wave-C placement pass
/// (`tests/ced/placement.rs`) characterized every unit (`CpuOnly`,
/// `CpuAndGpu`, `CpuAndNeuralEngine`, `All`) across all four sizes. Every
/// unit agrees with the `CpuOnly` reference at ≥ 0.99999 cosine and is
/// NaN-free, and warm latency is flat across units (~0.6–0.8 s/clip,
/// dominated by the Rust mel front end, not the CoreML forward) — so
/// `CpuAndGpu` is not faster here, contra the spec's original expectation.
/// The default `All` arm is in fact the numerically
/// *tightest* vs the committed PyTorch fp32 goldens
/// (`tests/ced/parity_logits.rs`: worst cos ~0.99999988, max|Δlogit| ~0.03),
/// and unlike siglip's vision tower the `CpuAndNeuralEngine` arm did not
/// collapse either, so `All` stays the default. Only a *measured* per-size
/// divergence would promote this to a per-[`CedModel`] table; Wave-C found
/// none, so one shared default stands.
pub const DEFAULT_COMPUTE: ComputeUnits = ComputeUnits::All;

/// Declared feature names on the CED `.mlmodelc` (pinned by
/// `tests/ced/model_io.rs`). Wave A DECLARES these; the Wave-B export must
/// emit exactly them (we own the conversion), or they change with the probe —
/// the recorded rework seam.
mod names {
  pub const MEL: &str = "mel";
  pub const LOGITS: &str = "logits";
}

#[cfg(feature = "serde")]
fn default_compute() -> ComputeUnits {
  DEFAULT_COMPUTE
}

/// Construction options for the CED [`Classifier`] (rust-options-pattern): a
/// single `compute` knob with one source of truth shared by
/// `const new`/`Default`.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
#[cfg_attr(feature = "serde", derive(serde::Serialize, serde::Deserialize))]
pub struct ClassifierOptions {
  #[cfg_attr(feature = "serde", serde(default = "default_compute"))]
  compute: ComputeUnits,
}

impl Default for ClassifierOptions {
  fn default() -> Self {
    Self::new()
  }
}

impl ClassifierOptions {
  /// Options matching the module default: [`DEFAULT_COMPUTE`].
  pub const fn new() -> Self {
    Self {
      compute: DEFAULT_COMPUTE,
    }
  }

  /// Which hardware CoreML may schedule the graph on.
  #[inline]
  pub const fn compute(&self) -> ComputeUnits {
    self.compute
  }

  /// Builder form of [`Self::set_compute`].
  #[must_use]
  #[inline]
  pub const fn with_compute(mut self, compute: ComputeUnits) -> Self {
    self.set_compute(compute);
    self
  }

  /// Sets [`Self::compute`] in place.
  #[inline]
  pub const fn set_compute(&mut self, compute: ComputeUnits) -> &mut Self {
    self.compute = compute;
    self
  }
}

/// CED sound-event classifier: 16 kHz mono `&[f32]` in, ranked AudioSet
/// predictions out. Loads any of the four [`CedModel`] sizes — they share one
/// mel→logits contract, so this type is size-agnostic and stores no identity.
///
/// The front-end is a Rust log-mel port (the private `mel` submodule); the
/// fp16 CoreML transformer maps the believed `[1, 64, 1001]` mel to `[1, 527]`
/// PRE-sigmoid logits, and sigmoid + ranking run in Rust.
///
/// Point [`Self::from_file`] / [`Self::load`] at the size you staged, composing
/// the path with [`CedModel::mlmodelc_path`]:
///
/// ```no_run
/// use coremlit::audio::ced::{CedModel, Classifier};
/// let models_root = "Models/ced";
/// Classifier::from_file(CedModel::Small.mlmodelc_path(models_root))?;
/// # Ok::<(), coremlit::audio::ced::Error>(())
/// ```
///
/// `&self` inference (no mutable scratch): the FFT plan and filterbank are
/// built once at load and per-call buffers are local, so fan-out means one
/// [`Classifier`] per worker over a `Send` [`crate::Model`] (`crate::Model` is
/// deliberately `!Sync`).
#[derive(Debug)]
pub struct Classifier {
  /// A [`Checked`], never a bare [`Model`]: [`ced_contract`] is the only
  /// contract this door states and [`Checked::new`] is the only way one is
  /// built, so removing the check from [`Self::load`] does not compile.
  model: Checked,
  mel: MelExtractor,
}

impl Classifier {
  /// Loads the CED `.mlmodelc` from `model_path` with custom `options` — the
  /// primary constructor.
  ///
  /// The model is checked against this door's load contract and held as a
  /// crate-internal `Checked` wrapper whose only constructor runs that check.
  /// The contract IS this door's statement of what it needs, so there is
  /// nothing here to restate:
  ///
  /// ```text
  /// input   mel     f32  [1, 64, 1001]  every axis Exactly
  /// output  logits  f32  [1, 527]       every axis Exactly
  /// state   none
  /// ```
  ///
  /// The four sizes are contract-identical, so every number above is this
  /// door's own and none is read back off the artifact. The ground truth lives
  /// in `tests/ced/model_io.rs`, which now also loads each staged size THROUGH
  /// this constructor.
  ///
  /// The contract is COMPLETE over the members of [`crate::ModelDescription`]
  /// that can make a conformant prediction fail, not just over the features
  /// this door sends: a graph carrying `mel` plus another REQUIRED input clears
  /// every per-feature clause and then fails on every prediction, and a STATE
  /// buffer is not an input at all — it lives in its own dictionary, so a
  /// stateful ML Program declaring exactly `mel` and `logits` plus a state
  /// clears the input set too, and then meets [`Self::raw_scores`], which
  /// predicts through the stateless API CoreML does not let a stateful model be
  /// called with.
  ///
  /// No model is bundled: the `.mlmodelc` is a directory artifact, distributed
  /// via Hugging Face and staged gitignored under `Models/ced/` (Wave B).
  ///
  /// # Errors
  /// [`Error::Load`] if CoreML rejects the model; [`Error::ContractMismatch`]
  /// if a named feature's type or geometry mismatches;
  /// [`Error::UnsatisfiableInput`] if it requires an input this door never
  /// sends; [`Error::UnsatisfiableState`] if it declares a state buffer.
  pub fn load(model_path: impl AsRef<Path>, options: ClassifierOptions) -> Result<Self> {
    let model = Model::load(model_path, options.compute())?;
    let model = Checked::new(model, &ced_contract()).map_err(contract_violation)?;

    Ok(Self {
      model,
      mel: MelExtractor::new(),
    })
  }

  /// Loads the CED `.mlmodelc` with [`ClassifierOptions::new`].
  ///
  /// # Errors
  /// As [`Self::load`].
  pub fn from_file(model_path: impl AsRef<Path>) -> Result<Self> {
    Self::load(model_path, ClassifierOptions::new())
  }

  /// Scores one fixed window: the `[527]` **PRE-sigmoid** logits — the parity
  /// seam and the power-user escape (custom thresholds in logit space).
  ///
  /// `samples_16k` is 16 kHz mono and must be `1..=`[`WINDOW_SAMPLES`] long; a
  /// shorter input is zero-padded to the fixed window (the believed sub-window
  /// policy, probe-pinned in Wave B); a longer input is rejected — never
  /// silently truncated (route long clips to [`Self::classify_windows`] /
  /// [`Self::classify_long`]).
  ///
  /// # Errors
  /// [`Error::EmptyAudio`] if `samples_16k` is empty; [`Error::AudioTooLong`]
  /// if it exceeds [`WINDOW_SAMPLES`]; [`Error::NonFiniteInput`] if any sample
  /// is NaN/infinite (it would silently poison the mel); [`Error::Tensor`] /
  /// [`Error::Prediction`] on a tensor or CoreML failure;
  /// [`Error::OutputShape`] if the predicted `logits` shape diverges from
  /// `[1, `[`NUM_CLASSES`]`]`; [`Error::NonFiniteOutput`] if the model output
  /// has a NaN/infinite logit (model corruption — never reaches sigmoid).
  pub fn raw_scores(&self, samples_16k: &[f32]) -> Result<Vec<f32>> {
    validate_window_input(samples_16k)?;

    let mut features = vec![0.0f32; N_MELS * N_FRAMES];
    self.mel.extract_into(samples_16k, &mut features)?;

    // Freq-major mel [64, 1001] maps directly onto the row-major believed
    // `mel [1, 64, 1001]` contract.
    let input = MultiArray::from_slice(&[1, N_MELS, N_FRAMES], &features)?;
    let mut outputs = self.model.predict_with(&[(names::MEL, &input)])?;
    let logits = outputs
      .take(names::LOGITS)
      .ok_or_else(|| crate::PredictionError::MissingOutput(names::LOGITS.to_string()))?;
    if logits.shape() != [1, NUM_CLASSES] {
      return Err(Error::OutputShape(OutputShape::new(
        logits.shape().to_vec(),
        vec![1, NUM_CLASSES],
      )));
    }

    let mut row = vec![0.0f32; NUM_CLASSES];
    logits.copy_into::<f32>(&mut row)?;
    check_finite_logits(&row)?;
    Ok(row)
  }

  /// Classifies one window: the top `k` classes, descending confidence, ties
  /// broken by ascending class index (the soundevents contract) — ties in the
  /// raw logit are broken by ascending class index; distinct logits that
  /// saturate to equal confidences keep logit order. Runs the min-heap over
  /// raw logits and maps sigmoid at extraction. `k == 0` returns an empty vec
  /// without running the model; `k > `[`NUM_CLASSES`] saturates.
  ///
  /// # Errors
  /// As [`Self::raw_scores`]; [`Error::UnknownClassIndex`] is defensive-only.
  pub fn classify(&self, samples_16k: &[f32], k: usize) -> Result<Vec<EventPrediction>> {
    if k == 0 {
      validate_window_input(samples_16k)?;
      return Ok(Vec::new());
    }
    let logits = self.raw_scores(samples_16k)?;
    prediction::top_k_from_scores(logits.into_iter().enumerate(), k, prediction::sigmoid)
  }

  /// All [`NUM_CLASSES`] classes, **ranked** (descending confidence,
  /// soundevents tie-break) — caller-side thresholding. Note this deliberately
  /// differs from soundevents' `classify_all`, which returns model order; the
  /// spec (§4) pins the ranked form.
  ///
  /// # Errors
  /// As [`Self::classify`].
  pub fn classify_all(&self, samples_16k: &[f32]) -> Result<Vec<EventPrediction>> {
    self.classify(samples_16k, NUM_CLASSES)
  }

  /// The long-clip primitive: per-window sigmoid confidences + their
  /// [`Span`]s, ALWAYS exposed — so time-localized tagging ("when did the dog
  /// bark") is a caller-side read of `windows[i].value().as_slice()[class]`
  /// against `windows[i].span()`, no second API needed.
  ///
  /// Slices `samples_16k` at the plan's offsets and runs one
  /// [`Self::raw_scores`] per span (a short tail is zero-padded by the mel
  /// front-end). Runs sequentially: [`crate::Model`] is `!Sync`, so windows
  /// share one classifier on one thread.
  ///
  /// # Errors
  /// [`Error::EmptyAudio`] if `samples_16k` is empty; [`Error::Windowing`] if
  /// the plan exceeds [`WindowPlan::max_windows`]
  /// ([`WinditError::TooManyWindows`]) or the span/result buffer cannot be
  /// allocated ([`WinditError::AllocFailed`]); otherwise any per-window
  /// [`Self::raw_scores`] error (a [`Error::NonFiniteInput`] index is relative
  /// to the offending window's start).
  pub fn classify_windows(
    &self,
    samples_16k: &[f32],
    plan: &WindowPlan,
  ) -> Result<Vec<WindowConfidences>> {
    if samples_16k.is_empty() {
      return Err(Error::EmptyAudio);
    }
    let spans = plan.spans(samples_16k.len())?;
    // Fallible reservation: the cap already bounds `spans.len()`, but the result
    // vector is still caller-geometry-sized, so reserve it checked rather than
    // risk an infallible `with_capacity` abort under memory pressure.
    let mut out = Vec::new();
    out.try_reserve_exact(spans.len()).map_err(|_| {
      Error::Windowing(WinditError::AllocFailed {
        elements: spans.len(),
      })
    })?;
    for span in spans {
      let logits = self.raw_scores(&samples_16k[span.start()..span.end()])?;
      out.push(WindowConfidences::new(
        Confidences::from_logits(&logits),
        span,
      ));
    }
    Ok(out)
  }

  /// The composed long-clip convenience: scores each planned window and folds
  /// the per-window confidences (`aggregation`, in confidence space) into one
  /// clip-level [`Confidences`], then returns its [`Confidences::top_k`]`(k)`.
  ///
  /// The fold streams through a shared O([`NUM_CLASSES`]) accumulator — the
  /// per-window vectors are never all held at once, so a long clip that plans
  /// many windows does not retain one 527-float vector per window; use
  /// [`Self::classify_windows`] when per-window access is wanted. `k == 0`
  /// returns an empty vec without running the model OR any windowing, so the
  /// [`WindowPlan::max_windows`] cap does not apply to it (it does the same
  /// finite-sample check the model path would).
  ///
  /// # Errors
  /// [`Error::EmptyAudio`] if `samples_16k` is empty; [`Error::Windowing`] if
  /// the plan exceeds [`WindowPlan::max_windows`]
  /// ([`WinditError::TooManyWindows`]) or a buffer cannot be allocated
  /// ([`WinditError::AllocFailed`]); otherwise any per-window
  /// [`Self::raw_scores`] error. ([`Error::EmptyWindows`] is unreachable — a
  /// nonempty clip always plans at least one span.)
  pub fn classify_long(
    &self,
    samples_16k: &[f32],
    k: usize,
    plan: &WindowPlan,
    aggregation: ChunkAggregation,
  ) -> Result<Vec<EventPrediction>> {
    if samples_16k.is_empty() {
      return Err(Error::EmptyAudio);
    }
    if k == 0 {
      // k == 0 skips ALL windowing (and thus the cap), matching the pre-stream
      // behavior; it still rejects a NaN/±∞ clip rather than wave it through.
      check_finite_samples(samples_16k)?;
      return Ok(Vec::new());
    }
    let spans = plan.spans(samples_16k.len())?;
    let mut acc = aggregate::Accumulator::new(aggregation);
    for span in spans {
      let logits = self.raw_scores(&samples_16k[span.start()..span.end()])?;
      acc.push(&Confidences::from_logits(&logits));
    }
    let confidences = acc.finish()?;
    confidences.top_k(k)
  }

  /// Runs one throwaway [`Self::raw_scores`] on a fixed synthetic window to
  /// fully specialize the prediction path, so the first user-facing request is
  /// warm. Construction pays the model load / device specialization; what it
  /// does NOT pay is the first prediction's own graph specialization — calling
  /// `prewarm` once, after construction and before serving, moves that
  /// one-time cost off the first real clip. Then reuse this same classifier
  /// for every request (`&self` — it stays resident).
  ///
  /// The warm-up runs a fixed 1 s 440 Hz tone (zero-padded to the fixed
  /// window), so it neither reads caller audio nor allocates a full-window
  /// buffer up front.
  ///
  /// # Errors
  /// As [`Self::raw_scores`]; a failure here surfaces a broken model at
  /// prewarm time rather than on the first request.
  pub fn prewarm(&self) -> Result<()> {
    let sr = SAMPLE_RATE_HZ as f32;
    let signal: Vec<f32> = (0..SAMPLE_RATE_HZ as usize)
      .map(|i| 0.5 * (std::f32::consts::TAU * 440.0 * (i as f32 / sr)).sin())
      .collect();
    self.raw_scores(&signal)?;
    Ok(())
  }
}

/// Reject a per-window input the pipeline must not see: empty (nothing to
/// classify), longer than the fixed window (never silently truncated — long
/// clips are windowed explicitly), or carrying a NaN/±∞ sample (it would
/// silently poison the mel). Free fn so the guards are hermetically testable
/// without a model.
fn validate_window_input(samples: &[f32]) -> Result<()> {
  if samples.is_empty() {
    return Err(Error::EmptyAudio);
  }
  if samples.len() > WINDOW_SAMPLES {
    return Err(Error::AudioTooLong(AudioTooLong::new(
      samples.len(),
      WINDOW_SAMPLES,
    )));
  }
  check_finite_samples(samples)
}

/// Reject a NaN/±∞ sample ([`Error::NonFiniteInput`]) — it would silently
/// poison the mel. The finite-scan shared by [`validate_window_input`] (the
/// single-window path) and `Classifier::classify_long`'s `k == 0` early
/// return, which must skip `validate_window_input`'s `AudioTooLong` bound (a
/// long clip is expected to exceed [`WINDOW_SAMPLES`]) but must still not
/// wave a NaN/∞ clip through as an empty result.
fn check_finite_samples(samples: &[f32]) -> Result<()> {
  if let Some(index) = samples.iter().position(|v| !v.is_finite()) {
    return Err(Error::NonFiniteInput(index));
  }
  Ok(())
}

/// Classify a NaN/∞ the CoreML runtime produced as model-output corruption
/// ([`Error::NonFiniteOutput`]) before it can reach sigmoid — a NaN logit
/// would silently rank via `total_cmp` and poison downstream aggregation.
fn check_finite_logits(logits: &[f32]) -> Result<()> {
  if let Some(index) = logits.iter().position(|v| !v.is_finite()) {
    return Err(Error::NonFiniteOutput(index));
  }
  Ok(())
}

/// The load contract this door states: `mel` `[1, 64, 1001]` f32 in,
/// `logits` `[1, 527]` f32 out, no state.
///
/// Data rather than a sequence of checks, and the ONLY thing
/// [`Classifier::load`] does beyond calling [`Model::load`]. The four
/// hand-written comparisons this replaced — a presence test and a
/// shape-and-dtype test per feature — were each a check `load` could forget to
/// make, and deleting any of them failed no runnable test. A [`Checked`] field
/// turns that mutation into a compile error; what remains here is the door's
/// own numbers.
///
/// Every axis is [`Dim::Exactly`], and that buys more than the numbers.
/// [`crate::FeatureInfo::shape`] reports the DEFAULT shape of a flexible
/// input, so a `RangeDims` graph converted at `[1, 64, 1001]` declares this
/// contract's exact numbers — and a flexible input is what takes a graph off
/// the accelerator. An all-`Exactly` contract therefore requires the whole
/// feature to be [`crate::ShapeConstraint::Fixed`], which is the only thing
/// that separates the two. Nothing here is read back off the artifact: the
/// four CED sizes are contract-identical, so every number is this door's.
///
/// Built rather than `const` because a [`LoadContract`] owns its axes.
fn ced_contract() -> LoadContract {
  LoadContract::new(
    vec![FeatureContract::new(
      names::MEL,
      DataType::F32,
      vec![
        Dim::Exactly(1),
        Dim::Exactly(N_MELS),
        Dim::Exactly(N_FRAMES),
      ],
    )],
    vec![FeatureContract::new(
      names::LOGITS,
      DataType::F32,
      vec![Dim::Exactly(1), Dim::Exactly(NUM_CLASSES)],
    )],
    StateContract::None,
  )
}

/// Map a [`ContractViolation`] into this module's error vocabulary.
///
/// The two "unsatisfiable" clauses keep their own variants — they are about
/// what the door cannot SUPPLY, not about a feature's declared shape — and the
/// per-feature clauses all land in [`Error::ContractMismatch`], which already
/// carries a feature name and a rendered expected/actual pair. An output the
/// model declares OPTIONAL is one of those: it is a fact about the named
/// feature's declaration, so "expected a required output, got optional" is the
/// shape that pair was made for.
///
/// `ContractViolation::rendered` performs that reduction, so a clause added to
/// the checker later lands in the `Feature` arm rather than breaking this
/// function and its five siblings at once.
fn contract_violation(violation: ContractViolation) -> Error {
  match violation.rendered() {
    Rendered::UnsatisfiableInput(name) => Error::UnsatisfiableInput(name),
    Rendered::UnsatisfiableState(name) => Error::UnsatisfiableState(name),
    Rendered::Feature(feature) => Error::ContractMismatch(ContractMismatch::new(
      feature.feature(),
      feature.clone().expected(),
      feature.actual(),
    )),
  }
}