1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
//! Per-window confidence aggregation for long clips:
//! [`ChunkAggregation`] `{ Mean, Max }` + [`aggregate_windows`] — soundevents'
//! chunked-inference semantics exactly.
//!
//! Both the batch [`aggregate_windows`] and `Classifier::classify_long` fold
//! through one streaming `Accumulator`, so Mean/Max are single-sourced: the
//! long-clip path never materializes a per-window vector for every window, yet
//! its result is bit-identical to aggregating the materialized slice.
//!
//! Aggregation runs in **confidence space** (per-window sigmoid confidences),
//! never logit space: sigmoid is nonlinear, so the two disagree, and
//! soundevents defines the contract as sigmoid-then-aggregate (pinned by a
//! mutation-red test in the sibling `tests.rs`). Mean is the equal-weight
//! arithmetic mean (f32 accumulation, one divide at the end — a single window
//! aggregates to itself bit-exactly); Max is the elementwise peak.
//!
//! windit's aggregation engine is deliberately NOT used here: its built-ins
//! are renormalizing unit-vector policies, the wrong domain for independent
//! per-class probabilities (spec §2).
use crate;
/// Controls how a long clip's per-window confidences combine into one
/// clip-level [`Confidences`] — soundevents' `ChunkAggregation`, spelling and
/// semantics (wire spellings `"mean"` / `"max"`, golden-pinned).
/// Streaming Mean/Max fold shared by [`aggregate_windows`] and
/// `Classifier::classify_long`, one window at a time — `classify_long` folds
/// each window's confidences in and never materializes the per-window vectors,
/// so a long clip retains O([`NUM_CLASSES`]) state rather than one 527-float
/// vector per window.
///
/// Bit-identical to the batch fold by construction: the SAME op sequence — copy
/// the first window, then `+=` (Mean) / `max` (Max) in window order, one
/// trailing divide gated on `count > 1` — over the same f32 storage, so the
/// golden aggregation values do not shift. Do NOT "improve" to f64 accumulation:
/// "f32 accumulation, one divide at the end" is the golden-pinned contract.
///
/// [`NUM_CLASSES`]: crate::audio::ced::NUM_CLASSES
pub
/// Combine per-window confidences into one clip-level [`Confidences`] under
/// `aggregation`, folding them through the shared `Accumulator`. Every
/// window's vector is [`NUM_CLASSES`]-long by type invariant, so no length
/// reconciliation is needed; finiteness is preserved (mean/max of finite
/// `[0, 1]` values is finite).
///
/// A single window aggregates to itself bit-exactly (soundevents divides only
/// when the window count exceeds 1).
///
/// # Errors
/// [`Error::EmptyWindows`] if `windows` is empty. (Unreachable through
/// `classify_long` — a nonempty clip always plans at least one span.)
///
/// [`NUM_CLASSES`]: crate::audio::ced::NUM_CLASSES