1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
//! Text utilities: zlib compression-ratio repetition signal, string
//! normalization, and the streaming word-timing prefix/suffix helpers.
//! Ports `Utilities/TextUtilities.swift`
//! `TextUtilities.compressionRatio(of:)` (both overloads),
//! `Utilities/Extensions+Public.swift` `String.normalized`/
//! `String.trimmingSpecialTokenCharacters()`, and two functions from
//! `Utilities/TranscriptionUtilities.swift`: `findLongestCommonPrefix`/
//! `findLongestDifferentSuffix`.
//!
//! `Array.batched` (`Core/WhisperKit.swift:739`, concurrent-worker audio
//! batching) is unrelated to the above and stays out of scope: Plan 3 uses
//! `slice::chunks` directly wherever batching is needed, and this sync,
//! single-threaded port has no concurrent-worker fan-out to batch for.
use Cell;
use autoreleasepool;
use ;
use UnicodeCategories;
use crateWordTiming;
thread_local!
/// Clears this thread's swallowed-compression-error flag. The fallback ladder
/// calls this at the start of each decode attempt so the flag it reads at that
/// attempt's fact merge reflects only that attempt's compressions.
pub
/// Reads and clears this thread's swallowed-compression-error flag: `true` iff
/// [`zlib_compressed_len`] erased at least one OS compression error since the
/// last [`clear_compression_error_swallowed`].
pub
/// The crate-private compression fault seam (mirrors
/// [`crate::audio::whisper::backend::mock::MockBackend::fail_on_call`]): tests script the
/// `call`-th [`zlib_compressed_len`] on a thread to fail exactly as a genuine
/// OS compression error would, so the swallowed-error provenance path can be
/// exercised without an input that forces Foundation's zlib API to error.
pub
/// zlib-compresses `bytes` with Apple's libcompression — the exact
/// `NSData.compressed(using: .zlib)` API Swift WhisperKit's
/// `TextUtilities.compressionRatio` calls
/// (`Utilities/TextUtilities.swift:14-28,33-53`) — and returns the
/// compressed byte length. `None` mirrors Swift's `catch { return
/// .infinity }`: on a genuine OS compression error the caller substitutes
/// [`f32::INFINITY`] rather than `unwrap`-panicking.
///
/// # Why Apple's compressor, not `flate2`/`miniz_oxide`
/// This length feeds the temperature-fallback repetition signal
/// ([`compression_ratio_of_tokens`] -> `decode::finalize_decoding_result`),
/// whose fallback *decision* must match Swift byte-for-byte (coremlit
/// issue #9). Apple's `.zlib` algorithm emits **raw DEFLATE (RFC 1951)** —
/// no zlib wrapper (its output begins e.g. `db c3 …`, not `78 …`) — and
/// compresses markedly harder than `miniz_oxide` on repetitive input.
/// `flate2::write::ZlibEncoder` emits RFC 1950 (a 2-byte header + 4-byte
/// Adler-32 trailer, +6 bytes) and saturates at a weaker ratio; both gaps
/// push our ratio *below* Swift's, flipping the fallback decision at
/// realistic thresholds. No `flate2` compression level reproduces Apple's
/// lengths, so this crate calls the identical Foundation API to get a
/// ratio equal to Swift's by construction. This is a safe `objc2` call —
/// `objc2` owns the FFI.
///
/// The [`autoreleasepool`] matches
/// [`crate::audio::whisper::tokenizer::nl_recognizer::redetect_language`]'s rationale: any
/// Objective-C method may autorelease internally, so a pool must sit on the
/// stack or temporaries leak on a thread with no Cocoa run-loop pool above
/// it. Empty input needs no special-casing: Apple's libcompression
/// compresses a zero-length buffer to a small non-empty result (2 bytes,
/// per the issue-9 objc2 probe), never throwing, so this returns `Some(2)`
/// for `&[]`. `None` is reserved for a genuine OS compression error — the
/// only case Swift's `catch` actually handles.
///
/// Returning `None` (whether from the real API or the crate-private
/// [`fault`] seam) also latches this thread's swallowed-error flag
/// ([`note_compression_error_swallowed`]), the record the fallback ladder reads
/// so an erased error is not silently converted to a reproducible-looking
/// transcript (see [`COMPRESSION_ERROR_SWALLOWED`]).
/// Compression ratio (`raw_bytes / compressed_bytes`) of `tokens`, encoded
/// as **little-endian `i32`** before zlib compression — ports Swift
/// `TextUtilities.compressionRatio(of textTokens: [Int])`
/// (`Utilities/TextUtilities.swift:14-28`): `Int32($0)` per token packed
/// into a `Data` buffer (platform-native byte order, little-endian on
/// every Apple target), then `(data as NSData).compressed(using: .zlib)`.
///
/// An empty `tokens` slice returns **`0.0`**, matching Swift's tokens
/// overload exactly: that overload has **no empty guard**
/// (`Utilities/TextUtilities.swift:14-28`), so it compresses an empty
/// `Data()`, which Apple's libcompression turns into 2 bytes (not an
/// error) — giving `0 / 2 == 0.0`. (Contrast [`compression_ratio_of_text`],
/// which Swift *does* guard.) Only a genuine OS compression error yields
/// [`f32::INFINITY`] here, mirroring Swift's `catch { return .infinity }` —
/// and Swift's tokens overload would land in that identical `catch` on the
/// same error, so the error path matches on both sides too.
///
/// The compression itself goes through the private `zlib_compressed_len`,
/// which calls the same Apple `NSData.compressed(using: .zlib)` API as
/// Swift — see that function for why the codec, not just the byte
/// encoding, must match Swift (coremlit issue #9). The `i32`-LE token
/// encoding below is unchanged: it already matched Apple; only the
/// compressor differed.
/// Compression ratio (`raw_bytes / compressed_bytes`) of `text`'s UTF-8
/// bytes — ports Swift `TextUtilities.compressionRatio(of text: String)`
/// (`Utilities/TextUtilities.swift:33-53`). Returns [`f32::INFINITY`] for
/// empty text via Swift's explicit `if text.isEmpty { return .infinity }`
/// guard (lines 34-36), or on a genuine compression error (Swift's
/// `catch`, lines 49-51). Swift also guards a fallible `text.data(using:
/// .utf8)` (lines 39-42); that path is unreachable here since a Rust
/// `&str` is always valid UTF-8.
///
/// **Empty-input asymmetry (deliberate, matches Swift).** Unlike
/// [`compression_ratio_of_tokens`] — whose Swift overload has *no* empty
/// guard and so returns `0.0` for empty input — this text overload *does*
/// guard empty and returns infinity. The two Swift overloads diverge here
/// on purpose (`:34-36` guards text; `:14-28` does not guard tokens); keep
/// this guard.
///
/// Compresses via the private `zlib_compressed_len` (Apple's `.zlib` = raw
/// DEFLATE), identical to Swift; see [`compression_ratio_of_tokens`].
/// Normalizes `text` for repetition/equality comparisons — ports Swift
/// `String.normalized` (`Utilities/Extensions+Public.swift:24-41`)
/// **exactly**, verified empirically by running the live Swift extension
/// standalone (see this task's report), not just by reading it: lowercase
/// the whole string, then replace the literal ASCII `-` character with a
/// space (a plain, non-regex substring replace — no other dash variant is
/// touched by this step), then **delete** (not replace) every character
/// in Unicode general category `P` (Punctuation: `Pc`/`Pd`/`Pe`/`Pf`/`Pi`/
/// `Po`/`Ps`) — this is the step that removes `_` (category `Pc`) and
/// every non-ASCII dash/quote/CJK punctuation mark, matching Foundation's
/// `CharacterSet.punctuationCharacters` — then collapse runs of the
/// literal space character down to one space, then trim
/// [`char::is_whitespace`] from both ends (this matches Foundation's
/// `.whitespacesAndNewlines`: both sets are exactly U+0009-U+000D,
/// U+0020, U+0085, U+00A0, U+1680, U+2000-U+200A, U+2028, U+2029, U+202F,
/// U+205F, U+3000).
///
/// **Deviation from this task's own brief** (source-corrected, per this
/// task's explicit mandate to verify against source): the brief's
/// semantics sketch ("dashes/underscores to spaces") is wrong for
/// underscores. Swift's dash step is a literal, non-regex
/// `replacingOccurrences(of: "-", with: " ")` that matches only the ASCII
/// hyphen; `_` is punctuation category `Pc` and gets *deleted* by the
/// punctuation step, not turned into a space. Confirmed by running the
/// actual Swift extension standalone: `"multi-word_test".normalized ==
/// "multi wordtest"`, not `"multi word test"` as the brief's own given
/// test asserted — see the task report for the probe script and full
/// output; this module's tests reflect the verified behavior.
/// Strips leading/trailing `<`/`|`/`>` characters, repeatedly, from both
/// ends — ports Swift `String.trimmingSpecialTokenCharacters()`
/// (`Utilities/Extensions+Public.swift:43-45`), which trims
/// `Constants.specialTokenCharacters` (`Core/Models.swift:1332`,
/// `CharacterSet(charactersIn: "<|>")`). This is a **character-class**
/// trim, not a fixed `"<|"`/`"|>"` substring strip: e.g. `"<<|x|>"` trims
/// to `"x"` (every wrapping character is a member of the set), not
/// `"<<|x"` (which a literal-prefix reading of this function's own
/// summary would wrongly produce, since `"<<|x|>"` does not start with
/// the literal two-character substring `"<|"`).
/// Longest run of word-by-word agreement between two decode passes over
/// the same audio span, comparing [`normalized`] text — ports
/// `TranscriptionUtilities.findLongestCommonPrefix`
/// (`Utilities/TranscriptionUtilities.swift:34-37`): `zip(words1,
/// words2).prefix(while: { $0.word.normalized == $1.word.normalized })`,
/// mapped to `$0.1`, i.e. the **returned elements come from `current`**,
/// the newer pass, not `previous`. A borrowed prefix of `current` replaces
/// Swift's array copy — identical contents, zero allocation. Stops at the
/// shorter of the two inputs, exactly like `zip`.
/// `current` past its agreement with `previous` — ports
/// `TranscriptionUtilities.findLongestDifferentSuffix`
/// (`Utilities/TranscriptionUtilities.swift:44-48`): `words2[commonPrefix.
/// count...]`. When `previous` and `current` share no common prefix at
/// all, the whole of `current` is returned.