1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
//! The structural check on a caller-supplied tokenizer's post-processor that
//! the tokenizers crate's own deserializer skips.
//!
//! # The boundary
//!
//! `tokenizers::processors::template::TemplateProcessing` is built two ways.
//! `TemplateProcessingBuilder::build` runs a `validate` step; deserializing a
//! `tokenizer.json` does **not** — the file goes through
//! `From<TemplateProcessingDeserializer>`, which only recomputes the added-token
//! counts. So a template that the builder would refuse still *parses*, and the
//! refusal is deferred to the first `encode`, where it is a **panic** inside the
//! dependency rather than an error:
//!
//! * a `SpecialToken` piece naming an id absent from `special_tokens` indexes
//! `self.special_tokens.0[id]` (a `HashMap` index) — "no entry found for key";
//! * a `Sequence` piece with id `B` in the SINGLE template indexes
//! `encodings[1]` when a single sequence was encoded — "index out of bounds";
//! * a single template with no `$A` at all silently drops the text and encodes
//! every input as its special tokens alone — the same degenerate answer the
//! `>= window` special-token-overhead rule refuses, reported here instead.
//!
//! # Which post-processor runs at which count
//!
//! Those rules are about ONE template, and which template a `TemplateProcessing`
//! applies is not fixed: `process_encodings` selects it by how many encodings it
//! was handed — `2 => pair`, `1 => single`, and every other count is a `todo!()`,
//! i.e. another panic ("not yet implemented"). That count is not always one.
//! `apply_template` emits ONE ENCODING PER PIECE of the template it applied — a
//! `Sequence` piece clones the encoding it names, a `SpecialToken` piece builds
//! one from its ids — and the merge back into a single encoding happens at the
//! TOKENIZER level, after the whole post-processor has run. A post-processor
//! `Sequence` meanwhile threads each member's output into the next. So a
//! three-piece template inside a `Sequence` hands three encodings to whatever
//! follows it.
//!
//! The guard therefore does not judge each template independently: it SIMULATES
//! the encoding count along the chain, for the single-sequence encode the doors
//! here perform. The count starts at 1; a `Sequence` threads it through its
//! members in order, nested ones included; a `TemplateProcessing` hands on its
//! selected template's piece count; and every kind that ADDS TOKENS —
//! `TemplateProcessing`, `RobertaProcessing`, `BertProcessing` — must be reached
//! at EXACTLY ONE encoding or it is refused. `ByteLevel` adds nothing at any
//! count and only trims offsets, so it passes the count through untouched. Any
//! FINAL count is fine — that one the tokenizer merges.
//!
//! # Why exactly one, and why `$A` exactly once
//!
//! Both halves are about the door's TRUNCATION, which the tokenizer applies to
//! the RAW encoding — before the post-processor — at
//! `max_length - added_tokens(false)`. That subtrahend is the ONLY overhead
//! number in the system: `with_truncation` reads it, `post_process` reads it
//! again on every `encode`, and the doors' special-token-overhead guard reads
//! the same one. For a post-processor `Sequence` it is the SUM of the members'
//! `added_single`, whichever mode each member actually runs in.
//!
//! * **Exactly one encoding.** A `TemplateProcessing` reached at two applies its
//! `pair` template and adds `added_pair`, which that reading does not report;
//! `RobertaProcessing` and `BertProcessing` reached at `n` wrap EVERY encoding
//! (`cls … sep` around the first, `sep … sep` around each of the rest) and add
//! `2n` while reporting a flat 2. Either way the truncation is sized on an
//! overhead the chain does not add. Measured on 0.23.1:
//! `Sequence[Template("[CLS] $A [SEP]"), RobertaProcessing]` declares 4 and
//! returns 12 ids at an 8-token window. Counts of 0, 3 and up have no template
//! at all — that is the `todo!()`. Refusing every count but one costs nothing
//! real: no door here encodes a pair, so a chain that manufactures a second
//! encoding is not a shape any of them needs.
//! * **`$A` exactly once.** Each placement emits its own copy of the text, and
//! `count_added` scores a `Sequence` piece as zero however often it appears,
//! so `Template("$A $A")` advertises NO overhead and returns twice the
//! truncated length — measured, a 3-token text comes back as 6 ids at a
//! 4-token window, which the door then refuses with a typed `TokenCount` for
//! ordinary text although construction succeeded. Zero placements is the other
//! end of the same rule: the text dropped entirely.
//!
//! Consequently the count a template hands on is its piece count, and a second
//! token-adding post-processor downstream is admitted only when the first was
//! the identity `$A` — one piece, one encoding. Nothing legitimate is refused by
//! that: no single-sequence tokenizer shape in common use chains two templating
//! processors. RoBERTa and BERT carry one `TemplateProcessing`, or their
//! dedicated `RobertaProcessing` / `BertProcessing`; GPT-2 and the other
//! byte-level BPEs a bare `ByteLevel`; the SentencePiece families (Gemma, T5)
//! one `TemplateProcessing`. Where a `Sequence` appears at all it pairs a
//! `ByteLevel` with a single template, which this admits.
//!
//! [`check_post_processor`] is the tokenizers builder's skipped `validate`,
//! applied to the template each `TemplateProcessing` in the chain would really
//! select, plus the placement and count rules that `validate` has no reason to
//! carry — the builder owns no truncation window. What it proves is bounded and
//! exact: the single-sequence encode reaches no `todo!()` and no undeclared-key
//! index inside the dependency's post-processing, the text is placed EXACTLY
//! ONCE, and `added_tokens(false)` is EXACTLY the overhead the chain adds. Those
//! three are what make the doors' `>= window` overhead guard sound as written —
//! a raw encoding truncated to `max_length - added` post-processes to at most
//! `max_length`. Everything else about a `tokenizer.json`'s internal consistency
//! remains the tokenizers crate's contract; no door here re-implements its
//! deserializer. A post-processor kind this module does not recognize is the one
//! place that reasoning stops: it carries no template to judge and no overhead
//! this module can know, so it passes and is credited with preserving the count,
//! as every kind that exists today does.
//!
//! A door runs it BEFORE reading `added_tokens(false)` off the post-processor.
//! Not because that reading is unsafe: `count_added` scores an undeclared
//! `SpecialToken` id as **zero**, so the special-token-overhead guard is simply
//! BLIND to the templates this module refuses — it cannot substitute for this
//! check, and this check catches them whichever order the two run in. The order
//! is a diagnostic choice. A count derived from a malformed template is not a
//! fact about the tokenizer, so a file that breaks both rules is reported by its
//! structural defect rather than by a number the caller would try to shrink.
//!
//! # Why serialization
//!
//! `TemplateProcessing` exposes no accessor for `special_tokens`, and its
//! `get_single` renders a `Debug` string. Its serde representation is the
//! `tokenizer.json` shape itself, so the structure is read back out through
//! `serde_json::to_value` — the same JSON the file supplied.
use Tokenizer;
/// Which structural rule of a post-processor was broken.
///
/// Payload of each door's `Error::PostProcessorTemplate` variant
/// ([`crate::embeddings::clap::error::Error::PostProcessorTemplate`] /
/// [`crate::embeddings::siglip::error::Error::PostProcessorTemplate`]).
/// Applies the tokenizers builder's skipped `validate` to `tokenizer`'s
/// post-processor, plus the count and placement rules the door's truncation
/// needs — see the module docs for the exact boundary.
///
/// A tokenizer with no post-processor passes. `ByteLevel` adds no tokens and
/// maps encodings 1:1, so it passes at any count, and so does a kind added after
/// this was written. `RobertaProcessing` and `BertProcessing` map encodings 1:1
/// but WRAP each of them, so they pass only at a count of one. A `Sequence`
/// post-processor is walked in order, because it applies each of its members to
/// the previous one's output.
///
/// # Errors
/// [`PostProcessorTemplate`] naming the first rule the chain breaks, for the
/// SINGLE-sequence encode every door here performs.
pub
/// One serialized post-processor, given the number of encodings reaching it;
/// answers with the number it hands on.
///
/// Dispatch is on the `type` tag: a `Sequence` threads the count through its
/// members in order (nested `Sequence`s included), a `TemplateProcessing` is
/// judged, `RobertaProcessing` and `BertProcessing` are admitted only at one,
/// and every other kind preserves the count.
/// The rules, over the SINGLE template — the only one a `TemplateProcessing` in
/// an admitted chain ever applies.
///
/// The count is refused first: the dependency selects `pair` at two and panics
/// at anything but one or two, and neither is a mode whose overhead
/// `added_tokens(false)` reports, so no rule below is worth judging for a
/// template that would not run in single mode. `$B` stays refused because
/// applying it to one encoding indexes `encodings[1]`, and `$A` must appear
/// exactly once because every placement emits its own copy of the text.
///
/// The `single` template being absent or not an array leaves no pieces, so it
/// falls out of the no-input rule below rather than needing a shape check of its
/// own — which keeps the guard closed on a shape this module does not expect.