pub enum SpmPrefixScheme {
Once,
AfterEachSpecial,
}Expand description
Where the dummy prefix (add_dummy_prefix / add_space_prefix) is placed
once the input is split on added tokens.
The two reference implementations of this same vocabulary format genuinely
disagree, and both were measured rather than inferred — so neither is “the”
behavior and the loader that built the tokenizer has to say which one its
vocabulary was produced for. Encoding "[INST]Write" and "a[INST]b":
Once (HF) | AfterEachSpecial (llama.cpp) | |
|---|---|---|
[INST]Write | ▁, [INST], Write | [INST], ▁Write |
a[INST]b | ▁a, [INST], b | ▁a, [INST], ▁b |
Getting this wrong is invisible from the outside — every id stays in range and decodes back to the original string — while every chat prompt (which is exactly a text/marker alternation) reaches the model as pieces it was not trained on.
The default is AfterEachSpecial, because
SpmTokenizer::new takes a GGUF-style vocabulary and llama.cpp is the
reference for those. A vocabulary lifted out of a HuggingFace
tokenizer.model is the case that has to be declared, and its one loader
does declare it.
Variants§
Once
Prefix the whole text once, before splitting on added tokens
(HuggingFace / sentencepiece).
SentencePiece normalizes — and therefore prefixes — the input and only
then splits, so only the stretch beginning at byte 0 can carry a marker.
A leading added token leaves the marker standing alone, with no text to
attach to. Measured with
AutoTokenizer.from_pretrained("mistral-7b-v0.3", use_fast=False) and
add_special_tokens=False: "[INST]Write" -> [29473, 3, 6006]
(▁, [INST], bare Write) and "a[INST]b" -> [1032, 3, 29494]
(▁a, [INST], bare b).
This is HuggingFace’s corrected behavior (legacy = false in
tokenizer_config.json). Which scheme a bundled .spm vocabulary needs
is not determined by the fact that it came from a tokenizer.model —
it is that per-checkpoint legacy flag: Mistral V2 sets legacy = false
and needs Once, but Mistral V1 sets legacy = true and needs
AfterEachSpecial despite also being extracted
from a tokenizer.model — see spm_prefix_scheme in pretrained.rs,
which reads that flag off per vocabulary rather than assuming it.
AfterEachSpecial
Prefix the first stretch and every stretch that follows an added
token (llama.cpp’s is_prev_special).
llama-vocab.cpp’s LLAMA_VOCAB_TYPE_SPM arm walks the fragment buffer
with bool is_prev_special = true (“prefix with space if first token”),
prepending ' ' to a raw-text fragment whenever the flag is set and
re-arming it on every special-token fragment. A special token at the very
start therefore emits no standalone marker — there is no text
fragment before it to prefix.
Correct for every GGUF-loaded vocabulary, because llama.cpp is what
actually runs those files — and so the default, see the type’s docs.
Also correct for a bundled .spm vocabulary whose checkpoint declares
legacy = true (Mistral V1) — see Once’s docs.
Trait Implementations§
Source§impl Clone for SpmPrefixScheme
impl Clone for SpmPrefixScheme
Source§fn clone(&self) -> SpmPrefixScheme
fn clone(&self) -> SpmPrefixScheme
1.0.0 (const: unstable) · Source§fn clone_from(&mut self, source: &Self)
fn clone_from(&mut self, source: &Self)
source. Read moreimpl Copy for SpmPrefixScheme
Source§impl Debug for SpmPrefixScheme
impl Debug for SpmPrefixScheme
Source§impl Default for SpmPrefixScheme
impl Default for SpmPrefixScheme
Source§fn default() -> SpmPrefixScheme
fn default() -> SpmPrefixScheme
impl Eq for SpmPrefixScheme
Source§impl PartialEq for SpmPrefixScheme
impl PartialEq for SpmPrefixScheme
impl StructuralPartialEq for SpmPrefixScheme
Auto Trait Implementations§
impl Freeze for SpmPrefixScheme
impl RefUnwindSafe for SpmPrefixScheme
impl Send for SpmPrefixScheme
impl Sync for SpmPrefixScheme
impl Unpin for SpmPrefixScheme
impl UnsafeUnpin for SpmPrefixScheme
impl UnwindSafe for SpmPrefixScheme
Blanket Implementations§
Source§impl<T> BorrowMut<T> for Twhere
T: ?Sized,
impl<T> BorrowMut<T> for Twhere
T: ?Sized,
Source§fn borrow_mut(&mut self) -> &mut T
fn borrow_mut(&mut self) -> &mut T
Source§impl<T> CloneToUninit for Twhere
T: Clone,
impl<T> CloneToUninit for Twhere
T: Clone,
Source§impl<T> IntoEither for T
impl<T> IntoEither for T
Source§fn into_either(self, into_left: bool) -> Either<Self, Self> ⓘ
fn into_either(self, into_left: bool) -> Either<Self, Self> ⓘ
self into a Left variant of Either<Self, Self>
if into_left is true.
Converts self into a Right variant of Either<Self, Self>
otherwise. Read moreSource§fn into_either_with<F>(self, into_left: F) -> Either<Self, Self> ⓘ
fn into_either_with<F>(self, into_left: F) -> Either<Self, Self> ⓘ
self into a Left variant of Either<Self, Self>
if into_left(&self) returns true.
Converts self into a Right variant of Either<Self, Self>
otherwise. Read more