pub struct QwenChatProvider { /* private fields */ }Expand description
A real local chat backend: a Qwen2 GGUF model + its tokenizer, driven
through el_runtime::InferenceSession.
The model weights are loaded once at construction and kept resident
(ADR-018): each chat renders the whole conversation to Qwen2.5 ChatML and
reuses one persistent provenance-gated session — a follow-up turn reuses the
cached KV prefix and prefills only the new suffix (continue_prompt, AC-3
cross-turn incremental prefill), while a fresh turn uses load_prompt —
then generate (grammar mask → safety steer → guard + checkpointed rollback
→ greedy commit), instead of rebuilding the engine and re-reading the GGUF
every turn.
On-device safety (ADR-005 Lightweight tier + the ADR-012 control loop) is
on by default; see with_safety. The resident model
lives behind a Mutex, so the provider stays Send + Sync and concurrent
chat calls serialize on the one conversation.
Implementations§
Source§impl QwenChatProvider
impl QwenChatProvider
Sourcepub fn from_paths(
model_path: impl AsRef<Path>,
tokenizer_path: impl AsRef<Path>,
) -> Result<Self>
pub fn from_paths( model_path: impl AsRef<Path>, tokenizer_path: impl AsRef<Path>, ) -> Result<Self>
Load a Qwen2 GGUF model and its tokenizer.json from local paths.
Sourcepub fn with_safety(self, mode: SafetyMode) -> Self
pub fn with_safety(self, mode: SafetyMode) -> Self
Select the on-device safety tier (ADR-005). SafetyMode::Off disables
the steerer and the ADR-012 control loop (the plain single-pass decode);
SafetyMode::Lightweight (the default) runs the token-anchor guard +
hard-ban steerer + checkpointed rollback. SecDecoding/Csd need model
assets not shipped here and fall back to the Lightweight wiring.
Sourcepub fn with_extra_guard_words<I, S>(self, words: I) -> Self
pub fn with_extra_guard_words<I, S>(self, words: I) -> Self
Add extra words to the chunk guard’s unsafe patterns (resolved to token
ids via this model’s tokenizer). Primarily a test/demo hook: e.g.
--guard-word banana lets you watch the ADR-012 rollback / fail-closed
refusal fire on a benign word, without needing the model to emit genuinely
harmful content. Guard-only — these are not added to the hard-ban list, so
the trajectory loop (not silent suppression) is what engages.
Sourcepub fn with_expert_model(self, path: impl AsRef<Path>) -> Self
pub fn with_expert_model(self, path: impl AsRef<Path>) -> Self
Enable model-backed contrastive steering (ADR-013) with a safety
expert GGUF (same tokenizer/family as the chat model). Steering runs
only inside the early-token window. Pointing this at the chat model itself
gives ~zero contrast (a no-op); a safety-tuned Qwen GGUF gives real
steering. No effect under --safety off.
Sourcepub fn with_steer_alpha(self, alpha_milli: i32) -> Self
pub fn with_steer_alpha(self, alpha_milli: i32) -> Self
Contrastive steering strength ×1000 (1000 = 1.0×). Only meaningful with
with_expert_model.
Sourcepub fn end_session(&self) -> Result<()>
pub fn end_session(&self) -> Result<()>
End the current conversation, releasing its KV / output / prompt / buffered
events while keeping the model resident (ADR-018 separation of
conversation and model lifecycles; the AC-4 explicit release / PRD line 131
“KV caches … cleared on session end”). The next chat starts a fresh
conversation on the same loaded weights — no reload. A no-op if no
conversation has started yet. To free the weights too, drop the provider
(Rust ownership).
Trait Implementations§
Source§impl LlmProvider for QwenChatProvider
impl LlmProvider for QwenChatProvider
Source§fn chat(&self, req: &ChatRequest) -> Result<ChatResponse>
fn chat(&self, req: &ChatRequest) -> Result<ChatResponse>
Source§fn chat_stream(
&self,
req: &ChatRequest,
on_token: &mut dyn FnMut(ChatToken),
) -> Result<()>
fn chat_stream( &self, req: &ChatRequest, on_token: &mut dyn FnMut(ChatToken), ) -> Result<()>
on_token for each fragment as it arrives.
Returns when generation is complete or on error. The final call will
have ChatToken::is_final == true.Auto Trait Implementations§
impl !Freeze for QwenChatProvider
impl RefUnwindSafe for QwenChatProvider
impl Send for QwenChatProvider
impl Sync for QwenChatProvider
impl Unpin for QwenChatProvider
impl UnsafeUnpin for QwenChatProvider
impl UnwindSafe for QwenChatProvider
Blanket Implementations§
Source§impl<T> BorrowMut<T> for Twhere
T: ?Sized,
impl<T> BorrowMut<T> for Twhere
T: ?Sized,
Source§fn borrow_mut(&mut self) -> &mut T
fn borrow_mut(&mut self) -> &mut T
impl<ST, DT> CastableFrom<ST, Initialized, Initialized> for DT
impl<ST, DT> CastableFrom<ST, Uninit, Uninit> for DT
impl<T> ErasedDestructor for Twhere
T: 'static,
Source§impl<T> Instrument for T
impl<T> Instrument for T
Source§fn instrument(self, span: Span) -> Instrumented<Self>
fn instrument(self, span: Span) -> Instrumented<Self>
Source§fn in_current_span(self) -> Instrumented<Self>
fn in_current_span(self) -> Instrumented<Self>
Source§impl<T> IntoEither for T
impl<T> IntoEither for T
Source§fn into_either(self, into_left: bool) -> Either<Self, Self>
fn into_either(self, into_left: bool) -> Either<Self, Self>
self into a Left variant of Either<Self, Self>
if into_left is true.
Converts self into a Right variant of Either<Self, Self>
otherwise. Read moreSource§fn into_either_with<F>(self, into_left: F) -> Either<Self, Self>
fn into_either_with<F>(self, into_left: F) -> Either<Self, Self>
self into a Left variant of Either<Self, Self>
if into_left(&self) returns true.
Converts self into a Right variant of Either<Self, Self>
otherwise. Read more