pub struct AskConfig {Show 15 fields
pub k_summary: u32,
pub k_raw: u32,
pub escalation_threshold: f64,
pub mmr_threshold: f64,
pub max_context_tokens: u32,
pub response_tokens: u32,
pub timeout_secs: u32,
pub min_score: f64,
pub continue_history_turns: u32,
pub rewriter_timeout_secs: u32,
pub compress_hits_enabled: bool,
pub summarize_hits_enabled: bool,
pub summarize_model: Option<String>,
pub backend: Option<BackendConfig>,
pub rewriter_backend: Option<BackendConfig>,
}Fields§
§k_summary: u32§k_raw: u32§escalation_threshold: f64§mmr_threshold: f64§max_context_tokens: u32§response_tokens: u32§timeout_secs: u32§min_score: f64§continue_history_turns: u32§rewriter_timeout_secs: u32Separate, shorter timeout for the rewriter LLM call (Phase 3.3).
Rewriter output is small (~80 tokens) and falling back to the raw
question on failure is non-fatal, so we don’t want to burn the full
timeout_secs budget waiting on a slow/unreachable Ollama before
the user sees any response.
compress_hits_enabled: bool§summarize_hits_enabled: bool§summarize_model: Option<String>§backend: Option<BackendConfig>Per-stage backend override for the answer-generation model.
None = inherit the smart slot (config.llm).
rewriter_backend: Option<BackendConfig>Per-stage backend override for the query rewriter.
None = inherit the answer stage’s backend (self.backend), falling
through to the smart slot (config.llm) only if that is also unset.
Implementations§
Source§impl AskConfig
impl AskConfig
Sourcepub fn effective_backend(&self, llm: &LlmConfig) -> BackendConfig
pub fn effective_backend(&self, llm: &LlmConfig) -> BackendConfig
Effective backend for answer generation. An explicit per-stage
override wins; otherwise the stage inherits the smart slot
(config.llm) with this stage’s own timeout baked in, so a slow
backend cannot silently fall back to the factory’s 120s default.
Sourcepub fn effective_rewriter_backend(&self, llm: &LlmConfig) -> BackendConfig
pub fn effective_rewriter_backend(&self, llm: &LlmConfig) -> BackendConfig
Effective backend for the query rewriter. An explicit
rewriter_backend override wins outright. Otherwise the rewriter
follows the answer stage’s backend (self.backend), and only falls
through to the smart slot (llm) when the answer stage has no
override either — the rewriter is not an independent pinning point,
only an independent timeout. Whichever source it resolves from, it
always keeps its own much tighter rewriter_timeout_secs budget:
the rewriter’s output is small and falling back to the raw question
on timeout is non-fatal.