pub struct ContextManager {Show 15 fields
pub max_tokens: usize,
pub reserve_ratio: f32,
pub truncation_ratio: f32,
pub min_keep_rounds: usize,
pub token_safety_margin: f32,
pub max_output_lines: usize,
pub max_output_bytes: usize,
pub compression_token_threshold: usize,
pub compression_enabled: bool,
pub max_tool_calls_per_turn: usize,
pub progressive_compression: bool,
pub rounds_per_summary: usize,
pub max_summary_segments: usize,
pub merge_count: usize,
pub max_merges_per_segment: usize,
}Expand description
Manages the context window, truncating history when approaching token limits.
Fields§
§max_tokens: usizeModel’s context window size in tokens.
reserve_ratio: f32Ratio of context window to reserve for LLM response (default 0.2 = 20%).
truncation_ratio: f32Fraction of max_tokens at which truncation triggers (default 0.7).
min_keep_rounds: usizeMinimum conversation rounds to keep after truncation (default 3).
token_safety_margin: f32Safety multiplier for token estimates (default 1.3).
max_output_lines: usizeMax output lines for tool results.
max_output_bytes: usizeMax output bytes for tool results.
compression_token_threshold: usizeToken threshold for triggering compression.
compression_enabled: boolWhether compression is enabled.
max_tool_calls_per_turn: usizeMaximum tool calls per turn (default 30).
progressive_compression: boolWhether progressive segmented compression is enabled (default true).
rounds_per_summary: usizeNumber of full rounds per summary segment (default 3).
max_summary_segments: usizeMaximum number of summary segments to keep (default 5).
merge_count: usizeNumber of segments to merge at a time (default 2).
max_merges_per_segment: usizeMaximum merges per segment before discarding (default 2).
Implementations§
Source§impl ContextManager
impl ContextManager
pub fn new(context_window: Option<u64>, config: Option<&ContextConfig>) -> Self
Sourcepub fn truncation_threshold(&self) -> usize
pub fn truncation_threshold(&self) -> usize
Maximum tokens at which truncation is triggered.
Uses truncation_ratio (default 0.7) to trigger earlier than the
absolute limit, leaving headroom for estimation errors and LLM response.
Sourcepub fn available_tokens(&self) -> usize
pub fn available_tokens(&self) -> usize
Maximum tokens available for input (total - reserved for response).
Note: this is the absolute upper bound; truncation actually triggers
earlier via truncation_threshold().
Sourcepub fn truncate_tool_output(&self, content: &str) -> String
pub fn truncate_tool_output(&self, content: &str) -> String
Truncate tool output using the configured limits.
Sourcepub fn estimate_context_tokens(
&self,
messages: &[ChatCompletionRequestMessage],
last_known_prompt_tokens: Option<u32>,
snapshot_message_count: usize,
) -> usize
pub fn estimate_context_tokens( &self, messages: &[ChatCompletionRequestMessage], last_known_prompt_tokens: Option<u32>, snapshot_message_count: usize, ) -> usize
Hybrid token estimation: precise API anchor + incremental estimation.
When last_known_prompt_tokens is Some and history hasn’t been truncated
(i.e. messages.len() >= snapshot_message_count), uses the API-reported
prompt_tokens as an exact baseline and adds estimated tokens for only
the new messages appended since the snapshot. This is far more accurate
than full heuristic estimation because the baseline is from the model’s
own tokenizer.
Falls back to full heuristic estimation with token_safety_margin when:
- No calibration data yet (first call, before any API response)
- History was truncated/compressed (messages.len() < snapshot)
Sourcepub fn maybe_truncate(
&self,
messages: &mut Vec<ChatCompletionRequestMessage>,
last_known_prompt_tokens: Option<u32>,
snapshot_message_count: usize,
) -> TruncationResult
pub fn maybe_truncate( &self, messages: &mut Vec<ChatCompletionRequestMessage>, last_known_prompt_tokens: Option<u32>, snapshot_message_count: usize, ) -> TruncationResult
Check if history needs truncation and perform it if necessary.
Returns TruncationResult with removed messages for async compression.
Strategy (progressive compression):
- Uses
truncation_threshold()(default 70% of max_tokens) as trigger point. - When over threshold, performs exactly one compression action per call:
- Priority 1: Compress the oldest
rounds_per_summaryfull rounds into a new summary segment. - Priority 2: Merge the oldest
merge_countsummary segments into one. - Priority 3: Discard the oldest summary segment (if merge limit reached).
- Priority 4: Fall back to aggressive truncation (old behavior).
- Priority 1: Compress the oldest
- If
progressive_compressionis disabled, falls back to single-shot truncation.