Skip to main content

ContextManager

Struct ContextManager 

Source
pub struct ContextManager {
Show 15 fields pub max_tokens: usize, pub reserve_ratio: f32, pub truncation_ratio: f32, pub min_keep_rounds: usize, pub token_safety_margin: f32, pub max_output_lines: usize, pub max_output_bytes: usize, pub compression_token_threshold: usize, pub compression_enabled: bool, pub max_tool_calls_per_turn: usize, pub progressive_compression: bool, pub rounds_per_summary: usize, pub max_summary_segments: usize, pub merge_count: usize, pub max_merges_per_segment: usize,
}
Expand description

Manages the context window, truncating history when approaching token limits.

Fields§

§max_tokens: usize

Model’s context window size in tokens.

§reserve_ratio: f32

Ratio of context window to reserve for LLM response (default 0.2 = 20%).

§truncation_ratio: f32

Fraction of max_tokens at which truncation triggers (default 0.7).

§min_keep_rounds: usize

Minimum conversation rounds to keep after truncation (default 3).

§token_safety_margin: f32

Safety multiplier for token estimates (default 1.3).

§max_output_lines: usize

Max output lines for tool results.

§max_output_bytes: usize

Max output bytes for tool results.

§compression_token_threshold: usize

Token threshold for triggering compression.

§compression_enabled: bool

Whether compression is enabled.

§max_tool_calls_per_turn: usize

Maximum tool calls per turn (default 30).

§progressive_compression: bool

Whether progressive segmented compression is enabled (default true).

§rounds_per_summary: usize

Number of full rounds per summary segment (default 3).

§max_summary_segments: usize

Maximum number of summary segments to keep (default 5).

§merge_count: usize

Number of segments to merge at a time (default 2).

§max_merges_per_segment: usize

Maximum merges per segment before discarding (default 2).

Implementations§

Source§

impl ContextManager

Source

pub fn new(context_window: Option<u64>, config: Option<&ContextConfig>) -> Self

Source

pub fn truncation_threshold(&self) -> usize

Maximum tokens at which truncation is triggered. Uses truncation_ratio (default 0.7) to trigger earlier than the absolute limit, leaving headroom for estimation errors and LLM response.

Source

pub fn available_tokens(&self) -> usize

Maximum tokens available for input (total - reserved for response). Note: this is the absolute upper bound; truncation actually triggers earlier via truncation_threshold().

Source

pub fn truncate_tool_output(&self, content: &str) -> String

Truncate tool output using the configured limits.

Source

pub fn estimate_context_tokens( &self, messages: &[ChatCompletionRequestMessage], last_known_prompt_tokens: Option<u32>, snapshot_message_count: usize, ) -> usize

Hybrid token estimation: precise API anchor + incremental estimation.

When last_known_prompt_tokens is Some and history hasn’t been truncated (i.e. messages.len() >= snapshot_message_count), uses the API-reported prompt_tokens as an exact baseline and adds estimated tokens for only the new messages appended since the snapshot. This is far more accurate than full heuristic estimation because the baseline is from the model’s own tokenizer.

Falls back to full heuristic estimation with token_safety_margin when:

  • No calibration data yet (first call, before any API response)
  • History was truncated/compressed (messages.len() < snapshot)
Source

pub fn maybe_truncate( &self, messages: &mut Vec<ChatCompletionRequestMessage>, last_known_prompt_tokens: Option<u32>, snapshot_message_count: usize, ) -> TruncationResult

Check if history needs truncation and perform it if necessary. Returns TruncationResult with removed messages for async compression.

Strategy (progressive compression):

  1. Uses truncation_threshold() (default 70% of max_tokens) as trigger point.
  2. When over threshold, performs exactly one compression action per call:
    • Priority 1: Compress the oldest rounds_per_summary full rounds into a new summary segment.
    • Priority 2: Merge the oldest merge_count summary segments into one.
    • Priority 3: Discard the oldest summary segment (if merge limit reached).
    • Priority 4: Fall back to aggressive truncation (old behavior).
  3. If progressive_compression is disabled, falls back to single-shot truncation.

Auto Trait Implementations§

Blanket Implementations§

Source§

impl<T> Any for T
where T: 'static + ?Sized,

Source§

fn type_id(&self) -> TypeId

Gets the TypeId of self. Read more
Source§

impl<T> Borrow<T> for T
where T: ?Sized,

Source§

fn borrow(&self) -> &T

Immutably borrows from an owned value. Read more
Source§

impl<T> BorrowMut<T> for T
where T: ?Sized,

Source§

fn borrow_mut(&mut self) -> &mut T

Mutably borrows from an owned value. Read more
Source§

impl<T> From<T> for T

Source§

fn from(t: T) -> T

Returns the argument unchanged.

Source§

impl<T> Instrument for T

Source§

fn instrument(self, span: Span) -> Instrumented<Self>

Instruments this type with the provided Span, returning an Instrumented wrapper. Read more
Source§

fn in_current_span(self) -> Instrumented<Self>

Instruments this type with the current Span, returning an Instrumented wrapper. Read more
Source§

impl<T, U> Into<U> for T
where U: From<T>,

Source§

fn into(self) -> U

Calls U::from(self).

That is, this conversion is whatever the implementation of From<T> for U chooses to do.

Source§

impl<T> MaybeSend for T
where T: Send,

Source§

impl<T> PolicyExt for T
where T: ?Sized,

Source§

fn and<P, B, E>(self, other: P) -> And<T, P>
where T: Sized + Policy<B, E>, P: Policy<B, E>,

Create a new Policy that returns Action::Follow only if self and other return Action::Follow. Read more
Source§

fn or<P, B, E>(self, other: P) -> Or<T, P>
where T: Sized + Policy<B, E>, P: Policy<B, E>,

Create a new Policy that returns Action::Follow if either self or other returns Action::Follow. Read more
Source§

impl<T, U> TryFrom<U> for T
where U: Into<T>,

Source§

type Error = !

The type returned in the event of a conversion error.
Source§

fn try_from(value: U) -> Result<T, !>

Performs the conversion.
Source§

impl<T, U> TryInto<U> for T
where U: TryFrom<T>,

Source§

type Error = <U as TryFrom<T>>::Error

The type returned in the event of a conversion error.
Source§

fn try_into(self) -> Result<U, <U as TryFrom<T>>::Error>

Performs the conversion.
Source§

impl<T> WithSubscriber for T

Source§

fn with_subscriber<S>(self, subscriber: S) -> WithDispatch<Self>
where S: Into<Dispatch>,

Attaches the provided Subscriber to this type, returning a WithDispatch wrapper. Read more
Source§

fn with_current_subscriber(self) -> WithDispatch<Self>

Attaches the current default Subscriber to this type, returning a WithDispatch wrapper. Read more