Skip to main content

Module compression

Module compression 

Source
Expand description

Context compression for long agent conversations.

Long tool-heavy conversations balloon the message list sent to the LLM on every turn, slowing each call down and eventually exceeding the context window. This module provides a CompressionMiddleware that observes the per-LLM-call message list via Middleware::on_pre_llm and, once estimated tokens exceed a configurable threshold, summarises the earlier portion into a compact handoff message.

§Strategy

The compressor uses a hybrid retention approach:

  1. System prompt — always preserved verbatim.
  2. Recent N messages — kept verbatim (including tool results), placed at the end to leverage the model’s recency bias.
  3. Older history — condensed into a single handoff summary by an LLM call, cached via a stable-prefix hash to avoid redundant re-summarisation.

This balances information preservation (recent tool data stays intact) with token efficiency (older context is compressed).

§Architecture

session.chat_messages().to_vec()          ← full history clone
    ↓
CompressionMiddleware::on_pre_llm()       ← gate → split → cache check → summarise → assemble
    ↓
[system prompt] + [summary] + [recent N]  ← compressed copy for this LLM call only

The session’s stored history is never modified by automatic compression; only the per-call message copy is trimmed. The /compact CLI command can optionally write the compressed form back to the session.

§Feature gate

This module is behind the compression Cargo feature.

Re-exports§

pub use events::CompressionEvent;
pub use events::CompressionTrigger;

Modules§

events
Typed compression lifecycle events.

Structs§

AutoCompressionPolicy
Default policy: always compress when threshold is exceeded.
CompressionConfig
Tuning knobs for context compression.
CompressionMiddleware
Middleware that compresses conversation history before each LLM call.
ContextCompactor
Core compressor implementing the hybrid retention strategy.
RateLimitPolicy
Rate limiting policy: compress at most once every N seconds.
UserConfirmationPolicy
User confirmation policy: prompt user before compression.

Constants§

SUMMARY_PREFIX
Prefix injected in front of every summary produced by the compressor.

Traits§

CompressionPolicy
Compression policy trait — controls whether compression should proceed.

Functions§

is_summary_message
Returns true if msg is a user message whose content starts with SUMMARY_PREFIX — i.e. a summary injected by a previous compression pass.
language_instruction
Detect the dominant script of a text and return the appropriate language instruction for the summarisation prompt.
safe_cut_index
Walk cut backward until the boundary is tool-pairing safe.
serialize_block
Serialise a block of messages into a compact human-readable transcript.
split_system_prompt
Split a message slice into (system_messages, rest).
summarize
Summarise a transcript block via the LLM.
truncate_str
Truncate a string to max_chars, appending … if truncated.
truncate_summary_output
Truncate a summary to max_chars (front 80 % + rear 20 %).