Expand description
Context compression for long agent conversations.
Long tool-heavy conversations balloon the message list sent to the LLM on
every turn, slowing each call down and eventually exceeding the context
window. This module provides a CompressionMiddleware that observes the
per-LLM-call message list via
Middleware::on_pre_llm and, once
estimated tokens exceed a configurable threshold, summarises the earlier
portion into a compact handoff message.
§Strategy
The compressor uses a hybrid retention approach:
- System prompt — always preserved verbatim.
- Recent N messages — kept verbatim (including tool results), placed at the end to leverage the model’s recency bias.
- Older history — condensed into a single handoff summary by an LLM call, cached via a stable-prefix hash to avoid redundant re-summarisation.
This balances information preservation (recent tool data stays intact) with token efficiency (older context is compressed).
§Architecture
session.chat_messages().to_vec() ← full history clone
↓
CompressionMiddleware::on_pre_llm() ← gate → split → cache check → summarise → assemble
↓
[system prompt] + [summary] + [recent N] ← compressed copy for this LLM call onlyThe session’s stored history is never modified by automatic compression;
only the per-call message copy is trimmed. The /compact CLI command can
optionally write the compressed form back to the session.
§Feature gate
This module is behind the compression Cargo feature.
Re-exports§
pub use events::CompressionEvent;pub use events::CompressionTrigger;
Modules§
- events
- Typed compression lifecycle events.
Structs§
- Auto
Compression Policy - Default policy: always compress when threshold is exceeded.
- Compression
Config - Tuning knobs for context compression.
- Compression
Middleware - Middleware that compresses conversation history before each LLM call.
- Context
Compactor - Core compressor implementing the hybrid retention strategy.
- Rate
Limit Policy - Rate limiting policy: compress at most once every N seconds.
- User
Confirmation Policy - User confirmation policy: prompt user before compression.
Constants§
- SUMMARY_
PREFIX - Prefix injected in front of every summary produced by the compressor.
Traits§
- Compression
Policy - Compression policy trait — controls whether compression should proceed.
Functions§
- is_
summary_ message - Returns
trueifmsgis a user message whose content starts withSUMMARY_PREFIX— i.e. a summary injected by a previous compression pass. - language_
instruction - Detect the dominant script of a text and return the appropriate language instruction for the summarisation prompt.
- safe_
cut_ index - Walk
cutbackward until the boundary is tool-pairing safe. - serialize_
block - Serialise a block of messages into a compact human-readable transcript.
- split_
system_ prompt - Split a message slice into
(system_messages, rest). - summarize
- Summarise a transcript block via the LLM.
- truncate_
str - Truncate a string to
max_chars, appending…if truncated. - truncate_
summary_ output - Truncate a summary to
max_chars(front 80 % + rear 20 %).