1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
//! Context compression for long agent conversations.
//!
//! Long tool-heavy conversations balloon the message list sent to the LLM on
//! every turn, slowing each call down and eventually exceeding the context
//! window. This module provides a [`CompressionMiddleware`] that observes the
//! per-LLM-call message list via
//! [`Middleware::on_pre_llm`](agent_base::Middleware::on_pre_llm) and, once
//! estimated tokens exceed a configurable threshold, summarises the *earlier*
//! portion into a compact handoff message.
//!
//! # Strategy
//!
//! The compressor uses a **hybrid retention** approach:
//!
//! 1. **System prompt** — always preserved verbatim.
//! 2. **Recent N messages** — kept verbatim (including tool results), placed at
//! the end to leverage the model's recency bias.
//! 3. **Older history** — condensed into a single handoff summary by an LLM
//! call, cached via a stable-prefix hash to avoid redundant re-summarisation.
//!
//! This balances information preservation (recent tool data stays intact) with
//! token efficiency (older context is compressed).
//!
//! # Architecture
//!
//! ```text
//! session.chat_messages().to_vec() ← full history clone
//! ↓
//! CompressionMiddleware::on_pre_llm() ← gate → split → cache check → summarise → assemble
//! ↓
//! [system prompt] + [summary] + [recent N] ← compressed copy for this LLM call only
//! ```
//!
//! The session's stored history is **never** modified by automatic compression;
//! only the per-call message copy is trimmed. The `/compact` CLI command can
//! optionally write the compressed form back to the session.
//!
//! # Feature gate
//!
//! This module is behind the `compression` Cargo feature.
pub use ;
pub use CompressionConfig;
pub use ;
pub use ;
pub use CompressionMiddleware;
pub use ;
pub use ;