harpe
Deterministic context synthesis and token budgeting for AI agents.
Part of the Perseus suite by Perseus Computing LLC.
Why Harpe exists
In multi-turn agent runs, prompts grow until they hit model limits. When that happens, naive truncation usually lops off system rules, tool schemas, or early task context, causing the agent to hallucinate or stall.
Harpe solves this by enforcing hard token budgets before calling an LLM. It sorts context items by priority and applies deterministic eviction rules so your agents never exceed token limits or drop critical instructions.
Key design goals
- Microsecond speed: Median pruning latency of 1.72 µs across 128k context allocations adds zero noticeable overhead to agent turns.
- Predictable eviction: Policies like
LowestPriority,Fifo,TruncateTail, andLeastRelevantensure prompt assembly is fully reproducible. - Zero allocation steady state: Token accounting avoids heap churn during rapid tool-calling loops.
- No network calls: Pure Rust that runs entirely on your own machine.
Installation
Add harpe to your Cargo.toml:
[]
= "0.1.0-alpha.2"
Or via Cargo:
Example
use ;
Benchmark results
Measured on bare-metal Linux x86_64 (rustc 1.85.0, opt-level = 3, lto = "fat"):
| Metric | Measurement | Test condition |
|---|---|---|
| p50 Pruning Latency | 1.72 µs | 128k context ceiling |
| p95 Pruning Latency | 4.88 µs | High contention allocation |
| p99 Pruning Latency | 7.14 µs | Peak context window boundary |
| Ceiling Reduction | 69.50% | Multi-turn dynamic compression |
| Allocations / Cycle | 0 | Steady state pre-allocated pool |
Run benchmarks yourself:
License
Copyright © 2026 Perseus Computing LLC. Released under the MIT License.