tower-http-cache
Tower middleware for HTTP response caching with pluggable storage backends (in-memory, Redis, and more). tower-http-cache brings a production-grade caching layer to Tower/Axum/Hyper stacks, with stampede protection, stale-while-revalidate, header allowlisting, compression, and policy controls out of the box.
Features at a Glance
- β
Drop-in
CacheLayer: wrap any Tower service; caches GET/HEAD by default. - π Stampede protection: deduplicates concurrent misses and serves stale data while recomputing.
- β± Flexible TTLs: positive/negative TTL, refresh-before-expiry window, stale-while-revalidate.
- π Auto-refresh: proactively refreshes frequently-accessed cache entries before expiration.
- π¬ Chunk Caching: memory-efficient caching for large files with range request support.
- π·οΈ Cache Tags: group and invalidate related cache entries together.
- π― Multi-Tier: hybrid L1/L2 caching for optimal performance and capacity.
- π Admin API: REST endpoints for cache introspection and management.
- π€ ML-Ready Logging: structured logs with request correlation for ML training.
- π¦ Pluggable storage: in-memory (Moka) and Redis backends.
- π Policy guards: min/max body size, cache-control respect/override, custom method/status filters.
- π§° Custom keys: built-in extractors (path, path+query) plus custom closures.
- π Observability hooks: optional metrics counters and tracing spans.
Installation
[]
= "0.6"
# Enable Redis support if required
= { = "0.6", = ["redis-backend"] }
# With admin API support
= { = "0.6", = ["admin-api"] }
Quick Start
use Duration;
use ServiceBuilder;
use *;
let cache_layer = builder
.ttl
.negative_ttl
.stale_while_revalidate
.refresh_before
.min_body_size
.max_body_size
.respect_cache_control
.build;
let svc = new
.layer
.service;
Chunk Caching for Large Files
Efficiently cache and serve large files with byte-range support - perfect for video streaming:
use *;
use StreamingPolicy;
use Duration;
let cache_layer = builder
.policy
.build;
Benefits:
- 90% memory reduction for large file workloads
- Instant seeking for video streaming (no re-download)
- Range requests served directly from memory
- Only cache accessed chunks (partial file caching)
Example:
See examples/chunk_cache_demo.rs for a complete working example.
Using the Redis backend
use Duration;
use ConnectionManagerConfig;
use *;
async
Enabling Auto-Refresh
Auto-refresh proactively refreshes frequently-accessed cache entries before they expire, reducing cache misses and latency for hot endpoints:
use Duration;
use *;
use AutoRefreshConfig;
let cache_layer = builder
.ttl
.refresh_before
.auto_refresh
.build;
// Initialize auto-refresh with the service instance
cache_layer.init_auto_refresh.await?;
Using Cache Tags
Group related cache entries and invalidate them together:
use *;
use TagPolicy;
let cache_layer = builder
.policy
.build;
// Later: invalidate all entries with a tag
backend.invalidate_by_tag.await?;
backend.invalidate_by_tags.await?;
Tag-based invalidation works on InMemoryBackend, and on MultiTierBackend
over one. RedisBackend keeps no reverse tag index, so get_keys_by_tag,
list_tags and invalidate_by_tag return CacheError::Unsupported rather
than a silent Ok(0). Tags themselves do cross the Redis wire as of 0.6.0, so
a CacheRead from Redis carries the tags its entry was stored with. TagIndex
is also process-local, so invalidating a tag clears only the calling process's
index.
Multi-Tier Caching
Combine fast in-memory cache with larger distributed storage:
use MultiTierBackend;
let backend = builder
.l1 // Hot data (fast)
.l2 // Cold storage (large)
.promotion_threshold // Promote after 3 L2 hits
.promotion_strategy
.write_through
.build;
let cache_layer = builder
.ttl
.build;
Smart Streaming & Large File Handling
Automatically prevent large files from overwhelming your cache:
use StreamingPolicy;
let cache_layer = builder
.policy
.build;
Features:
- Automatic early detection via
Content-Lengthandsize_hint() - Content-Type based filtering (skip PDFs, videos, archives by default)
- Protects multi-tier caches (large files excluded from L1)
- Prevents memory exhaustion from large response bodies
- Fully configurable per content-type and size
Admin API
Enable cache introspection and management endpoints:
use AdminConfig;
let admin_config = new
.with_require_auth
.with_auth_token
.with_enabled;
// Available handler functions (wire into your Axum router):
// tower_http_cache::admin::routes::handle_health
// tower_http_cache::admin::routes::handle_stats
// tower_http_cache::admin::routes::handle_hot_keys
// tower_http_cache::admin::routes::handle_list_tags
// tower_http_cache::admin::routes::handle_invalidate
ML-Ready Structured Logging
Enable structured logging for ML model training:
use MLLoggingConfig;
let cache_layer = builder
.policy
.build;
// Logs will be emitted in JSON format:
// {
// "timestamp": "2025-11-10T12:00:00Z",
// "request_id": "550e8400-...",
// "operation": "cache_hit",
// "latency_us": 150,
// "tags": ["user:123"],
// "tier": "l1"
// }
Configuration Highlights
| Policy | Description |
|---|---|
ttl / negative_ttl |
cache lifetime for successful and error responses |
stale_while_revalidate |
serve stale data while a refresh is in progress |
refresh_before |
proactively refresh the cache shortly before expiry |
auto_refresh |
automatically refresh frequently-accessed entries before expiration |
tag_policy |
configure cache tags and invalidation groups |
multi_tier |
enable multi-tier caching with L1/L2 backends |
ml_logging |
enable ML-ready structured logging |
allow_streaming_bodies |
opt into caching streaming responses |
min_body_size / max_body_size |
enforce size bounds for cached bodies |
header_allowlist |
restrict which headers are stored alongside cached bodies |
method_predicate / statuses |
customize cacheable methods and status codes |
For the full API surface, see the generated docs: cargo doc --open.
Benchmarks
Benchmarks are powered by Criterion and can be reproduced with:
Run benchmarks locally only. Never add cargo bench to CI β Criterion
executes the whole suite, which takes minutes and produces timings that are
meaningless on shared runners. CI compiles the benches instead, via
cargo test --no-run --benches.
Latest results (macOS / M3 Pro / Rust 1.85, redis-backend disabled unless noted).
The codec/* rows were measured against 0.5.x's bincode codec; 0.6.0 replaced it
with postcard and renamed those benches to codec/postcard_*. They have not been
re-measured.
| Group | Benchmark | Median | Notes |
|---|---|---|---|
layer_throughput |
baseline_inner |
1.41 ms | Underlying service without caching |
cache_hit |
0.67 Β΅s | Cached GET; body already materialized | |
cache_miss |
0.68 Β΅s | Miss with immediate store | |
key_extractor |
path |
23.8 ns | GET/HEAD path only |
path_and_query |
97.4 ns | Path + query concatenation | |
custom_hit |
84.7 ns | User extractor returning Some |
|
custom_miss |
1.35 ns | User extractor returning None |
|
backend/in_memory |
get_small_hit |
309 ns | 1 KiB entry |
get_large_hit |
327 ns | 128 KiB entry | |
set_small |
676 ns | 1 KiB write | |
set_large |
660 ns | 128 KiB write | |
stampede |
cache_layer |
5.92 ms | 64 concurrent requests with caching |
no_cache |
5.76 ms | Same workload without layer | |
stale_while_revalidate |
stale_hit_latency |
33.6 ms | Serve-stale branch |
strict_refresh_latency |
33.7 ms | Force refresh branch | |
codec/bincode (0.5.x) |
encode_small |
362 ns | 1 KiB payload |
decode_small |
381 ns | 1 KiB payload | |
encode_large |
146 Β΅s | 128 KiB payload | |
decode_large |
174 Β΅s | 128 KiB payload | |
negative_cache |
initial_miss |
14.0 Β΅s | First miss populates negative entry |
stored_negative_hit |
21.9 ms | TTL-expired negative pathways | |
after_ttl_churn |
5.66 Β΅s | Subsequent positive hit |
Full raw output, including outlier analysis, is captured in initial_benchmark.md.
Testing & Tooling
# Library unit tests + integration tests
# Redis integration tests
REDIS_URL=redis://127.0.0.1:6379/
# Redis smoke test (launches example service, verifies cache hit/miss behaviour)
# Examples
Feature Flags
| Feature | Description | Default |
|---|---|---|
in-memory |
Enables the Moka-powered in-memory backend | β |
redis-backend |
Enables the Redis backend, codec, and async utilities | β |
admin-api |
Enables admin REST API endpoints (requires axum) | β |
serde |
Derives serde traits for cached entries/codecs |
β |
compression |
Adds optional gzip compression for cached payloads | β |
metrics |
Emits metrics counters (hit/miss/store/etc.) |
β |
tracing |
Adds tracing spans around cache operations | β |
legacy-bincode1-read |
Reads cache entries written by 0.5.x (see below) | β |
Upgrading from 0.5.x
The on-the-wire cache format changed
Entries in Redis are now written as a 21-byte versioned envelope β "THC" magic,
a format byte, a codec byte, and the expiry and stale timestamps as
little-endian u64 β followed by a postcard-encoded payload. 0.5.x wrote a
bare bincode 1 record. Tags are part of the payload now; they never crossed the
Redis wire before.
Upgrading does not cold-start your cache. 0.6.0 reads 0.5.x entries through
the legacy-bincode1-read feature, which is on by default. Entries are rewritten
in the new format as they are refreshed.
Rolling back to 0.5.x is also safe. A 0.5.x binary reading a 0.6.0 entry gets a clean decode error, which the cache layer already treats as a miss. The cost is a cold cache, not corrupted responses β that is what the envelope header buys. Without it, the old reader would have silently accepted the new bytes and ignored the trailing remainder.
BincodeCodec is renamed PostcardCodec. A deprecated alias keeps the old name
working through 0.6.x.
Turning legacy-bincode1-read off
The reader is hand-written against the bincode 1 layout and pulls no dependency, so leaving it on costs only dead code. It exists so 0.7.0 can delete it cleanly, not to dodge an advisory.
Turning it off is safe at any time and cannot lose data: an entry it would have read becomes a miss, the response is recomputed, and the entry is rewritten in the new format. The only cost is a colder cache while 0.5.x-written entries are re-populated. Since entries are self-expiring, once every 0.5.x entry has aged past its TTL plus its stale window the feature is doing nothing anyway.
= { = "0.6", = false, = ["in-memory", "serde"] }
The feature and the module behind it are removed in 0.7.0.
CacheBackend no longer uses #[async_trait]
The trait uses native async fn in traits (RPITIT). Every method is declared as
fn name(..) -> impl Future<Output = ..> + Send; the + Send is required
because the cache layer boxes backend futures into a Send future.
If you implement CacheBackend yourself, the migration is one line per impl β
delete the attribute. Method bodies are unchanged:
-#[async_trait]
impl CacheBackend for MyBackend {
async fn get(&self, key: &str) -> Result<Option<CacheRead>, CacheError> {
// unchanged
}
}
Leaving #[async_trait] in place produces error[E0195]: lifetime parameters or bounds on method 'get' do not match the trait declaration. The same one-line
change applies to any overridden default method (get_keys_by_tag,
invalidate_by_tag, invalidate_by_tags, list_tags).
CacheBackend was already non-dyn-compatible because of its Clone supertrait,
so no working code used it as a trait object. MSRV is unchanged β RPITIT
stabilised in Rust 1.75, well below this crate's floor.
Also breaking
CacheErroris#[non_exhaustive]and gained anUnsupportedvariant. Exhaustive matches need a_arm;CacheError::is_unsupported()avoids matching at all.RedisBackend::get_keys_by_tagandlist_tagsreturnUnsupportedinstead ofOk(vec![]), so the defaultedinvalidate_by_tagpropagates an error where it previously returnedOk(0).- The
memcached-backendfeature andMemcachedBackendare removed. - redis
0.32->1.6changedConnectionManagerConfig's defaults from no timeouts to a 500 ms response timeout and a 1 s connection timeout. See the Redis backend example above.
See CHANGELOG.md for the full list.
Minimum Supported Rust Version
MSRV: 1.85 for the default feature set, matching the crate's rust-version
field. redis-backend requires 1.88 β redis 1.6 declares it, and the feature
also reaches url -> idna -> icu_*, which require the same. Both floors are
enforced by separate CI jobs.
The MSRV will only increase with a minor version bump and will be documented in release notes.
Status
tower-http-cache is under active development. Expect API adjustments while we stabilize the 0.x series. Contributions and feedback are welcomeβfeel free to open an issue or PR! ***
License
This project is dual-licensed under either:
- Apache License, Version 2.0 (LICENSE-APACHE or http://www.apache.org/licenses/LICENSE-2.0)
- MIT License (LICENSE-MIT or http://opensource.org/licenses/MIT)
You may choose either license to suit your needs. Unless explicitly stated otherwise, any contribution intentionally submitted for inclusion in the crate shall be dual-licensed as above, without additional terms or conditions.
Contributing
- Fork and clone the repository.
- Install prerequisites (
cargo,rustup, and Docker if you plan to run Redis tests). - Run the checks:
- Open a pull request with a succinct summary, test evidence, and (when applicable) benchmark output via
cargo bench. Run benchmarks locally and paste the output; do not add a bench step to CI.
Bug reports and feature requests are welcome in the issue tracker. For larger design changes, please start a discussion thread to align on API shape before submitting code.