Two-layer cache: L1Cache is an exact-match TTL LRU, L2Cache is a per-tenant semantic cache over bge-small embeddings.
L1Cache
L2Cache