a3s-memory
Pluggable memory storage for A3S.
Provides the MemoryStore trait and two default implementations. Agents that need to persist and recall knowledge across sessions depend on this crate directly — nothing else required.
The crate also provides a separate VectorIndex capability for caller-owned
semantic retrieval. Its in-memory backend is dependency-free; the optional
SQLite backend preserves index history and revision CAS across process
restarts. Vector indexes are not memory stores: callers remain responsible for
document admission, embedding generation, lifecycle, and result fusion.
Default stores also enforce a small amount of memory hygiene: normalized exact duplicates are merged into the existing item, tags/metadata/importance are consolidated, and pruning protects curated memories such as pinned, frequently recalled, consolidated, or conflict-tracking items. Semantic equivalence is not inferred from keyword overlap at the storage layer.
Design
The crate follows a minimal core + external extensions pattern:
Core (stable, non-replaceable):
MemoryStore— storage backend traitMemoryItem— the unit of memoryMemoryType— episodic / semantic / procedural / workingRelevanceConfig— scoring parameters
Extensions (replaceable via MemoryStore):
InMemoryStore— default, ephemeral (testing and non-persistent use)FileMemoryStore— persistent, atomic writes, in-memory index
Three-tier session memory (AgentMemory) and context injection (MemoryContextProvider) live in a3s-code, not here. This crate only owns the storage layer.
Durable repository V2
The additive repository module provides the policy-free integrity kernel for
evidence-backed durable memory. It adds exact tenant/principal/scope
namespaces, candidate-to-active lifecycle transitions, immutable evidence
references, typed relations, complete revision history, optimistic concurrency,
and idempotent atomic change sets. Candidate activation requires new decision
evidence so an LLM annotation cannot silently become serving state. Reads are
pure; admission and use are recorded explicitly against the exact node revision
observed by the host. Hosts rebuilding derived projections can request a
complete bounded namespace snapshot: the repository filters an exact status
set, orders nodes deterministically, streams a stable SHA-256 view identity,
and rejects a node- or canonical-byte-over-budget view instead of silently
truncating it. Snapshot responses can be recomputed against their original
request, so hosts need not trust a custom backend's claimed digest.
Backends may additionally opt into
MemoryRepository::namespace_change_token. Equal bounded tokens prove that
no novel successful change set altered that exact namespace between two reads,
allowing callers to avoid a redundant snapshot without weakening their normal
snapshot and publication proofs. Tokens contain only a versioned monotonic
sequence; they do not contain namespace identifiers or memory content.
InMemoryRepository updates the sequence at the same write-lock
linearization point as node state, and FileMemoryRepository reconstructs the
same sequence by journal replay. Idempotent replay, failed changes, admission,
and use do not advance it. Custom backends return None unless they explicitly
implement the contract.
The deterministic V2 lexical profile preserves lowercased alphanumeric words
and adds overlapping bigrams for contiguous Chinese, Japanese, and Korean text.
This makes ordinary same-language CJK phrase variation searchable without a
model or egress while avoiding noisy single-character matching. Hosts can bind
MEMORY_LEXICAL_QUERY_PROFILE_V1 to detect algorithm drift. The profile is not
cross-language semantic retrieval: translated or no-overlap paraphrases still
require a caller-owned evaluated retrieval extension.
InMemoryRepository is the executable reference implementation.
FileMemoryRepository adds local durability through a checksummed write-ahead
journal: validated operations are appended and synced before publication,
replayed deterministically after restart, and protected by a single-writer
directory lock. A torn final record is discarded; checksum corruption in a
committed record fails closed.
The existing MemoryItem and MemoryStore API remains source-compatible while
hosts migrate to V2. Extraction, consolidation policy, embeddings, context
admission, and scheduling remain owned by the host runtime. See
docs/MEMORY_KERNEL_V2.md for invariants and release
gates.
Usage
[]
= { = "0.1", = "../memory" }
Store and retrieve
use ;
use Arc;
let store = new;
let item = new
.with_importance
.with_tag
.with_type;
store.store.await?;
let results = store.search.await?;
Persistent storage
use ;
let store = new.await?;
// Directory layout:
// memory/
// index.json ← in-memory index, persisted atomically
// items/{id}.json ← one file per memory item
Custom backend
Implement MemoryStore to use any storage system (SQLite, vector DB, etc.):
use ;
Caller-owned vector search
InMemoryVectorIndex stores caller-supplied vectors in immutable partition
snapshots and performs exact bounded top-k search. It does not use SQLite,
persist data, call an embedding model, or spawn a background task. The index is
released when its final owner is dropped.
use ;
let index = new?;
index
.replace_partition_if_revision
.await?;
let result = index
.search
.await?;
Dimensions are selected at index construction. Cosine indexes normalize
records and queries on admission, reject zero/non-finite vectors, and return
the immutable index revision that produced each result page. Replacing one
partition atomically publishes its complete new record set while sharing all
unchanged partition blocks. InMemoryVectorIndex additionally advertises
index_revision_cas: conditional replacement and removal compare the expected
global index revision at the same linearization point as publication. A delayed
writer therefore fails with RevisionConflict instead of overwriting or
deleting a newer generation. Custom backends remain source-compatible and
default to partition_atomic; their conditional methods fail closed until the
backend implements the CAS contract. Because the precondition is the global
index revision, an unrelated partition mutation can conservatively reject a
prepared update.
Backends may also expose VectorIndex::change_token() as exact continuity
evidence for one index history and revision. InMemoryVectorIndex assigns a
fresh opaque SHA-256 history digest when it is constructed; clones retain it and
every effective mutation advances the token revision. Two independently
constructed indexes therefore have different tokens even when their counters
and byte sizes happen to match. Custom indexes return None by default. A
durable backend may preserve a history digest across process restarts only
while it can prove the same linear mutation history; recreation, rollback, or
divergent restore requires a new identity. The token is not a content snapshot,
distributed lease, or remote-durability claim.
Correctness-sensitive callers should use the asynchronous
VectorIndex::observe() method. It returns one self-consistent status/token
pair and can report storage failures. The synchronous status() and
change_token() methods remain compatibility views and can be stale for a
durable or remote implementation.
With the sqlite feature, SqliteVectorIndex::open(path, descriptor) adds a
locally durable exact-search backend. It stores the descriptor, opaque history
identity, revision, stable logical byte accounting, partition integrity
digests, and vector rows in one SQLite database. Mutations use IMMEDIATE
transactions, so independent processes share one global revision-CAS
linearization point. Reopen validates the descriptor, accounting, record
shape, and integrity digests before serving; mismatches and corruption fail
closed. On Unix and Windows, the history identity is also bound to the database
file identity: copying or atomically replacing the file forks the token on its
next open without changing its content revision. Backup restore must replace
the closed database file; in-place overwrite and concurrent out-of-band file
operations are unsupported. Other targets conservatively fork the token on
every open when a stable file identity is unavailable. Blocking SQLite work
runs on Tokio's blocking pool. This is local durability and fencing, not a
distributed lease or remote replicated store.
use ;
let index = open
.await?;
let observation = index.observe.await?;
Run the locked release qualification for 25,000 records at 384 dimensions with
cargo run --example vector_search_benchmark --release. It emits JSON evidence
and fails when exact top-20 search exceeds the 30 ms p95 budget.
Relevance scoring
Search combines lexical match strength (exact phrase, term, tag, and memory-type matches) with the relevance score below. Exact or more specific query matches are kept ahead of generic high-importance memories, while equally specific results still benefit from importance and recency.
score = importance × importance_weight + decay × recency_weight
decay = exp(−age_days / decay_days)
Default: importance_weight = 0.7, recency_weight = 0.3, decay_days = 30.
use ;
let config = RelevanceConfig ;
let score = item.relevance_score_at;
Deduplication and pruning
InMemoryStore, FileMemoryStore, and the optional SQLite store collapse exact
durable duplicates after normalizing case and whitespace. Punctuation remains
significant. The first memory id remains canonical; later duplicates raise
importance, merge tags and list-style metadata such as supersedes /
conflicts_with, and record duplicate_count metadata.
Use MemoryStore::store_and_return() when the caller needs the canonical item
that now represents the fact. Semantic consolidation belongs to an upstream
model or caller with enough context to make that judgment. Such callers can use
MemoryItem::merge_duplicate() explicitly, or persist relation metadata and
let the owning memory runtime apply it.
PrunePolicy removes old, low-importance items and can enforce a maximum item
count, but it hard-protects curated memories: keep / pinned / protected
tags or metadata, repeatedly accessed items, and memories carrying
supersedes / conflicts_with relation metadata.
What this crate does NOT own
| Concern | Lives in |
|---|---|
| Three-tier session memory (working / short-term / long-term) | a3s-code |
MemoryConfig (max_short_term, max_working) |
a3s-code |
MemoryStats |
a3s-code |
| Context injection into agent prompts | a3s-code |
| Workspace scanning, code chunking, embeddings, and hybrid ranking | a3s-code |
Tests
The test suite covers MemoryItem, RelevanceConfig, InMemoryStore,
FileMemoryStore, and the reusable V2 repository contract. Both V2 backends
run the same behavior suite; file tests additionally cover restart recovery,
single-writer locking, torn writes, corruption, and durable concurrency. V2
tests cover
namespace isolation, evidence admission, idempotent replay, atomic rollback,
revision preservation, pure queries, explicit usage records, bounded input,
complete namespace-snapshot identity and overflow behavior, deterministic
word/CJK-bigram retrieval across both V2 backends, and concurrent writers.
The optional namespace change-token suite additionally covers atomic
advancement, namespace isolation, idempotent and failed writes, access-event
stability, concurrent single-winner updates, restart reconstruction, tamper
rejection, redaction, and source-compatible custom-backend opt-in.
The vector suites prove atomic observation, clone continuity,
effective-mutation advancement, no-op stability, serialization validation, and
distinct histories for independently constructed indexes with colliding
logical status. Enabling the sqlite feature additionally proves restart
continuity, cross-connection single-winner CAS, stable cross-backend byte
accounting, descriptor-drift rejection, content-integrity checks, and
fail-closed corruption handling, as well as the SQLite V1 store contract.
The sqlite-vec gate separately proves that concurrent first connections
register the vec0 auto-extension before either SQLite connection is opened.
Community
Join us on Discord for questions, discussions, and updates.
License
MIT