# kkernel Design
**Last reviewed**: 2026-06-06
## ADR Compliance
### System Architecture (kernel/MCP split) (ADR-003)
- `kkernel` is the admin/management binary; `khive-mcp` is the MCP stdio server.
- They share the `khive-runtime` crate but have separate entry points.
- Pack crates do not depend on `kkernel` or its coordinator module — they receive
a single-backend `KhiveRuntime` from the MCP layer.
- The anti-pattern of packs depending on the coordinator is explicitly guarded by the
module boundary: `coordinator` is `kkernel`-internal.
### Multi-backend configuration (ADR-009, ADR-028)
- MCP/exec host boot constructs `BackendRegistry` from `khive.toml`, deduplicates
physical databases, and routes each pack to its assigned runtime. Canonical
`main` remains the sole attachment and blob-GC liveness authority after every
secondary passes boot inventory.
- `kkernel backend list` and `kkernel backend info` are a narrower legacy admin
surface: they currently expose only the synthetic default `main` backend, not
the full config-driven boot registry.
### VCS and sync (ADR-010, ADR-020)
- `kkernel sync` delegates to `khive_vcs::sync::run_sync` — the NDJSON-to-SQLite
rebuild logic is owned by `khive-vcs`, not by `kkernel`.
- `kkernel kg fetch` (alias: `kkernel kg sync`) fetches a remote KG archive and
populates the local remote cache under `.khive/kg/remotes/<remote>/`.
### ADR-015: Schema migrations
- Ordinary schema preparation applies the canonical prefix automatically.
V21 is application-assisted when legacy blob references exist: the async
MCP/kkernel host installs bounded hydration, authenticates pack-owned
artifacts, and finalizes attachment-only GC before constructing a runtime.
- `kkernel db migrate` loads the configured storage topology and reuses the
core-only host coordinator: physical aliases are deduplicated, distinct
secondaries are prepared before canonical `main`, and no pack runtime,
embedder, or pack-auxiliary DDL is constructed. `--backend` is an active
configured-backend selector; advancing `main` retains its secondary
prerequisites. `kkernel db check` loads the same target set and inspects each
migration ledger read-only without writing.
- The V21 rollout requires Phase-4a fleet convergence and old-binary drain,
followed by quiescence of every Phase-4a application reader/writer during
cutover. A GC-only Phase-4a worker's completed-V21 compatibility is narrower
than serving compatibility. Phase-4b serving starts only after the same
planned topology validates exact-current.
### Pack standard (vocabulary + handlers) (ADR-017)
- `pack_introspect` module builds an in-memory `VerbRegistry` from all `inventory!`-
registered packs and exposes `list_packs()` and `pack_handler(name)`.
- Handler `visibility` distinguishes MCP-exposed `Verb` entries from internal
`Subhandler` entries (e.g. `memory.recall_embed`).
### Verb namespace contract (ADR-023)
- The kg substrate pack owns 26 verbs: the bare names (no dot prefix) `create`, `get`,
`list`, `stats`, `update`, `delete`, `restore`, `search`, `link`, `neighbors`, `traverse`,
`query`, `merge`, `propose`, `review`, `withdraw`, `resolve`, `verbs`, `context`
(ADR-089), `whoami`, `scan`, `db_diagnostics` (ADR-091), plus its one documented
sub-namespace, `stream.append` / `stream.read` / `stream.stat` (ADR-174 §2).
- Every other pack must prefix verbs with `<pack>.` (e.g. `memory.recall`).
- Sub-variants use underscore, not nested dots: `memory.recall_embed`, not
`memory.recall.embed`.
- Enforced by the integration test in `tests/verb_namespace_contract.rs`.
### Dynamic pack loading (self-registration via inventory!) (ADR-027)
- Pack crates self-register using `inventory::submit!`. The linker drops crates whose
symbols aren't referenced, so `kkernel/lib.rs` and the contract test binary both
include explicit `use PackName as _` anchors to prevent dead-stripping.
- `PackRegistry::discovered_names()` returns all self-registered pack names at runtime.
### SubstrateCoordinator (ADR-029)
- `coordinator/mod.rs` implements D1 (BackendRegistry), D2 (LocatorCache), and D3
(fan-out search with RRF).
- D4 (cross-backend traversal), D5 (WAL cascade delete), and D6 (health map) are
deferred; sub-modules (`edges`, `traversal`, `curation`, `health`) are reserved.
- See `docs/coordinator.md` for implementation phase detail.
### KG validation and init (ADR-034, ADR-035)
- `kkernel kg validate` runs seven unconditional structural checks: six error-severity
input/schema/identity/reference/taxonomy checks plus warning-severity sort order. A conditional
error-severity note-kind check also runs when `notes.ndjson` is present. Configurable rules from
`rules.toml` run afterward.
- `kkernel kg init` creates `.khive/kg/` and writes `khive.toml` with defaults.
- See `docs/kg-rules.md` for the rule TOML format.
### KG status and fetch/sync alias (ADR-036, ADR-037)
- `kkernel kg status` computes a content hash of the DB state and the NDJSON files and
reports whether they match.
- `kkernel kg fetch` has a `visible_alias = "sync"` so `kkernel kg sync --repin <remote>`
reaches the same handler.
### Embedding model lifecycle (ADR-043)
- `kkernel engine list/status` expose `_embedding_models` table data.
- `kkernel engine migrate` and `kkernel engine drift-check` are deferred to follow-up
#2873 (EmbedMigrationWorker and lattice_transport integration).
- No MCP verbs are exposed for engine management — these are operator-only commands.
### Vector store capabilities and orphan sweep (ADR-044)
- `kkernel vector capabilities` emits the sqlite-vec baseline capability flags.
A comparison test pins every field to `SqliteVecStore::capabilities()` in `khive-db`,
including `supports_orphan_sweep: true`. It prints JSON by default or text with `--human`
without opening a database; `--engine` only labels this capability report.
- `kkernel vector sweep [--namespace <ns>...] [--max-delete <n>] [--dry-run]
[--engine <name>] [--db <path>]` calls the backend's orphan sweep for vectors whose subject
has no live entity, note, or knowledge atom. Soft-deleted subjects count as orphans.
- `--engine` selects an exact configured `[[engines]].name`; without configured engine entries,
it selects a canonical runtime model name. Omitting it sweeps all configured model stores,
once per distinct model even when several engine names share it.
- Repeated `--namespace` flags restrict the sweep; omitting them includes all namespaces.
`--max-delete` defaults to `1000` and is a shared deletion budget across selected stores.
Zero deletes no vectors, and values above `4294967295` fail before opening the database.
- `--dry-run` reports orphan counts without deleting vectors. Counting still runs under the
backend's writer transaction; the deletion budget does not limit scanning.
- JSON output contains `namespaces`, `dry_run`, `max_delete`, aggregate `scanned`, `deleted`,
`would_delete`, and `max_delete_hit`, plus a `stores` array. Each store identifies its
`engine_names`, `model`, and `namespaces` with those four result fields. `would_delete` counts
all orphans before the cap; `max_delete_hit` means that count exceeded the available budget.
### Proposal lifecycle (ADR-046)
- The kg pack exposes `propose`, `review`, and `withdraw` verbs as part of the
23 kg-substrate verbs (the bare names plus the `stream` sub-namespace). These are
validated by the contract test.
### Serial non-atomic ops-file dispatch (ADR-099 Amendment 4)
`exec.rs` owns an opt-in `--ops-file --serial` scheduling policy for backends
whose reader or model resources cannot safely serve multiple handlers at once.
The complete source is first validated into the same byte-bounded stable
snapshot, including a whole-snapshot typed-JSON structural preflight, before
the runtime is built or any operation can write. Execution retains one local
`KhiveMcpServer` and therefore one loaded model instance. Each existing
100-op/32 MiB logical chunk is parsed and dispatched as one complete parallel
batch through the same server path, but that batch's trusted scheduler cap is
one. The next handler starts only after the current handler completes.
This is scheduling, not chain or transaction semantics: `$prev` remains
invalid, operations commit independently, and progress/save/strict/
reconciliation, write-key conflict preflight, and aggregate response budget stay
on the existing logical chunk boundary. The default path remains bounded
parallel. `--serial` conflicts with `--atomic`, whose prepared whole-file
transaction is owned by `atomic_apply.rs` below.
The server's established request-read deadline remains scoped to the full
logical batch and is not silently renewed per operation. Long-running trusted
local model work must select the documented bounded
`KHIVE_REQUEST_READ_TIMEOUT_SECS` override explicitly; serial scheduling itself
does not weaken cancellation policy.
### Atomic `exec --ops-file --atomic` execution path (ADR-099 Slice B3)
`atomic_apply.rs` is the CLI-boundary orchestrator for `kkernel exec --ops-file --atomic`.
It runs, in order:
1. Parse-time admissibility (`khive_request::atomic::check_atomic_admissible`, B1), loaded-verb
resolution against a metadata-only in-memory registry for the configured pack set, and the
op-count guard. Configured-pack membership classifies only statically rejected operations; it
does not narrow the full discovered execution registry used by atomic prepare. This may build
an ephemeral in-memory runtime for authoritative pack metadata, but runs BEFORE opening or
touching the target database. Unknown/unloaded verbs are therefore distinguishable from loaded
verbs that are merely atomic-ineligible without scraping prose.
2. The async prepare pass: KG-substrate verbs via `khive_runtime::atomic_prepare::prepare_op`;
`gtd.transition`/`gtd.complete` via two `prepare_gtd_*` adapters in this module that wrap
the SAME decide functions (`khive_pack_gtd::handlers::prepare_transition`/
`prepare_complete`) the canonical non-atomic `handle_transition`/`handle_complete` call —
the decide logic is not duplicated, only adapted into an `AtomicOpPlan`. The adapters live
in `kkernel` (not `khive-pack-gtd`) because `AtomicOpPlan`/`PlanStatement` are
`khive-runtime` atomic-plan vocabulary this CLI orchestrator owns.
3. The synchronous commit pass (`khive_runtime::atomic_runner::run_atomic_unit`, B2).
4. The async post-commit reindex pass
(`khive_runtime::atomic_prepare::apply_post_commit_effects_with_report`).
Its typed embedding outcomes are matched back to the originating update plan, so an atomic
update whose embedding input was bounded carries the same per-result `warnings` advisory as
the canonical non-atomic handler.
The commit pass is the retry-safety boundary. Once it returns `Committed`, a reindex, canonical
result-rendering, or save-file publication failure cannot make the database mutation retryable:
the base DML is already durable. Reindex and rendering failures return success with
`atomic.committed=true`,
`atomic.status="committed_degraded"`, `atomic.retryable=false`, and typed entries in
`atomic.degradations` (`post_commit_reindex` or `result_rendering`). A render-degraded operation
keeps `ok=true` and the committed summary count, carries `result=null`, and repeats the typed
non-retry marker on that result entry. This prevents automation from replaying the mutation while
still making index repair or result re-read work explicit. An ordinary committed run and every
pre-commit/rollback shape remain unchanged.
`--save-file` preflights its sibling temp file before execution. On successful publication, the
stdout manifest preserves the envelope's complete top-level `atomic` block. A write, flush, or
rename failure after commit instead appends the `save_file_publish` degradation, prints the full
committed/non-retryable envelope to stdout, and then returns `Err` so the file request still exits
non-zero. That error is about the sink only; the stdout commit fact is authoritative and forbids
replay.
Preflight and prepare failures cross back to `exec.rs` as a typed `AtomicExecFailure` containing
the unchanged terminal message plus a result envelope over the real ops-file entries. The shared
refusal annotator emits `verb-refused` for unknown/unloaded preflight entries,
`gate-refusal` for typed secret-gate prepare failures, `policy-refusal` for structured
immutable stream record refusals, and `strict-op-failure` for otherwise
unclassified rollback entries when the operator supplied `--strict`.
The structured immutable stream policy guard also stamps
`domain_disposition: "not_committed"` beside the atomic result's string `error`,
because prepare refused before the atomic unit applied any domain write.
Before prepare, the atomic runtime installs the same aggregate pack edge rules as canonical
server startup. Before each task-note `update` plan is built, the boundary invokes
`VerbRegistry`'s shared `KindHook` normalizer/validator; dependency `link` plans run their
shared validator as well. This preserves typed and content/description-mirror parity with
canonical KG dispatch. An explicit update-kind mismatch is rejected before the hook runs.
The hook and plan share one note snapshot, and the plan's update statement rechecks that
snapshot revision/deletion marker in the transaction. Since every plan is still prepared before any write, core migration V15 supplies the
transaction-time task dependency cycle triggers: a later statement sees earlier writes in the
unit, and a trigger error rolls the entire unit back. The same triggers serialize opposite
concurrent writers, while leaving unrelated notes and edge relations untouched.
Atomic v1 does not maintain a projected note image across two generic updates of the same
target in one file. Both prepare from the original snapshot; after the first advances its
revision, the second observes zero affected rows, fails closed, and rolls the entire unit
back. This is deliberate until projected-state preparation can preserve all ordered patch
semantics without weakening concurrent-write guards.
**Verbs without a prepare implementation.** `propose`/`review`/`withdraw` are listed in
`khive_types::pack::ATOMIC_ADMISSIBLE_VERBS` (ADR-099 D3 intends them to eventually gain a
seam) but have none yet; the B3 fix rejects them at the same pre-target-runtime
`check_atomic_admissible` guard, as `AtomicRejectionReason::KnownUnimplemented`, before they
ever reach the target `KhiveRuntime::new`/`prepare_one` path. `prepare_op`'s own
`prepare_governance_unimplemented` fallback is unreachable through this CLI path and remains
only as defense-in-depth for other `prepare_op` callers.
`merge` joined this deferred bucket in the B3 fix too: a full-parity atomic merge prepare was
drafted and unit-tested against `atomic_prepare` directly, but its edge-conflict resolution
cannot be expressed in ADR-099's static predicate/guard plan shape, so it is rejected here
rather than shipped partially-scoped (`merge is not yet supported under --atomic; use the
non-atomic merge verb`).
The returned envelope is additive-only and lives entirely outside `dispatch_request_local`'s
response shape — non-atomic `--ops-file` runs (and every other exec path) are untouched.
**`apply_gtd_audit_post_commit_effects`** applies every `PostCommitEffect::GtdAudit` by
calling the SAME `ensure_audit_schema`/`write_audit_record_with_status` functions the
canonical `handle_transition`/`handle_complete` handlers call (GAP-5). It lives in `kkernel` rather
than `khive-runtime::atomic_prepare` because those two functions are owned by
`khive-pack-gtd`, which depends on `khive-runtime` — not the other way around; `kkernel` is
the first crate in the dependency graph that can see both. Non-`GtdAudit` effects are
ignored here (they are `atomic_prepare::apply_post_commit_effects`'s job, called
separately). Best-effort by construction: the callee logs append failures and returns a
boolean outcome. The atomic result builder exposes that outcome as `audit_persisted`, so the
already-committed unit cannot be failed while audit degradation remains visible.
An idempotent same-status `gtd.transition` is different: atomic v1 emits a guarded
no-effect assertion and no post-commit audit effect, even when the caller supplied
`note`. The assertion revalidates the exact prepare revision/deletion/status snapshot
inside the commit transaction, so an earlier same-unit transition, update, or delete
cannot make the no-op stale. It returns the base no-op shape and persists no note event.
Canonical dispatch owns the guarded note-bearing no-op behavior; callers needing it must
use that path until atomic projected-state support exists.
**`validate_atomic_args`** closes an ADR-099 B3 parity gap: the canonical (non-atomic)
handlers deserialize their args through a `#[serde(deny_unknown_fields)]` param struct, so a
typo like `conten` (for `content`) is rejected rather than silently ignored. The pre-fix
`--atomic` path had no equivalent gate — each `prepare_*` fn only read the keys it knew
about, so a typo'd key was dropped on the floor and the op reported `ok:true` with every
OTHER field reset to its current value, silently losing the caller's intended change. The
fix reuses (rather than reimplements) the canonical param structs: `kkernel` already depends
on `khive-pack-kg`/`khive-pack-gtd` directly, and their param structs
(`UpdateParams`/`DeleteParams`/`LinkParams`/`TransitionParams`/`CompleteParams`) are
re-exported `pub` specifically for this seam. Deserializing an op's args through the same
struct the canonical handler uses reproduces its `deny_unknown_fields` rejection and exact
error message for free, with no duplicated key list to drift out of sync — the deserialized
value itself is discarded; `prepare_*` still reads the raw `Value` map. `merge`, `create`,
and the read/governance verbs are out of scope: they are already rejected earlier at
`check_atomic_admissible`, or are not part of the v1 admissible set at all.
**`delete_expected_kind`/`update_expected_kind`** resolve a caller-supplied `kind=...`
string into the `AtomicDeleteKind`/`AtomicUpdateKind` enum `prepare_delete`/`prepare_update`
enforce, via the SAME canonical `resolve_kind_spec` the non-atomic `handle_delete`/
`handle_update` call. The two functions are exact mirrors of each other (same reasoning,
same shape, differing only in which `AtomicOpPlan`-adjacent enum they target). `kind` absent
resolves to `Ok(None)` (no check, parity with the canonical handlers' own optional
discriminator); `Event`/`Proposal` are a fail-loud rejection before `prepare_delete`/
`prepare_update` ever runs — those substrates are not v1-admissible for atomic
delete/update at all.
### `exec` local-dispatch fallback server (ADR-067, ADR-028 §8)
`build_local_fallback_server` (`src/exec.rs`) is the server constructor for both of
`kkernel exec`'s non-daemon dispatch paths: the daemon-unreachable/mismatch fallback inside
`run_exec_inline_with_forward`, and the `--ops-file` bulk-apply path (which deliberately
never attempts the daemon fast path at all — bulk apply needs cross-op atomicity the daemon
doesn't provide). `KhiveMcpServer::new` alone only ever builds a single-backend runtime, with
no visibility into a `khive.toml` `[[backends]]` declaration; before this fix, both
local-dispatch paths always used that single-backend constructor, so a config declaring a
separate backend for e.g. the `session` pack was invisible to them — the in-process fallback
silently wrote that pack's data into the `main` backend instead of its declared one. The fix
makes both paths agree with the daemon's own boot logic (`khive_mcp::serve::build_server`):
an empty `khive_cfg.backends` still takes the plain single-backend constructor (byte-identical
`config_id`, since `compute_config_id` skips the topology fold for an empty list); otherwise
both delegate to `build_server_multi_backend_with_db_anchor`, the same captured-anchor
constructor the production MCP boot path uses. `cli_db_override` is the raw, pre-resolution
`--db`/`KHIVE_DB` value, required for `--db :memory:` multi-backend override handling; passing
the wrong value would silently ignore an operator's in-memory isolation request. `db_anchor` is
the canonical anchor captured alongside `cfg`, threaded through so fallback construction never
re-reads a changed `HOME`.
Non-atomic `--ops-file` chunks cross into that server through
`dispatch_typed_json_batch_local_for_exec`. The JSONL reader has already enforced
its 96 MiB line, 512 MiB file, 32 MiB chunk, and 100-op ceilings, so reserializing
the decoded values and re-running the raw request parser would incorrectly apply
the unrelated 1 MiB MCP/HTTP/daemon/inline-string cap a second time. The typed seam
only replaces that redundant serialization/parse step: `parse_typed_json_batch`
still enforces JSON-form nesting, `$prev`, reserved-envelope, and count rules, and
the resulting `ParsedRequest` enters the same `run_parsed` identity, gate, audit,
presentation, strict-refusal, rendering, and response-envelope pipeline. Public
raw request paths continue to call `parse_request` and retain the 1 MiB limit.
### `exec` daemon-bypass second-writer contract (#548; ADR-067, ADR-099)
**Decision:** keep `exec --save-file` and both `exec --ops-file` modes on the in-process path.
They do not forward through, stop, or lease a live daemon. If that daemon has the same SQLite
file open, the CLI runtime is a second process with an independent connection pool. ADR-067's
writer task and `KHIVE_WRITE_QUEUE=1` serialize writes only inside one pool/process; they do not
coordinate the CLI process with the daemon. The boot/recovery guard in `exec.rs` protects runtime
construction and schema initialization only and is dropped before request execution.
Cross-process writes therefore use SQLite's WAL writer lock as their only serialization seam.
Every normal file-backed connection waits for that lock for `KHIVE_BUSY_TIMEOUT_SECS` (30 seconds
by default). If the other process still holds the lock when that process-local timeout expires,
the operation fails with SQLite `BUSY`/`database is locked`. `kkernel exec` does not automatically
retry: a generic replay is unsafe for non-idempotent verbs and could duplicate a write whose
outcome the caller has not established.
The observable failure contract stays mode-specific:
- `--save-file` may carry read or write verbs. A contended op's error detail is written to that
op's JSONL result row; stdout carries the save manifest, whose `summary` reports the failure
counts. Under `--strict`, otherwise-unclassified failures receive their stable refusal reason
before the JSONL rows are serialized and checksummed, so the manifest failure projection and
saved rows agree. `--strict` makes any failed op a non-zero exit; a fully failed invocation is
non-zero even without `--strict`.
- Non-atomic `--ops-file` preserves per-op commits and continues collecting results. A contended
op can fail after earlier ops committed; the final summary identifies every failed op.
`--strict` makes any failure a non-zero exit, while a fully failed file is always non-zero.
- Combined non-atomic `--ops-file --save-file` keeps those incremental database commits but makes
only destination-file publication atomic. Once a chunk has been dispatched, every termination
emits a reconciliation manifest. Success publishes the complete JSONL and preserves the ordinary
manifest shape. A failure before the manifest is finalized discards the incomplete temp file,
preserves any prior destination, and
emits an additive `status="aborted"` manifest naming the confirmed `committed_chunks` plus any
unverified `dispatched_chunk`; that unverified chunk may already have database effects. Its
`summary` covers confirmed rows and `unconfirmed_ops` accounts for the remainder without
classifying unknown outcomes as aborted. The manifest, not existence of a newly published result
file, is the recovery boundary. A batch that runs to completion publishes the ordinary manifest
first, so the later non-zero exits from `--strict` or from an all-failed file retain that
manifest rather than replacing it with an aborted one; every op outcome is known on those paths.
- `--ops-file --atomic` holds one bounded, DML-only `BEGIN IMMEDIATE` during its commit pass, per
ADR-099. Daemon writes wait behind that unit and can fail if its hold exceeds their own busy
timeout. Admissibility and prepare failures print a typed per-op result envelope over the real
ops-file entries before returning their established non-zero exit. A plan-level guard or SQL
failure likewise prints the additive envelope with `atomic.committed=false` and
`atomic.rolled_back=true`; the CLI prints that envelope and exits non-zero for the rollback,
with or without `--strict`. Under `--strict`, otherwise-unclassified not-committed rows receive
the stable `strict-op-failure` reason. A durable `committed_degraded` unit still exits zero. Callers
must inspect `atomic.committed` before retrying. The op-count guard bounds the unit; operators should run a
large atomic file against an idle daemon or in a maintenance window.
Routing these modes through the daemon remains rejected. `--save-file` is a trusted local output
sink not represented in `DaemonRequestFrame`; bulk apply deliberately avoids the daemon frame cap,
and atomic bulk apply would require a new daemon transaction protocol. Refusing based on a daemon
liveness probe would be racy. Any future daemon-routed design must be an explicit protocol change,
not a liveness-dependent fallback. Until then, operators configure busy timeouts independently in
both processes, inspect the manifest and the result file when one was published, and retry only
operations known to be idempotent.
### `exec.rs` regression test notes
- `strict_mode_rejects_before_daemon_forward_when_comm_and_no_actor`: `run_exec_inline` must
enforce the strict-actor gate (`KHIVE_REQUIRE_ATTRIBUTED_ACTOR=1`) BEFORE forwarding to the
daemon. Prior to this fix, `enforce_strict_actor_mode` was only called in the in-process
fallback path (after the daemon fast-path returned), so an attacker or misconfigured
operator could start a no-actor daemon, then run strict-mode `kkernel exec`, which would
forward through it and exit 0 — bypassing the gate. The fix moves the check before the
daemon block.
- `atomic_update_null_and_type_semantics_match_canonical_no_op_behavior`: atomic `update`
null/type semantics must match canonical's actually-reachable behavior. Empirically
verified against live `handle_update` that `name=null`/`description=null` are canonical
no-ops, not rejections — canonical's field type is `Option<Value>`, and serde's derived
`Deserialize` for `Option<T>` intercepts a literal JSON `null` at the outer `Option`
boundary and maps it straight to `None` regardless of the inner type, so canonical's own
literal-null branches in `string_value`/`description_patch` are unreachable through normal
struct deserialization. This test deliberately does NOT implement the naive expectation
("`update(name=null)` REJECTED") since that doesn't match the live system. What canonical
DOES still reject is a non-null, non-string `name` (e.g. `name: 123`) — pre-fix, atomic
silently treated that as absent too, reporting success for an invalid update.
### `reindex` memory Vamana epoch protocol (#812, ADR-107 §4)
`begin_reindex_epoch` (`src/reindex.rs`) durably marks the reindex-in-progress epoch
BEFORE any vector mutation in the pass. The previous design bumped the durable epoch only
as a best-effort step AFTER every vector mutation had already committed; a crash between the
last commit and that bump — or a silently swallowed bump error — left an already-warm daemon
with no durable signal at all, serving pre-reindex vectors indefinitely. Bumping first closes
that gap: a daemon that observes this epoch mid-reindex (via
`khive_pack_memory::ann::maybe_check_durable_epoch`, sampled from the recall path) rebuilds
conservatively against whatever partial corpus is on disk at that moment — never worse than
trusting a stale index forever — and the completion bump at the end of the pass forces one
more rebuild once the corpus reaches its final, fully re-embedded state. Together the two
durable bumps form an in-progress/completed epoch protocol: any observer landing anywhere
between them still converges. `kkernel reindex` runs directly against a raw `KhiveRuntime`
without a pack-registry boot applying `MemoryPack::SCHEMA_PLAN`, so this function also
ensures the `memory_ann_epoch` table exists before its first bump.
**Fail-closed, not warn-and-continue**: an error here (schema creation OR the epoch write
itself) aborts the whole reindex before any mutation runs. A swallowed failure here is
exactly the bug ADR-107 §4 was written to close.
Namespace Vamana snapshot invalidation is also part of the graph reindex result.
A failure to acquire its writer or delete matching snapshots sets
`vamana_snapshot_invalidation_failed: true` in the JSON report and names the failed
stage in the human report. The default command result is nonzero; `--best-effort`
retains the flag and warning while allowing a zero exit. Earlier FTS/vector writes
remain committed, and the command still attempts the memory completion epoch and
remaining passes. An absent `retrieval_snapshots` table remains a successful no-op.
This does not change the separate active-memory snapshot deletion's best-effort,
defense-in-depth role or its completion epoch's existing failure reporting.
## Consistency Notes
- `kkernel db migrate --dry-run` delegates to `cmd_db_check` rather than implementing
a separate dry-run path. Both reuse the migration target planner; the `--check`
flag exits nonzero for any behind, ahead, or invalid target.
- The `cmd_vector_capabilities` function hard-codes baseline sqlite-vec values
rather than instantiating a runtime — this is intentional for the v1 implementation
but means the output does not reflect operator-configured backends.
- `coordinator/mod.rs` cannot be split into sub-files because `SubstrateCoordinator`
has a `#[cfg(test)]` field (`fail_backend_id`) that tests access directly; making it
`pub(crate)` would expose the failure-injection mechanism to integration tests.
- Edge weights are validated in the [0.0, 1.0] closed interval with finite-number
checks. NaN and infinity are explicitly rejected with descriptive error messages.