nedb-engine 3.1.0

NEDB v2 — content-addressed DAG storage engine with NQL and HTTP server (nedbd binary)
Documentation

nedb-engine

NEDB v2 — content-addressed DAG storage engine with NQL and HTTP server

crates.io License: MIT GitHub

This crate ships the nedbd binary — the NEDB v2 DAG HTTP server. Install it, point it at a data directory, and any language can speak to it over HTTP/JSON.

cargo install nedb-engine
nedbd ./data              # AOF engine (pure Rust, v1-compatible)
nedbd --dag ./data        # DAG engine (v2, content-addressed, recommended)

⚠️ New in 2.8.6 — Durability & Recovery (read this if you store anything you care about)

Three defects found by killing a real engine at every persistence boundary and by filling a real filesystem to zero free blocks. All three are fixed. If you are on 2.8.5 or earlier, upgrade.

1. A failed flush silently discarded acknowledged writes

IdIndex::flush_write_buf cleared every buffered entry regardless of whether its disk write succeeded. So a flush that hit ENOSPC threw the entry away, and no later flush retried it.

Reproduced on a full 22 MiB filesystem: 30 rows acknowledged by put() -> Ok, then list() returned 0 after reopen — while verify() reported all 30 objects healthy. The content-addressed objects were durable; the id-index entries that make them findable were gone.

before:  try_flush_all() -> (no return value)      reopen -> 0 rows, verify() = 30 ok
after:   try_flush_all() -> Err("id-index leaf rows/buf_25: No space left on device (os error 28)")
         ...free space, retry -> Ok                reopen -> 30 rows

Fixed: an entry leaves the WAL only when its write actually landed. Failures stay buffered and retry on the next flush.

2. Flush errors were unobservable — new try_flush_all()

flush_all() returns () and logged fsync failures to stderr, so a caller could not tell a durable flush from a failed one. Anything that takes a destructive or externally-visible action on the strength of a persisted record needs to know.

// Use this when the outcome matters:
db.try_flush_all()?;      // Result<()> — id-index WAL + segment sync + MANIFEST

// Still available, still logs, nowhere to propagate (ticker / Drop):
db.flush_all();

Also new: Db::try_flush_manifest() and IdIndex::try_flush_write_buf().

3. repair could not repair, and since() claimed "caught up" while behind

The cold scan rebuilt seq_index, per-collection tips, the Merkle head and MANIFEST — but never the id index. A database whose WAL never reached disk came back with every object verifying and list() empty, and nedb-cli repair printed success without fixing it, because start_cold_scan() is a deliberate no-op on a warm store.

nedb-cli repair ./data
# repaired: 203 id-index entr(ies) rebuilt, 203 node(s) verified, flushed
let restored = db.repair()?;   // rebuild id index from objects; highest seq wins

Every object carries its own coll, id and seq, so the id index is fully derivable — a lost WAL is recoverable and nothing is invented. repair() also recomputes head and tips, so a repaired database reopens warm instead of coming back up cold with an empty head.

Separately, since() set has_more = hit_limit alone. On a warm boot the seq index is empty by design (that is why warm start is O(1)), so every lookup missed and since() returned zero nodes with has_more = false — indistinguishable from genuinely up to date. A consumer following the documented drain loop stopped one call in, with every record unread.

Fixed: has_more is true whenever the cursor is behind the log head. ScanStatus gains seq_index_ready — replication consumers should gate on that, not on scan_complete, which is true on a warm boot precisely because the scan was skipped.

Known sharp edge (documented, not changed)

since()'s cursor is exclusive and seqs start at 0, so since(0, _) returns (0, head] and the very first write in a database (seq 0) is unreachable through any cursor value. Ten writes drain as nine records. Changing the convention would break existing consumers; a replica seeded from since() alone starts one record short.


What is NEDB v2?

NEDB v2 replaces the append-only log (AOF) with a content-addressed Merkle DAG:

  • Every document version is an immutable, BLAKE2b-hashed object. Nothing is ever overwritten.
  • Every write produces a new chain head — a BLAKE2b commitment over the entire database history.
  • Time-travel: read any document AS OF seq N to see exactly what it contained at that point.
  • Causal provenance: documents link to their causal parents via caused_by hashes. TRACE caused_by walks the full causal graph backward.
  • TRAVERSE: named graph relations via __links__. FROM person WHERE _id = "robert" TRAVERSE parent_of returns linked nodes.
  • O(1) warm start: a MANIFEST file stores seq + Merkle head so restarts never re-scan the entire object store.
  • Instant cold start: the daemon accepts connections immediately; background scan loads objects incrementally.
  • AES-256-GCM at rest: optional symmetric encryption with a double-envelope key structure (TMK wraps DEK).

Install

cargo install nedb-engine

The nedbd binary lands in ~/.cargo/bin/. Make sure that's on your PATH.

To verify:

nedbd --doctor      # diagnose your NEDB environment

Usage

nedbd [OPTIONS] [data_dir]
Argument Default Description
data_dir ./nedb-data Directory for database files
--dag off Use the v2 content-addressed DAG engine
--cast off Enable natural-language query planning (requires --features cast)
--doctor Diagnose environment, print fix commands

Environment variables

Variable Default Description
NEDBD_HOST 127.0.0.1 Bind address
NEDBD_PORT 7070 HTTP port
NEDBD_TOKEN Bearer token for auth (optional)
NEDB_TMK 64-char hex master key for AES-256-GCM encryption at rest
NEDBD_MEMORY 1 = pure in-memory mode (no disk I/O)
NEDBD_DAG 1 = same as --dag flag
NEDBD_CAST 1 = same as --cast flag
NEDBD_CAST_MODEL Explicit path to a model.cast container

HTTP API

All endpoints return JSON. Auth: Authorization: Bearer <token> if NEDBD_TOKEN is set.

GET    /health
GET    /v1/databases
POST   /v1/databases                    {name, init?}
GET    /v1/databases/<name>
DELETE /v1/databases/<name>
POST   /v1/databases/<name>/put         {coll, id, doc, caused_by?}
POST   /v1/databases/<name>/query       {nql}
POST   /v1/databases/<name>/link        {frm, rel, to}
POST   /v1/databases/<name>/neighbors   {node, rel}
POST   /v1/databases/<name>/cast        {prompt, execute?}    # feature = "cast"
GET    /v1/databases/<name>/verify
POST   /v1/databases/<name>/checkpoint

Cast — natural language into NQL

Optional. Feature-gated, off by default.

curl -X POST localhost:7070/v1/databases/shop/cast \
  -H 'Content-Type: application/json' \
  -d '{"prompt":"orders over 100"}'
{
  "prompt": "orders over 100",
  "nql": "FROM orders WHERE total > 100",
  "valid": true,
  "collection": "orders",
  "collection_known": true,
  "collections": ["orders"],
  "executed": false,
  "seq": 3,
  "head": "262fd9…"
}

A 3.33M-parameter model (nedb-cast-slm) turns a short English prompt into NQL. It runs locally on CPU in a few milliseconds. No API key, no network call, no tokens billed to anyone.

Why this lives in the engine

The hard part of natural-language querying isn't the model — it's knowing the schema. A client has to fetch the collection list and pass it in, and it's stale the moment it arrives. The engine already holds the live list.

So the planner is constrained against real collections at the instant of the call:

{ "prompt": "show me all stylists",
  "nql": "FROM stylists",
  "valid": true,
  "collection_known": false,
  "error": "collection \"stylists\" does not exist in \"shop\"" }

HTTP 422. Not an empty result set — an empty result set reads as "no matching rows", which would be a lie. The query was fine; the collection was imaginary.

Every nedbd client inherits this for free instead of reimplementing it.

Three safety properties

1 · The model never executes anything. It emits text. That text goes to the same nql::query path a hand-typed query uses. There is no second executor to audit.

2 · Validation is parsing, not pattern-matching. nql::parse and nql::execute share one code path, so they cannot disagree about what is well-formed. Unparseable output returns 422 with the offending text attached.

3 · execute defaults to false. You get a plan for review. Opt in explicitly:

curl -X POST localhost:7070/v1/databases/shop/cast \
  -d '{"prompt":"orders over 100","execute":true}'
# → { …, "executed": true, "count": 2, "rows": [ … ] }

That default is not decoration. Here is a real miss, from a real run:

prompt   "paid orders over 100"
nql      FROM orders WHERE status = "paid" LIMIT 100      ← wrong
correct  FROM orders WHERE status = "paid" AND total > 100

It read "over 100" as LIMIT 100 and dropped the second predicate. The row count still came back 2, because both paid orders happened to exceed 100 — a count-only check would have called that a pass. A human reading LIMIT 100 catches it instantly. An auto-executing client does not.

Multi-predicate WHERE is the model's weakest clause (85.1% eval / 61.2% holdout). Ship accordingly.

The failure valid cannot catch

A literal the model invented rather than copied:

"memories about pricing"  ->  FROM memories SEARCH "handoff"

That query parses. It names a real collection. It returns real rows. Both valid and collection_known are true — and it answers a question nobody asked. Measured on the released checkpoint:

terms in vocabulary copied correctly
release flow · guardrail · handoff yes 3/3
pricing · deadlines · kubernetes no 0/3 — all became "handoff"

So the response carries a drift field when a quoted literal is absent from the prompt:

{ "nql": "FROM memories SEARCH \"handoff\"",
  "valid": true,
  "collection_known": true,
  "drift": "generated the literal \"handoff\", which does not appear in the prompt — likely outside the model's vocabulary and substituted. Verify before trusting these results." }

It is advisory, never fatal — the plan may still be what you wanted, and discarding a valid query would be its own kind of lie. But an unattended caller should treat it as a third gate:

if plan["valid"] and plan["collection_known"] and not plan.get("drift"):
    rows = await db.query(plan["nql"])

Same root cause as truncated digits (height 4000004000): no copy mechanism over prompt tokens. Verified at 24/24 on real model output — 3 true positives, 21 true negatives, zero false alarms, including the case that matters most (correctly inferred enum values like "refunded orders"status = "refunded" stay silent).

Enabling it

Off by default because it pulls a model dependency and expects weights at runtime — most deployments want neither.

# 1. build with the feature
cargo install nedb-engine --features cast

# 2. get the weights (~13 MB)
curl -L -o model.cast \
  https://github.com/aiassistsecure/nedb-cast-slm/releases/download/v10.30.90/model.cast

# 3. run
nedbd --dag --cast ./data
#   cast     enabled — 3.33M params, vocab 581, ./data/model.cast

Model search order, first hit wins:

  1. $NEDBD_CAST_MODEL
  2. <data_dir>/model.cast
  3. $CAST_HOME/model.cast
  4. ~/.cache/nedb-cast-slm/v10.30.90/model.cast — where the Python and npm packages cache it, so a machine that has run either one is already ready

The container is checksum-verified on load. A silently corrupt model would emit plausible-but-wrong queries, which is the worst possible failure mode for a query planner.

Without the feature the route still exists and returns 501 — so a client can detect the capability instead of guessing. Missing weights on a --cast build is loud but non-fatal: the daemon starts and serves everything else normally.

As a library

use nedb_engine::cast::Caster;

let caster = Caster::load(&data_dir)?;
let out = caster.cast_checked("orders over 100", &db.id_index.collections());

if out.collection_known && nedb_engine::nql::parse(&out.nql).is_ok() {
    let (rows, count) = nedb_engine::nql::query(&db, &out.nql)?;
}

Decoding is greedy and deterministic — a DSL has exactly one right answer, so sampling could only hurt. The same prompt always produces the same plan, which makes the endpoint safe to cache.


NQL — NEDB Query Language

-- Basic queries
FROM person LIMIT 10
FROM driver WHERE status = "active"
FROM driver WHERE rating >= 4.5 ORDER BY rating DESC

-- Time-travel
FROM will WHERE _id = "evans_will_2019" AS OF 5

-- Causal trace (walks caused_by links backward)
FROM event WHERE _id = "probate_filing" TRACE caused_by

-- Graph traversal
FROM person WHERE _id = "robert" TRAVERSE parent_of

Example: The Will — causal DAG in action

import json, os, urllib.request

def put(db_url, coll, id_, doc):
    body = json.dumps({"coll": coll, "id": id_, "doc": doc}).encode()
    req = urllib.request.Request(f"{db_url}/put", body,
          {"Content-Type": "application/json"}, method="POST")
    return json.loads(urllib.request.urlopen(req).read())["doc"]

BASE = "http://localhost:7070/v1/databases/will"

# Robert writes his will
robert = put(BASE, "person", "robert", {"name": "Robert Evans", "role": "testator"})
will_v1 = put(BASE, "will", "evans_will", {
    "house": "mark", "business": "lisa",
    "caused_by": [robert["_hash"]],   # causal link to testator record
})

# Amendment — chains off v1
will_v2 = put(BASE, "will", "evans_will", {
    "house": "mark", "business": "lisa", "vintage_car": "mark",
    "caused_by": [will_v1["_hash"]],
})

# TRACE proves the amendment links back to the original
query = json.dumps({"nql": 'FROM will WHERE _id = "evans_will" TRACE caused_by'}).encode()
req = urllib.request.Request(f"{BASE}/query", query,
      {"Content-Type": "application/json"}, method="POST")
trace = json.loads(urllib.request.urlopen(req).read())["rows"]
# → both versions + Robert's testator record, in causal order

Python companion

The nedb-engine PyPI package ships the same server binary bundled in the wheel, plus:

  • Pure-Python AOF engine (NEDB class)
  • Embedded v2 DAG API (nedb._native.NedbCore) on supported platforms
  • nedbd console script
pip install nedb-engine

# Use embedded API (Linux/macOS/Windows CPython)
python3 -c "from nedb._native import NedbCore; db = NedbCore(); ..."

# Use HTTP mode (any platform, including MSYS2/MinGW)
NEDB_URL=http://localhost:7070 python3 your_script.py

# Diagnose
nedbd --doctor

License

MIT — free for any use, including commercial and production. © 2026 INTERCHAINED LLC.


Built by INTERCHAINED, LLC × Claude Sonnet 4.6