Expand description
The keyspaces. Extensibility = new tags, never new structures.
Big-endian everywhere: byte order IS numeric order, so “everything of X” is a range scan from a prefix. Tag first, so spaces never interleave.
0x00 catalog | 0x01 node(id) -> payload | 0x03 edge(src,type,dst) -> props 0x04 redge(dst,type,src) mirror | 0x05 vec(field,id) -> embedding (phase 2)
Constants§
- CAT_
FIELD_ AGGREGATE - Disposable SQL aggregate summaries. Kept before append-heavy data tags. Recovery must discard these: a salvaged source can differ from its summary.
- CAT_
KEY_ ORDER - Catalogue subspace for the SQL covering
_keyorder index. - TAG_
CATALOG - TAG_
CEDGE - ctx != 0 edges live in their own, wider keyspace. The base graph (ctx = 0) keeps 25-byte keys: measured, the flat +8B/key grew a 5M-node file by 640 MB and doubled 3-hop latency through OS-cache pressure alone – a price paid by workloads that never use perspectives. Split tags, each side pays its own way. Sort order per space is unchanged (separate tags).
- TAG_
CREDGE - TAG_
EDGE - TAG_EXT
- TAG_
FIELDIDX - SQL field (secondary) indexes: (coll, field, order-preserving value bytes, id) -> []. The VALUE ENCODING is the SQL layer’s business; the kernel only promises byte-ordered iteration (D4: an index is rows).
- TAG_
GEOM - Geometry rows (2i): 0x10 | field | id -> binary-encoded Geom. The kernel speaks TYPED geometry only; GeoJSON parsing is an API-layer concern (the core stays pure – no JSON dependency below the SQL line).
- TAG_
LABEL - TAG_NAV
- Vector navigation rows (2k): 0x11 | field | id -> [norm f32][2-bit code] [n u16][neighbor u64…]. Code co-located with links (FACT-02): one row read per visited node during a beam walk – ranking data arrives with the topology, no second lookup. One small-world graph per field: links carry bare ids, so a shared keyspace would wire two fields’ vectors into one another’s neighbourhoods.
- TAG_
NODE - TAG_
PROP - Property index: 0x0A | prop_id | value | node_id, empty value. 25B. The VALUE is an order-preserving 8-byte encoding, so every property query is a range scan: equality = one (prop,value) prefix, range = (prop,lo)..hi, top-k DESC = a descending-encoded prop scanned forward with LIMIT.
- TAG_
REDGE - TAG_
SEARCHSLOT - SQL SEARCH-index slot maps (phase 3c): dense u32 slots <-> node hashes, per search index. Kind 0: slot -> hash. Kind 1: hash -> slot. Kind 2: next-slot counter.
- TAG_
SPAT - Spatial cell postings (2i): 0x0F | field | level u8 | hilbert u64 | id -> [bbox f32x4 outward-rounded]. Hilbert order makes a neighbourhood’s postings contiguous on disk; the bbox filters without payload reads; a degenerate bbox IS the point (exact tier free for point data).
- TAG_
SQLMETA - SQL-layer metadata (phase 3): collection-name interning, schemas, edge-type names. The kernel never reads these rows; the tag is minted here so the keyspace vocabulary stays in ONE place (D4).
- TAG_
TEXT - Full-text postings (2h). Two shapes share the tag, split by segment id: HEAD (seg 0, mutable): 0x0C | field | 0u32 | term | 0x00 | docid -> tf varint FOLDED (seg>0, immut.): 0x0C | field | seg | term -> packed postings The head is row-per-posting so every write is BLIND (appending to a packed value would be read-modify-write – the e1 BM25 wound); folds pack head rows into value-per-term segments. 0x00 separates term from docid in head keys: tokenizer terms are alphanumeric UTF-8 (never 0x00), and the separator MUST sort before every text byte – with 0xFF, longer terms sorted ahead of their own prefixes (“handle” before “hand”) and the fuzzy walk’s seek skipped real terms; the oracle test caught it.
- TAG_
TEXTMETA - Per-(field, seg) metadata: doc_count + total_tokens (BM25 denominators) and the segment’s alive-bitmap (the deletion discipline: one value).
- TAG_
TEXTNORM - Per-(field, doc) token count – BM25’s |d|. Point lookups at scoring time only (candidates are few), so no packed norm blocks are needed.
- TAG_
VCODE - Vector fingerprint (2g): 0x0B | field | id -> [norm f32-LE][packed code].
Its own keyspace so
nearestscans codes ONLY – never the 6KB vectors. Codes of different fields are different WIDTHS (the width follows the field’s dimension), so the field must be in the key: a scan that mixed them would decode one field’s bytes with another field’s recipe. - TAG_VEC
- Embedding rows (2e): 0x05 | field | id -> f32-LE coordinates. FIELD FIRST, like every other index family here: each vector field owns a disjoint, prefix-scannable range, so two embedding columns of two different widths never meet in one scan.
- TEXT_
FIELD_ META_ SEG - TEXT_
TERM_ STATS_ SEG - Reserved metadata coordinates. Segment ids are allocated below these two values; keeping them under the existing metadata tag avoids minting a tag above append-heavy data keyspaces (the measured occupancy trap in D30).
Functions§
- catalog
- catalog_
field - A catalog slot subdivided PER FIELD: 0x00 | slot | field, 17 bytes.
It cannot collide with the 9-byte global slots – different lengths are
different keys – which is why an arbitrary field hash is safe here
while
catalog(field_hash)would not be: that would land on whichever global slot the hash happened to equal. - catalog_
field_ item - A per-item record within a catalog field: 0x00 | slot | field | item. Use this for sparse bookkeeping that must sort ahead of every data keyspace; putting it under a later tag changes right-edge append splits.
- catalog_
field_ prefix - ctx_
prefix - edge
- edge_
prefix - edge_
type_ prefix - enc_f64
- enc_
f64_ desc - Descending variants: an ascending scan over these yields DESC order, which is how top-k works on a forward-only iterator.
- enc_i64
- Order-preserving encodings: byte order of the 8-byte result == the natural order of the value, which is the whole trick that turns queries into ranges. i64: flip the sign bit. f64 (IEEE 754): negative -> flip ALL bits, non-negative -> flip the sign bit; total order matches numeric order (NaN sorts above +inf; callers who care filter NaN before indexing).
- enc_
i64_ desc - ext_
hash - FNV-1a 64. Only contract: same bytes -> same hash. Collision = wrong id returned by resolve(); the extkey VALUE stores the full external key so a collision is DETECTED (compare) rather than silently wrong.
- extkey
- field_
aggregate_ key - fieldidx_
key - (kind, hash) -> bytes. kind: 0 = collection name, 1 = table schema, 2 = edge-type name, 3 = index definition, 4 = edge-insertion counter (hash 0; the SQL layer’s monotonic edge seq), 6 = vector field name. (coll_hash, field_hash, value_bytes, id). id is the fixed 8-byte suffix – positional, never searched, so value bytes are unrestricted.
- fieldidx_
prefix - geom_
key - geom_
prefix - is_
field_ aggregate_ key - key_
order_ key - One live-row entry. NUL is escaped as
00 ff, then00 00terminates the UTF-8 key before the row hash tie-breaker. This is order preserving (a key sorts before every longer key it prefixes), supports embedded NUL, and makes duplicate logical keys distinct physical rows rather than an overwritten value. The key bytes remain covering and need no payload read. - key_
order_ prefix (catalog, key-order slot, collection, field)prefix.- label
- label_
prefix - nav_key
- nav_
prefix - node
- prop
- prop_
prefix - prop_
value_ prefix - redge
- redge_
prefix - redge_
type_ prefix - searchslot_
key - searchslot_
prefix - spat_
key - spat_
prefix - sqlmeta_
key - text_
field_ meta_ key - text_
head_ doc_ key - One document-membership row in the mutable text head. Empty terms do not exist, so the zero byte after segment 0 is an unambiguous namespace for streaming/counting the documents owned by the head.
- text_
head_ key - text_
meta_ key - text_
meta_ prefix - text_
norm_ key - text_
norm_ prefix - text_
prefix - text_
seg_ block_ key - text_
seg_ block_ prefix - Prefix and key for a bounded immutable posting block. Token bytes never
contain zero, so
[term|0x00|first_doc]remains prefix-seekable by term. - text_
seg_ doc_ key - text_
seg_ doc_ prefix - Per-document membership/length row for an immutable segment. Ownership metadata belongs under TEXTMETA so folded posting-row shape stays stable.
- text_
seg_ key - text_
term_ id_ key - text_
term_ lex_ key - text_
term_ lex_ prefix - Lexically ordered live-term directory used by prefix and typo expansion. Counts remain in the hashed point-lookup rows above; this second view keeps expansion from walking posting blocks or the number of immutable batches.
- u64_at
- Decode helpers for scans.
- vcode_
key - vcode_
prefix - vec_key
- vec_
prefix