Skip to main content

Module keys

Module keys 

Source
Expand description

The keyspaces. Extensibility = new tags, never new structures.

Big-endian everywhere: byte order IS numeric order, so “everything of X” is a range scan from a prefix. Tag first, so spaces never interleave.

0x00 catalog | 0x01 node(id) -> payload | 0x03 edge(src,type,dst) -> props 0x04 redge(dst,type,src) mirror | 0x05 vec(field,id) -> embedding (phase 2)

Constants§

CAT_FIELD_AGGREGATE
Disposable SQL aggregate summaries. Kept before append-heavy data tags. Recovery must discard these: a salvaged source can differ from its summary.
CAT_KEY_ORDER
Catalogue subspace for the SQL covering _key order index.
TAG_CATALOG
TAG_CEDGE
ctx != 0 edges live in their own, wider keyspace. The base graph (ctx = 0) keeps 25-byte keys: measured, the flat +8B/key grew a 5M-node file by 640 MB and doubled 3-hop latency through OS-cache pressure alone – a price paid by workloads that never use perspectives. Split tags, each side pays its own way. Sort order per space is unchanged (separate tags).
TAG_CREDGE
TAG_EDGE
TAG_EXT
TAG_FIELDIDX
SQL field (secondary) indexes: (coll, field, order-preserving value bytes, id) -> []. The VALUE ENCODING is the SQL layer’s business; the kernel only promises byte-ordered iteration (D4: an index is rows).
TAG_GEOM
Geometry rows (2i): 0x10 | field | id -> binary-encoded Geom. The kernel speaks TYPED geometry only; GeoJSON parsing is an API-layer concern (the core stays pure – no JSON dependency below the SQL line).
TAG_LABEL
TAG_NAV
Vector navigation rows (2k): 0x11 | field | id -> [norm f32][2-bit code] [n u16][neighbor u64…]. Code co-located with links (FACT-02): one row read per visited node during a beam walk – ranking data arrives with the topology, no second lookup. One small-world graph per field: links carry bare ids, so a shared keyspace would wire two fields’ vectors into one another’s neighbourhoods.
TAG_NODE
TAG_PROP
Property index: 0x0A | prop_id | value | node_id, empty value. 25B. The VALUE is an order-preserving 8-byte encoding, so every property query is a range scan: equality = one (prop,value) prefix, range = (prop,lo)..hi, top-k DESC = a descending-encoded prop scanned forward with LIMIT.
TAG_REDGE
TAG_SEARCHSLOT
SQL SEARCH-index slot maps (phase 3c): dense u32 slots <-> node hashes, per search index. Kind 0: slot -> hash. Kind 1: hash -> slot. Kind 2: next-slot counter.
TAG_SPAT
Spatial cell postings (2i): 0x0F | field | level u8 | hilbert u64 | id -> [bbox f32x4 outward-rounded]. Hilbert order makes a neighbourhood’s postings contiguous on disk; the bbox filters without payload reads; a degenerate bbox IS the point (exact tier free for point data).
TAG_SQLMETA
SQL-layer metadata (phase 3): collection-name interning, schemas, edge-type names. The kernel never reads these rows; the tag is minted here so the keyspace vocabulary stays in ONE place (D4).
TAG_TEXT
Full-text postings (2h). Two shapes share the tag, split by segment id: HEAD (seg 0, mutable): 0x0C | field | 0u32 | term | 0x00 | docid -> tf varint FOLDED (seg>0, immut.): 0x0C | field | seg | term -> packed postings The head is row-per-posting so every write is BLIND (appending to a packed value would be read-modify-write – the e1 BM25 wound); folds pack head rows into value-per-term segments. 0x00 separates term from docid in head keys: tokenizer terms are alphanumeric UTF-8 (never 0x00), and the separator MUST sort before every text byte – with 0xFF, longer terms sorted ahead of their own prefixes (“handle” before “hand”) and the fuzzy walk’s seek skipped real terms; the oracle test caught it.
TAG_TEXTMETA
Per-(field, seg) metadata: doc_count + total_tokens (BM25 denominators) and the segment’s alive-bitmap (the deletion discipline: one value).
TAG_TEXTNORM
Per-(field, doc) token count – BM25’s |d|. Point lookups at scoring time only (candidates are few), so no packed norm blocks are needed.
TAG_VCODE
Vector fingerprint (2g): 0x0B | field | id -> [norm f32-LE][packed code]. Its own keyspace so nearest scans codes ONLY – never the 6KB vectors. Codes of different fields are different WIDTHS (the width follows the field’s dimension), so the field must be in the key: a scan that mixed them would decode one field’s bytes with another field’s recipe.
TAG_VEC
Embedding rows (2e): 0x05 | field | id -> f32-LE coordinates. FIELD FIRST, like every other index family here: each vector field owns a disjoint, prefix-scannable range, so two embedding columns of two different widths never meet in one scan.
TEXT_FIELD_META_SEG
TEXT_TERM_STATS_SEG
Reserved metadata coordinates. Segment ids are allocated below these two values; keeping them under the existing metadata tag avoids minting a tag above append-heavy data keyspaces (the measured occupancy trap in D30).

Functions§

catalog
catalog_field
A catalog slot subdivided PER FIELD: 0x00 | slot | field, 17 bytes. It cannot collide with the 9-byte global slots – different lengths are different keys – which is why an arbitrary field hash is safe here while catalog(field_hash) would not be: that would land on whichever global slot the hash happened to equal.
catalog_field_item
A per-item record within a catalog field: 0x00 | slot | field | item. Use this for sparse bookkeeping that must sort ahead of every data keyspace; putting it under a later tag changes right-edge append splits.
catalog_field_prefix
ctx_prefix
edge
edge_prefix
edge_type_prefix
enc_f64
enc_f64_desc
Descending variants: an ascending scan over these yields DESC order, which is how top-k works on a forward-only iterator.
enc_i64
Order-preserving encodings: byte order of the 8-byte result == the natural order of the value, which is the whole trick that turns queries into ranges. i64: flip the sign bit. f64 (IEEE 754): negative -> flip ALL bits, non-negative -> flip the sign bit; total order matches numeric order (NaN sorts above +inf; callers who care filter NaN before indexing).
enc_i64_desc
ext_hash
FNV-1a 64. Only contract: same bytes -> same hash. Collision = wrong id returned by resolve(); the extkey VALUE stores the full external key so a collision is DETECTED (compare) rather than silently wrong.
extkey
field_aggregate_key
fieldidx_key
(kind, hash) -> bytes. kind: 0 = collection name, 1 = table schema, 2 = edge-type name, 3 = index definition, 4 = edge-insertion counter (hash 0; the SQL layer’s monotonic edge seq), 6 = vector field name. (coll_hash, field_hash, value_bytes, id). id is the fixed 8-byte suffix – positional, never searched, so value bytes are unrestricted.
fieldidx_prefix
geom_key
geom_prefix
is_field_aggregate_key
key_order_key
One live-row entry. NUL is escaped as 00 ff, then 00 00 terminates the UTF-8 key before the row hash tie-breaker. This is order preserving (a key sorts before every longer key it prefixes), supports embedded NUL, and makes duplicate logical keys distinct physical rows rather than an overwritten value. The key bytes remain covering and need no payload read.
key_order_prefix
(catalog, key-order slot, collection, field) prefix.
label
label_prefix
nav_key
nav_prefix
node
prop
prop_prefix
prop_value_prefix
redge
redge_prefix
redge_type_prefix
searchslot_key
searchslot_prefix
spat_key
spat_prefix
sqlmeta_key
text_field_meta_key
text_head_doc_key
One document-membership row in the mutable text head. Empty terms do not exist, so the zero byte after segment 0 is an unambiguous namespace for streaming/counting the documents owned by the head.
text_head_key
text_meta_key
text_meta_prefix
text_norm_key
text_norm_prefix
text_prefix
text_seg_block_key
text_seg_block_prefix
Prefix and key for a bounded immutable posting block. Token bytes never contain zero, so [term|0x00|first_doc] remains prefix-seekable by term.
text_seg_doc_key
text_seg_doc_prefix
Per-document membership/length row for an immutable segment. Ownership metadata belongs under TEXTMETA so folded posting-row shape stays stable.
text_seg_key
text_term_id_key
text_term_lex_key
text_term_lex_prefix
Lexically ordered live-term directory used by prefix and typo expansion. Counts remain in the hashed point-lookup rows above; this second view keeps expansion from walking posting blocks or the number of immutable batches.
u64_at
Decode helpers for scans.
vcode_key
vcode_prefix
vec_key
vec_prefix