pub struct Dictionary { /* private fields */ }Expand description
A read-only dictionary mapping terms ↔ role-specific IDs.
Holds the four sections as [ChunkedSection]s — local files keep each
section as one resident chunk; a lazily-opened remote file faults
individual chunks in on first touch. Lookups never re-parse a header.
Implementations§
Source§impl Dictionary
impl Dictionary
Sourcepub fn from_sections(sections: [Vec<u8>; 4]) -> Dictionary
pub fn from_sections(sections: [Vec<u8>; 4]) -> Dictionary
Rebuild from four serialized sections (shared, subjects, objects,
predicates), e.g. when reading a .rete file. The four sections share one
[ChunkCache] (unlimited cap here, so every section body stays resident
exactly as before).
Sourcepub fn from_chunked_sections(sections: [ChunkedSection; 4]) -> Dictionary
pub fn from_chunked_sections(sections: [ChunkedSection; 4]) -> Dictionary
Rebuild from four already-chunked sections (the remote lazy-open path).
Sourcepub fn term_count(&self) -> u32
pub fn term_count(&self) -> u32
Total number of distinct terms across all four sections.
Sourcepub fn set_cache_cap(&self, cap: u64)
pub fn set_cache_cap(&self, cap: u64)
Set the byte cap of the shared chunk cache (bounded-export phase 2).
u64::MAX (the default from every open) means unlimited — no eviction,
residency identical to the pre-budget design. A finite cap evicts
least-recently-used chunk bodies down to it on every fault, so a full
dump’s peak dictionary residency is bounded rather than the whole
decompressed dictionary. All four sections share one cache; setting it on
each is idempotent (they are the same Arc).
Sourcepub fn cache_stats(&self) -> CacheStats
pub fn cache_stats(&self) -> CacheStats
A snapshot of the shared chunk cache’s observability counters — for
RETE_OPEN_DEBUG accounting and the eviction tests.
Sourcepub fn has_quoted_triples(&self) -> bool
pub fn has_quoted_triples(&self) -> bool
Does this dataset contain RDF-star quoted triples? Meaningful only for a
freshly built dictionary (a read-back one carries the truth in the file
header’s FLAG_HAS_QUOTED_TRIPLES, not here).
Sourcepub fn load_incomplete(&self) -> bool
pub fn load_incomplete(&self) -> bool
Did any lazy chunk fetch fail since this dictionary was opened — or
since the last reset_load_failure?
Sourcepub fn reset_load_failure(&self)
pub fn reset_load_failure(&self)
Forget recorded fetch failures across every section (see
GraphIndex::reset_load_failure — the per-query reset for resident
sessions). Failed chunks were never cached, so they simply retry.
Sourcepub fn prefetch_all(&self)
pub fn prefetch_all(&self)
Batch-fault every unloaded chunk of every section (no-op for local dictionaries). Callers about to resolve every term — export, dump — call this once so the sweep coalesces into a few range reads instead of one fetch per chunk.
Sourcepub fn prefetch_terms(&self, node_ids: &[u32], predicate_ids: &[u32])
pub fn prefetch_terms(&self, node_ids: &[u32], predicate_ids: &[u32])
Batch-fault just the chunks needed to resolve a bounded result set:
node_ids are unified-node ids (the Val::Id(id) with id >= 0),
predicate_ids are predicate-space ids. Groups them by (section, chunk) and coalesces each section’s faults into a few range reads —
turning “one request per distinct output term” into “a handful per
query”. No-op for a local dictionary (every chunk is already resident).
Number of shared terms S.
Sourcepub fn subject_only_count(&self) -> u32
pub fn subject_only_count(&self) -> u32
Number of subject-only terms Su.
Sourcepub fn object_only_count(&self) -> u32
pub fn object_only_count(&self) -> u32
Number of object-only terms Oo.
Sourcepub fn node_count(&self) -> u32
pub fn node_count(&self) -> u32
Total distinct graph nodes N = S + Su + Oo.
Sourcepub fn subject_node(&self, sid: u32) -> u32
pub fn subject_node(&self, sid: u32) -> u32
Unified node ID for a subject-role ID. Shared and subject-only IDs are
already contiguous (1..=S+Su), so this is just sid - 1.
Sourcepub fn object_node(&self, oid: u32) -> u32
pub fn object_node(&self, oid: u32) -> u32
Unified node ID for an object-role ID: shared stay at oid-1, object-only
IDs are shifted past the subject-only block.
Sourcepub fn node_term(&self, node: u32) -> Option<String>
pub fn node_term(&self, node: u32) -> Option<String>
Resolve a unified node ID back to its term.
Sourcepub fn prefetch_subject_terms(&self, ids: &[u32])
pub fn prefetch_subject_terms(&self, ids: &[u32])
Batch-fault the dictionary chunks that decoding these subject-role ids will touch — one coalesced read per section instead of a blocking round trip per id. Purely a cache warmer (resident sections no-op).
Sourcepub fn prefetch_node_terms(&self, nodes: &[u32])
pub fn prefetch_node_terms(&self, nodes: &[u32])
prefetch_subject_terms for unified
node ids — the id space solution rows carry. Warms every chunk a batch
of row decodes (a FILTER’s string ops, the projection) is about to
fault, using node_term’s exact routing.
Sourcepub fn node_of_term(&self, term: &str) -> Option<u32>
pub fn node_of_term(&self, term: &str) -> Option<u32>
Unified node ID for a term, resolving via subject role then object role.
Sourcepub fn node_as_subject_id(&self, node: u32) -> Option<u32>
pub fn node_as_subject_id(&self, node: u32) -> Option<u32>
Subject-role ID of a node, or None if the node never appears as a
subject (an object-only term).
Sourcepub fn node_as_object_id(&self, node: u32) -> Option<u32>
pub fn node_as_object_id(&self, node: u32) -> Option<u32>
Object-role ID of a node, or None if the node never appears as an
object (a subject-only term).
Sourcepub fn subject_id(&self, term: &str) -> Option<u32>
pub fn subject_id(&self, term: &str) -> Option<u32>
Subject-role ID for term: shared 1..=S, else subject-only S+1...
Sourcepub fn object_id(&self, term: &str) -> Option<u32>
pub fn object_id(&self, term: &str) -> Option<u32>
Object-role ID for term: shared 1..=S, else object-only S+1...
Sourcepub fn predicate_id(&self, term: &str) -> Option<u32>
pub fn predicate_id(&self, term: &str) -> Option<u32>
Predicate ID for term (independent space).
pub fn subject_term(&self, id: u32) -> Option<String>
pub fn object_term(&self, id: u32) -> Option<String>
pub fn predicate_term(&self, id: u32) -> Option<String>
Sourcepub fn encode(&self, s: &str, p: &str, o: &str) -> Option<(u32, u32, u32)>
pub fn encode(&self, s: &str, p: &str, o: &str) -> Option<(u32, u32, u32)>
Encode a (s, p, o) term triple to its (subject_id, predicate_id, object_id). None if any term is unknown.
Sourcepub fn resolve_window(
&self,
ids: &[(u32, u32, u32)],
w: &mut WindowResolver,
emit: impl FnMut(&str, &str, &str),
)
pub fn resolve_window( &self, ids: &[(u32, u32, u32)], w: &mut WindowResolver, emit: impl FnMut(&str, &str, &str), )
Resolve a window of (subject_id, predicate_id, object_id) triples in
chunk order and emit each fully-resolved row through emit in the
window’s original order.
This is the batch twin of resolving each row with
subject_term /
predicate_term /
object_term and emitting on (Some, Some, Some)
— the exact thing it replaces inside Rete::dump_filtered_each. A row is
emitted only when all three of its terms resolve, and the bytes of each
term are identical to the per-row path, so the dump stays
byte-for-byte the same.
What changes is how the terms are decoded. Each row contributes three
jobs to the section that owns its id (shared/subject-only for the
subject, predicates for the predicate, shared/object-only for the
object); each section then resolves its jobs in one chunk-ordered,
single-pass walk (ChunkedSection::resolve_into). A window whose
object ids scatter across the object-only section therefore touches each
chunk once — one faulted body, one forward walk — instead of the
scattered re-faults and per-row re-walks the row-order path performed.
w carries the per-window scratch buffers so back-to-back windows cost
no allocation beyond growth.