pub struct Rete { /* private fields */ }Expand description
A read-only, in-memory view over a .rete file image.
Implementations§
Source§impl Rete
impl Rete
Sourcepub fn open(bytes: &[u8]) -> Result<Rete, FileError>
pub fn open(bytes: &[u8]) -> Result<Rete, FileError>
Parse a full file image (v0 loads everything; a range-reading client will fetch only the sections it needs — same container format).
Sourcepub fn set_service_client(&mut self, client: Box<dyn ServiceClient>)
pub fn set_service_client(&mut self, client: Box<dyn ServiceClient>)
Attach the client that executes SERVICE <endpoint> { … } blocks
(SPARQL 1.1 federated query) against remote SPARQL endpoints. Without
one, a non-SILENT SERVICE fails the query with a clear error and a
SERVICE SILENT degrades to one empty solution, per the spec.
pub fn header(&self) -> &Header
Sourcepub fn file_layout(&self) -> Vec<LayoutSegment>
pub fn file_layout(&self) -> Vec<LayoutSegment>
The file’s byte layout, for visualization: header, metadata, dictionary, each index permutation’s tile directory and individual tiles, pyramid summary, and named graphs — sorted by offset. Bytes not covered by any segment are container framing (section directories and length fields).
Sourcepub fn metadata(&self) -> Option<&[u8]>
pub fn metadata(&self) -> Option<&[u8]>
Raw bytes of the file’s metadata section, or None if it has none. The
CLI stores a JSON Dataset Card here; rete-core treats it as opaque.
Populated by Rete::open only — an Rete::open_ranged view returns
None here (the card is not fetched on the minimal query path).
pub fn dictionary(&self) -> &Dictionary
Sourcepub fn pyramid(&self) -> Option<&PyramidMeta>
pub fn pyramid(&self) -> Option<&PyramidMeta>
The pyramid metadata (summary graph + tiles), if the file has a pyramid.
Sourcepub fn pyramid_if_loaded(&self) -> Option<&PyramidMeta>
pub fn pyramid_if_loaded(&self) -> Option<&PyramidMeta>
The pyramid metadata only if already resident or previously faulted — never triggers a lazy range read. The query planner uses this for cardinality estimation so it is free for an in-memory file and never adds a fetch on the lazy remote path (which defers the pyramid by design).
Sourcepub fn predicate_stats(&self) -> &[PredStat]
pub fn predicate_stats(&self) -> &[PredStat]
Per-predicate planner statistics from the query-stats block — empty when
the file has none or the pyramid isn’t resident (the lazy path doesn’t
fault it just for stats). See crate::meta::PredStat.
Sourcepub fn char_sets(&self) -> &[CharSet]
pub fn char_sets(&self) -> &[CharSet]
The entity shapes (characteristic sets) from the pyramid — empty when the
file has none or the pyramid isn’t resident. See crate::meta::CharSet.
Sourcepub fn label_index(&self) -> &[LabelEntry]
pub fn label_index(&self) -> &[LabelEntry]
The label index from the pyramid — empty when the file has none or the
pyramid isn’t resident. See crate::meta::LabelEntry.
Sourcepub fn prefix_search(&self, prefix: &str, limit: usize) -> Vec<(String, String)>
pub fn prefix_search(&self, prefix: &str, limit: usize) -> Vec<(String, String)>
Prefix-search the label index: the subjects whose label starts with
prefix (case-insensitive), as (label, subject_iri), capped at limit.
Unlike the planner accessors, this faults the pyramid (where the index
lives) on the lazy path — a prefix search is an explicit read, not a free
estimate. Returns an empty vec when the file carries no label index.
Sourcepub fn has_text_index(&self) -> bool
pub fn has_text_index(&self) -> bool
Whether this file carries a full-text (TEXT_INDEX) section, i.e. it was
built with --text-index. Cheap — reads the header, never faults.
Sourcepub fn text_index_token_table_len(&self) -> Option<u64>
pub fn text_index_token_table_len(&self) -> Option<u64>
Byte length of the TEXT_INDEX section’s leading token table — the
length varint plus the compressed table, which is exactly what a first
text_search faults. None when the file carries
no text index (or its head could not be read).
This is the number to quote as the cost of a first search, NOT
header().text_index_len: that counts the postings blob too, which is
the bulk of the section and is only ever fetched one posting list at a
time. On the published epfl-infoscience.rete the section is 195 MB and
its token table 29 MB — a 6.5× difference.
Costs nothing on a resident open (measured while the bytes were in hand); a lazy/ranged open pays one ≤10-byte range read, memoized, and still never faults the table itself.
Sourcepub fn text_search(
&self,
words: &[&str],
prefix: Option<&str>,
limit: usize,
) -> Vec<String>
pub fn text_search( &self, words: &[&str], prefix: Option<&str>, limit: usize, ) -> Vec<String>
Full-text search over the literals: subject IRIs that carry every word
in words (whole-word, case-insensitive — AND semantics), optionally also
requiring a word that starts with prefix (token-prefix). Results are
ordered by subject id and capped at limit (0 = uncapped). Empty when the
file has no text index or nothing matches.
Like prefix_search, this faults the index on
the lazy remote path — a search is an explicit read, and only the queried
posting lists are fetched, not the whole index.
Sourcepub fn default_index(&self) -> &GraphIndex
pub fn default_index(&self) -> &GraphIndex
The default-graph permutation index.
Sourcepub fn dump(&self, graph: Option<&str>) -> Vec<(String, String, String)>
pub fn dump(&self, graph: Option<&str>) -> Vec<(String, String, String)>
Resolve every triple of a graph (None = default graph) back to terms.
Sourcepub fn dump_each<F>(&self, graph: Option<&str>, f: F)
pub fn dump_each<F>(&self, graph: Option<&str>, f: F)
Stream every triple of a graph (None = default) to f, resolving terms
a batch at a time — no full Vec materialization, so it is safe on graphs far
larger than RAM. rete export uses this to serialize 100M+ triple files
that dump() (which collects every term into a Vec<String>) would OOM on.
On a lazy/ranged open the dictionary is not prefetched whole: each batch faults only the chunks its own terms live in. Dumping ONE named graph therefore costs that graph’s terms, not the file’s. Dumping everything still reads every dictionary byte and leaves every chunk resident — that is inherent, and it, not the prefetch, is what sets the peak of a full export. See the comment in the body for both measurements.
To dump a slice — one predicate, one subject, one object — call
dump_filtered_each, which is this method
with a triple pattern and prunes tiles instead of filtering rows.
Sourcepub fn dump_filtered_each<F>(
&self,
graph: Option<&str>,
s: Option<&str>,
p: Option<&str>,
o: Option<&str>,
f: F,
)
pub fn dump_filtered_each<F>( &self, graph: Option<&str>, s: Option<&str>, p: Option<&str>, o: Option<&str>, f: F, )
dump_each restricted to a triple pattern — the
filtered dump. Any of s / p / o may be bound (canonical N-Triples
term tokens); all three unbound is exactly dump_each.
§Why this is cheap, and how much
It is the same access path a routed query takes, not a full scan with a
predicate test bolted on: GraphIndex::scan_iter picks the permutation
with the longest bound prefix, binary-searches that section’s tile
directory down to the span a bound leading id can live in, and then drops
every remaining tile whose recorded synopsis proves it cannot hold a
bound secondary component — all from directories, without fetching a
tile. So a filtered dump pays for the slice, which is the whole point
of issue #117.
Measured on cordis.rete (801 MB, a 417 MB dictionary, 26.4M quads,
six named graphs), lazily opened; “before” is the only way to get a
predicate slice out of a dump previously — dump the graph and throw away
what does not match, in the consumer:
graph=results p=s66#doi -> 337,811 rows
before 375.8 MB · 1797 req · 2105.8 MB RSS · 12.8 s · first row 409 ms
after 16.0 MB · 105 req · 182.6 MB RSS · 0.41 s · first row 14 ms
graph=results p=s66#isbn -> 40,761 rows
before 375.8 MB · 1797 req · 2105.7 MB RSS · 12.8 s
after 15.5 MB · 104 req · 155.2 MB RSS · 0.23 s§Where it does NOT help
The index is pruned; the dictionary is not. Term resolution still faults whichever chunks the surviving rows’ terms live in, and on a file whose payload is the literals those chunks are most of the file. Same file, a predicate whose objects are the long abstracts:
graph=projects p=s66#abstract -> 80,206 rows
before 260.7 MB · 1142 req · 1496.9 MB RSS · 5.97 s
after 213.4 MB · 1606 req · 1144.4 MB RSS · 2.91 s (1.2x, not 23x)An unfiltered dump is unchanged by construction — it is the floor, and
its peak is the resident dictionary, not this. Two ceilings sit above
both: a faulted dictionary chunk is a OnceCell nothing ever evicts, and
on a literal-heavy file the chunk directory alone can be a third of the
file before any dump work at all (#198).
§Permutations
Every one of the eight bound/unbound shapes routes inside
PermSet::CORE — {SPO, POS, OSP}, the minimum a legal file carries
(see perm_routing_never_leaves_core). A file built with
--permutations 3 therefore prunes identically to a six-permutation
one: same tiles, same rows, no fallback path to get wrong.
A bound term the dictionary does not know, or an absent graph IRI, yields nothing without touching the index.
§Order
Rows arrive in the routed permutation’s order, which is the price of
routing at all. Unfiltered (and subject-bound) that is SPO — canonical
(s, p, o), unchanged from dump_each. A bound
predicate routes to POS and streams (p, o, s); a bound object routes to
OSP. The set is identical to the unfiltered dump filtered; the order is
not, so a consumer that needs canonical order must sort. N-Quads, the
format rete export writes, is order-independent.
Sourcepub fn dump_plan(
&self,
graph: Option<&str>,
s: Option<&str>,
p: Option<&str>,
o: Option<&str>,
) -> DumpPlan
pub fn dump_plan( &self, graph: Option<&str>, s: Option<&str>, p: Option<&str>, o: Option<&str>, ) -> DumpPlan
What a dump_filtered_each over this graph
and pattern will fetch, before it starts — the dump twin of
rete cost’s query preview.
Costed from the tile directories, so it fetches no tile. It does read the dictionary far enough to resolve the bound terms (that is a real, unavoidable cost of the dump too) and to learn whether the graph exists.
Sourcepub fn dump_iter(
&self,
graph: Option<&str>,
) -> impl Iterator<Item = (String, String, String)>
pub fn dump_iter( &self, graph: Option<&str>, ) -> impl Iterator<Item = (String, String, String)>
The pull twin of Rete::dump_each: every triple of a
graph (None = default) as a lazy iterator of resolved terms.
Same constant-memory scan — one triple decoded and resolved per next(),
never a Vec of the whole graph — but the caller drives it, so it can
stop early or be suspended and resumed across a foreign-function
boundary. That is what the wasm/JS client’s batched quad cursor needs:
a callback cannot be paused mid-scan to hand control back to JavaScript,
an iterator can.
The dictionary is NOT prefetched: terms fault in as they are resolved, so
stopping after a handful of triples costs a handful of dictionary reads
rather than the whole dictionary — which on a lazy remote open of a large
file would be gigabytes before the first triple, defeating the point of a
pull API. dump_each does not prefetch it whole
either; it batches the same per-term faults. Peak memory is
O(faulted dictionary chunks + faulted index tiles), not O(triples).
Sourcepub fn query_iter(
&self,
graph: Option<&str>,
s: Option<&str>,
p: Option<&str>,
o: Option<&str>,
) -> impl Iterator<Item = (String, String, String)>
pub fn query_iter( &self, graph: Option<&str>, s: Option<&str>, p: Option<&str>, o: Option<&str>, ) -> impl Iterator<Item = (String, String, String)>
dump_iter generalized to a triple pattern: the
lazy pull form of query_in_graph. A bound term
unknown to the dictionary, or an unknown graph IRI, yields nothing.
Same constant-memory guarantee — one triple decoded and resolved per
next(), never a Vec of every match — and the same absence of a
dictionary prefetch, so stopping after a handful of matches costs a
handful of dictionary reads. query_in_graph is the eager twin; prefer
this one wherever the consumer can stop early (ASK, LIMIT, an RDF4J
getStatements that is about to be sliced).
Sourcepub fn query_batch(
&self,
graph: Option<&str>,
s: Option<&str>,
p: Option<&str>,
o: Option<&str>,
cursor: u64,
max_quads: usize,
) -> (Vec<(String, String, String)>, u64, bool)
pub fn query_batch( &self, graph: Option<&str>, s: Option<&str>, p: Option<&str>, o: Option<&str>, cursor: u64, max_quads: usize, ) -> (Vec<(String, String, String)>, u64, bool)
One bounded, resumable slice of a pattern’s matches inside one graph
(None = default): (triples, next_cursor, done). Start at cursor = 0
and feed the returned cursor back until done.
This is query_in_graph made streamable across a
boundary that cannot hold a Rust borrow — the wasm/JVM one. The engine
keeps no iterator alive between calls; the entire resume state is the
opaque u64 (see GraphIndex::scan_batch), so a client’s cursor
survives being suspended, and a cursor that is abandoned leaks nothing
but a u64.
The dictionary is faulted per batch, not whole: taking ten matches
off the front of a 9.8-billion-quad file reads ten matches’ worth of
terms. That is the difference between LIMIT 1 answering in bounded time
and it reading a 30 GB dictionary first.
max_quads is a floor (batches end on an a-group boundary), and every
call either returns at least one row or reports done.
Sourcepub fn dump_batch(
&self,
graph: Option<&str>,
cursor: u32,
max_quads: usize,
) -> (Vec<(String, String, String)>, u32, bool)
pub fn dump_batch( &self, graph: Option<&str>, cursor: u32, max_quads: usize, ) -> (Vec<(String, String, String)>, u32, bool)
Resolve one bounded slice of a graph’s triples. Returns
(triples, next_cursor, done); start with cursor = 0 and feed the
returned cursor back until done.
This is the PULL half of the dump, and it exists because Rete::dump_each is
push-based: it drives the scan itself and hands each triple to a
callback. A client that wants to pull (a Python generator, a JS async
iterator) has to hold the scan’s stack between calls — which means a
thread (unavailable on Pyodide and in wasm) or a self-referential struct
holding an iterator that borrows the Rete stored next to it, whose
soundness rests on drop order plus an unstated promise that this crate’s
lazily-faulted caches stay write-once. Neither is worth it for a dump.
Instead the scan is made RESUMABLE: SPO tiles are ordered by subject id,
so “every triple of subject sid, for ascending sid” visits exactly the
tiles a full scan visits, in the same order, and the whole resume state
collapses to one u32. No thread, no unsafe, no borrow held across
calls — and unlike (offset, limit) it is O(n) overall rather than
O(n²/limit), because nothing is ever re-scanned.
max_quads is the FLOOR of a batch, not a hard cut: a batch always ends
on a subject boundary, so no subject is split across two calls. Work per
call is bounded by a probe budget too — a sparse named graph whose few
subjects are spread over a huge id space returns early with done = false
rather than grinding through absent ids, and the caller just asks again.
Sourcepub fn named_graph_count(&self) -> usize
pub fn named_graph_count(&self) -> usize
How many named graphs this dataset has. On a lazy ranged open this reads only the section’s leading count varint — never the graphs.
Sourcepub fn named_graph_name_at(&self, i: usize) -> Option<&str>
pub fn named_graph_name_at(&self, i: usize) -> Option<&str>
The i-th named graph’s IRI (stored order). On a lazy ranged open this
walks entry HEADERS up to i — it never decodes a graph’s index.
Sourcepub fn named_graph_at(&self, i: usize) -> Option<(&str, &GraphIndex)>
pub fn named_graph_at(&self, i: usize) -> Option<(&str, &GraphIndex)>
The i-th named graph as (iri, index). On a lazy ranged open the
index is fetched and decoded on first access and memoised; check
index_incomplete after evaluating.
Sourcepub fn graph_names(&self) -> Vec<&str>
pub fn graph_names(&self) -> Vec<&str>
IRIs of the named graphs in this dataset (the default graph is unnamed).
On a lazy ranged open this walks the whole directory (headers only) —
prefer named_graph_count when the number
is all that’s needed.
Sourcepub fn graph_index(&self, iri: &str) -> Option<&GraphIndex>
pub fn graph_index(&self, iri: &str) -> Option<&GraphIndex>
The permutation index of a named graph, or None if absent.
Sourcepub fn match_ids(
&self,
pattern: (Option<u32>, Option<u32>, Option<u32>),
) -> Vec<(u32, u32, u32)>
pub fn match_ids( &self, pattern: (Option<u32>, Option<u32>, Option<u32>), ) -> Vec<(u32, u32, u32)>
Match a triple pattern in dictionary-ID space (subject/predicate/object IDs), returning integer triples — the fast path used by the BGP engine.
Sourcepub fn predicate_pairs(&self, predicate: &str) -> Vec<(u32, u32)>
pub fn predicate_pairs(&self, predicate: &str) -> Vec<(u32, u32)>
All (subject_node, object_node) pairs for a predicate, as unified node
IDs — no term resolution. The fast path for graph traversal.
Sourcepub fn open_ranged<R>(reader: &R) -> Result<Rete, FileError>where
R: RangeReader,
pub fn open_ranged<R>(reader: &R) -> Result<Rete, FileError>where
R: RangeReader,
Open via a RangeReader, fetching only the header and the named
section ranges — never a linear scan of the whole resource. A full query
open touches at most 4 ranges (header, dictionary, index, pyramid-meta).
Sourcepub fn open_ranged_lazy<R>(reader: R) -> Result<Rete, FileError>
pub fn open_ranged_lazy<R>(reader: R) -> Result<Rete, FileError>
Open via an owned RangeReader with lazy tile faulting (tiled
v0.2 files): fetches the header, dictionary, pyramid meta, named graphs,
and each permutation’s tile directory — but no default-graph tile
payloads. Tiles fault in (one range request each) the first time a scan
touches them, so a selective SPARQL query fetches O(touched tiles)
bytes instead of the whole index.
Failure contract: scans are infallible by design, so a failed tile
fetch yields an empty tile and sets a sticky flag — after evaluating,
callers MUST check index_incomplete and
surface an error instead of the (possibly partial) results.
Sourcepub fn index_incomplete(&self) -> bool
pub fn index_incomplete(&self) -> bool
Did any lazy fetch (index tile or dictionary chunk) fail since this
Rete was opened? When true, query results may be silently incomplete —
callers using Rete::open_ranged_lazy must check this after
evaluating and turn it into an error.
Sourcepub fn reset_load_failures(&self)
pub fn reset_load_failures(&self)
Forget recorded lazy-fetch failures — the start-of-evaluation reset for
a RESIDENT session (a browser worker holding one Rete across many
queries): it makes index_incomplete a
per-query verdict instead of a per-open one, so a single transient
network failure no longer fails every subsequent query on the session.
Sound because failed tiles/chunks are never cached — the next
evaluation simply retries the fetch.
Sourcepub fn set_memory_budget(&self, budget: Option<u64>) -> MemoryBudget
pub fn set_memory_budget(&self, budget: Option<u64>) -> MemoryBudget
Bound peak resident memory to roughly budget bytes by capping the
evictable dictionary chunk cache and the default-graph index tile cache
(bounded-export phase 2). None or u64::MAX means unlimited — nothing
is evicted and residency is identical to a plain open, which is what
--in-memory and every query/wasm path use.
Sound at any cap: the dictionary and the default-graph index both fault
their chunks/tiles through a loader, so an evicted body is simply
re-faulted on its next touch; the export dump reads each needed body
while holding it (the cache Arc handout invariant), so eviction never
drops bytes in use. Named-graph indexes are not capped here — a small
resident graph has no loader and could never reload an evicted tile — so
they stay unlimited and are released whole per slot instead (see
release_named_graph).
Returns the split actually applied, for RETE_OPEN_DEBUG logging.
Sourcepub fn dict_cache_stats(&self) -> CacheStats
pub fn dict_cache_stats(&self) -> CacheStats
The dictionary chunk cache’s observability counters (decoded chunks and
bytes, hits/misses, evictions, resident bytes) — for the bounded-export
work-accounting tests and RETE_OPEN_DEBUG.
Sourcepub fn index_cache_stats(&self) -> CacheStats
pub fn index_cache_stats(&self) -> CacheStats
The default-graph index tile cache’s observability counters.
Sourcepub fn dict_cache_cap(&self) -> u64
pub fn dict_cache_cap(&self) -> u64
The current dictionary chunk-cache byte cap (u64::MAX = unlimited).
Sourcepub fn release_named_graph(&mut self, iri: &str)
pub fn release_named_graph(&mut self, iri: &str)
Release a named graph’s decoded index after its slot has been exported,
dropping its GraphIndex (and, with it, that graph’s decoded tiles and
its own tile cache) so a many-graph dump does not accumulate one resident
index per graph — the next linear floor on many-graph files once the
dictionary is bounded (plan §6). A no-op for the default graph, an
unknown IRI, a resident (in-memory/eager) open, or a graph never opened.
Takes &mut self: the export loop finishes one slot’s dump (which only
borrows &self) before releasing, so no outstanding borrow of the index
exists. Correct because a released index is simply re-opened (headers +
tile directory) on a later access — the same lazy contract as a chunk
that was never faulted.
Sourcepub fn query_with_provenance(
&self,
s: Option<&str>,
p: Option<&str>,
o: Option<&str>,
) -> Vec<TripleProvenance>
pub fn query_with_provenance( &self, s: Option<&str>, p: Option<&str>, o: Option<&str>, ) -> Vec<TripleProvenance>
Evaluate a triple pattern and include the file/index provenance for every matched result. A bound term that is unknown to the dictionary yields no matches.
Sourcepub fn query(
&self,
s: Option<&str>,
p: Option<&str>,
o: Option<&str>,
) -> Vec<(String, String, String)>
pub fn query( &self, s: Option<&str>, p: Option<&str>, o: Option<&str>, ) -> Vec<(String, String, String)>
Evaluate a triple pattern given as optional term strings, returning matching triples resolved back to terms. A bound term that is unknown to the dictionary yields no matches.
Sourcepub fn query_in_graph(
&self,
graph: Option<&str>,
s: Option<&str>,
p: Option<&str>,
o: Option<&str>,
) -> Vec<(String, String, String)>
pub fn query_in_graph( &self, graph: Option<&str>, s: Option<&str>, p: Option<&str>, o: Option<&str>, ) -> Vec<(String, String, String)>
Match a triple pattern within a single graph — None is the default
graph, Some(iri) a named graph — resolving matches to canonical terms.
This is Rete::query (default-graph only) generalized to any graph: the
graph-scoped primitive a quad-aware consumer (e.g. an RDF4J Sail’s
getStatements) needs. An unknown graph IRI, or a bound term absent from
the shared dictionary, yields an empty result. All graphs share one
dictionary, so the pattern resolves once against that ID space.
Sourcepub fn query_quads(
&self,
s: Option<&str>,
p: Option<&str>,
o: Option<&str>,
) -> Vec<((String, String, String), Option<String>)>
pub fn query_quads( &self, s: Option<&str>, p: Option<&str>, o: Option<&str>, ) -> Vec<((String, String, String), Option<String>)>
Match a triple pattern across the default graph and every named graph,
tagging each match with its graph (None = default). The quad-level
companion to Rete::query; default-graph matches come first, then each
named graph in stored order.
Sourcepub fn query_ranged<R>(
reader: &R,
s: Option<&str>,
p: Option<&str>,
o: Option<&str>,
) -> Result<Vec<(String, String, String)>, FileError>where
R: RangeReader,
pub fn query_ranged<R>(
reader: &R,
s: Option<&str>,
p: Option<&str>,
o: Option<&str>,
) -> Result<Vec<(String, String, String)>, FileError>where
R: RangeReader,
Evaluate one triple pattern through a RangeReader by fetching only
the header, the dictionary, and — for a tiled (v0.2) file — the
selected permutation section’s tile directory plus the tile(s) the
bound leading id routes to; an unbound leading id fetches the section’s
tile body in one request. v0.1 files fetch the whole selected section.
Unknown bound terms return an empty result before touching the index.
Sourcepub fn route_pattern_ranged<R>(
reader: &R,
s: Option<&str>,
p: Option<&str>,
o: Option<&str>,
) -> Result<bool, FileError>where
R: RangeReader,
pub fn route_pattern_ranged<R>(
reader: &R,
s: Option<&str>,
p: Option<&str>,
o: Option<&str>,
) -> Result<bool, FileError>where
R: RangeReader,
Route one triple pattern to its permutation section without fetching
any payload bytes. Returns false when a bound term is unknown and the
index was skipped.