Skip to main content

Rete

Struct Rete 

Source
pub struct Rete { /* private fields */ }
Expand description

A read-only, in-memory view over a .rete file image.

Implementations§

Source§

impl Rete

Source

pub fn open(bytes: &[u8]) -> Result<Rete, FileError>

Parse a full file image (v0 loads everything; a range-reading client will fetch only the sections it needs — same container format).

Source

pub fn set_service_client(&mut self, client: Box<dyn ServiceClient>)

Attach the client that executes SERVICE <endpoint> { … } blocks (SPARQL 1.1 federated query) against remote SPARQL endpoints. Without one, a non-SILENT SERVICE fails the query with a clear error and a SERVICE SILENT degrades to one empty solution, per the spec.

Source

pub fn header(&self) -> &Header

Source

pub fn file_layout(&self) -> Vec<LayoutSegment>

The file’s byte layout, for visualization: header, metadata, dictionary, each index permutation’s tile directory and individual tiles, pyramid summary, and named graphs — sorted by offset. Bytes not covered by any segment are container framing (section directories and length fields).

Source

pub fn metadata(&self) -> Option<&[u8]>

Raw bytes of the file’s metadata section, or None if it has none. The CLI stores a JSON Dataset Card here; rete-core treats it as opaque. Populated by Rete::open only — an Rete::open_ranged view returns None here (the card is not fetched on the minimal query path).

Source

pub fn dictionary(&self) -> &Dictionary

Source

pub fn pyramid(&self) -> Option<&PyramidMeta>

The pyramid metadata (summary graph + tiles), if the file has a pyramid.

Source

pub fn pyramid_if_loaded(&self) -> Option<&PyramidMeta>

The pyramid metadata only if already resident or previously faulted — never triggers a lazy range read. The query planner uses this for cardinality estimation so it is free for an in-memory file and never adds a fetch on the lazy remote path (which defers the pyramid by design).

Source

pub fn predicate_stats(&self) -> &[PredStat]

Per-predicate planner statistics from the query-stats block — empty when the file has none or the pyramid isn’t resident (the lazy path doesn’t fault it just for stats). See crate::meta::PredStat.

Source

pub fn char_sets(&self) -> &[CharSet]

The entity shapes (characteristic sets) from the pyramid — empty when the file has none or the pyramid isn’t resident. See crate::meta::CharSet.

Source

pub fn label_index(&self) -> &[LabelEntry]

The label index from the pyramid — empty when the file has none or the pyramid isn’t resident. See crate::meta::LabelEntry.

Prefix-search the label index: the subjects whose label starts with prefix (case-insensitive), as (label, subject_iri), capped at limit. Unlike the planner accessors, this faults the pyramid (where the index lives) on the lazy path — a prefix search is an explicit read, not a free estimate. Returns an empty vec when the file carries no label index.

Source

pub fn has_text_index(&self) -> bool

Whether this file carries a full-text (TEXT_INDEX) section, i.e. it was built with --text-index. Cheap — reads the header, never faults.

Source

pub fn text_index_token_table_len(&self) -> Option<u64>

Byte length of the TEXT_INDEX section’s leading token table — the length varint plus the compressed table, which is exactly what a first text_search faults. None when the file carries no text index (or its head could not be read).

This is the number to quote as the cost of a first search, NOT header().text_index_len: that counts the postings blob too, which is the bulk of the section and is only ever fetched one posting list at a time. On the published epfl-infoscience.rete the section is 195 MB and its token table 29 MB — a 6.5× difference.

Costs nothing on a resident open (measured while the bytes were in hand); a lazy/ranged open pays one ≤10-byte range read, memoized, and still never faults the table itself.

Full-text search over the literals: subject IRIs that carry every word in words (whole-word, case-insensitive — AND semantics), optionally also requiring a word that starts with prefix (token-prefix). Results are ordered by subject id and capped at limit (0 = uncapped). Empty when the file has no text index or nothing matches.

Like prefix_search, this faults the index on the lazy remote path — a search is an explicit read, and only the queried posting lists are fetched, not the whole index.

Source

pub fn default_index(&self) -> &GraphIndex

The default-graph permutation index.

Source

pub fn dump(&self, graph: Option<&str>) -> Vec<(String, String, String)>

Resolve every triple of a graph (None = default graph) back to terms.

Source

pub fn dump_each<F>(&self, graph: Option<&str>, f: F)
where F: FnMut(&str, &str, &str),

Stream every triple of a graph (None = default) to f, resolving terms a batch at a time — no full Vec materialization, so it is safe on graphs far larger than RAM. rete export uses this to serialize 100M+ triple files that dump() (which collects every term into a Vec<String>) would OOM on.

On a lazy/ranged open the dictionary is not prefetched whole: each batch faults only the chunks its own terms live in. Dumping ONE named graph therefore costs that graph’s terms, not the file’s. Dumping everything still reads every dictionary byte and leaves every chunk resident — that is inherent, and it, not the prefetch, is what sets the peak of a full export. See the comment in the body for both measurements.

To dump a slice — one predicate, one subject, one object — call dump_filtered_each, which is this method with a triple pattern and prunes tiles instead of filtering rows.

Source

pub fn dump_filtered_each<F>( &self, graph: Option<&str>, s: Option<&str>, p: Option<&str>, o: Option<&str>, f: F, )
where F: FnMut(&str, &str, &str),

dump_each restricted to a triple pattern — the filtered dump. Any of s / p / o may be bound (canonical N-Triples term tokens); all three unbound is exactly dump_each.

§Why this is cheap, and how much

It is the same access path a routed query takes, not a full scan with a predicate test bolted on: GraphIndex::scan_iter picks the permutation with the longest bound prefix, binary-searches that section’s tile directory down to the span a bound leading id can live in, and then drops every remaining tile whose recorded synopsis proves it cannot hold a bound secondary component — all from directories, without fetching a tile. So a filtered dump pays for the slice, which is the whole point of issue #117.

Measured on cordis.rete (801 MB, a 417 MB dictionary, 26.4M quads, six named graphs), lazily opened; “before” is the only way to get a predicate slice out of a dump previously — dump the graph and throw away what does not match, in the consumer:

  graph=results  p=s66#doi     -> 337,811 rows
    before  375.8 MB · 1797 req · 2105.8 MB RSS · 12.8 s · first row 409 ms
    after    16.0 MB ·  105 req ·  182.6 MB RSS ·  0.41 s · first row  14 ms
  graph=results  p=s66#isbn    ->  40,761 rows
    before  375.8 MB · 1797 req · 2105.7 MB RSS · 12.8 s
    after    15.5 MB ·  104 req ·  155.2 MB RSS ·  0.23 s
§Where it does NOT help

The index is pruned; the dictionary is not. Term resolution still faults whichever chunks the surviving rows’ terms live in, and on a file whose payload is the literals those chunks are most of the file. Same file, a predicate whose objects are the long abstracts:

  graph=projects p=s66#abstract ->  80,206 rows
    before  260.7 MB · 1142 req · 1496.9 MB RSS · 5.97 s
    after   213.4 MB · 1606 req · 1144.4 MB RSS · 2.91 s   (1.2x, not 23x)

An unfiltered dump is unchanged by construction — it is the floor, and its peak is the resident dictionary, not this. Two ceilings sit above both: a faulted dictionary chunk is a OnceCell nothing ever evicts, and on a literal-heavy file the chunk directory alone can be a third of the file before any dump work at all (#198).

§Permutations

Every one of the eight bound/unbound shapes routes inside PermSet::CORE — {SPO, POS, OSP}, the minimum a legal file carries (see perm_routing_never_leaves_core). A file built with --permutations 3 therefore prunes identically to a six-permutation one: same tiles, same rows, no fallback path to get wrong.

A bound term the dictionary does not know, or an absent graph IRI, yields nothing without touching the index.

§Order

Rows arrive in the routed permutation’s order, which is the price of routing at all. Unfiltered (and subject-bound) that is SPO — canonical (s, p, o), unchanged from dump_each. A bound predicate routes to POS and streams (p, o, s); a bound object routes to OSP. The set is identical to the unfiltered dump filtered; the order is not, so a consumer that needs canonical order must sort. N-Quads, the format rete export writes, is order-independent.

Source

pub fn dump_plan( &self, graph: Option<&str>, s: Option<&str>, p: Option<&str>, o: Option<&str>, ) -> DumpPlan

What a dump_filtered_each over this graph and pattern will fetch, before it starts — the dump twin of rete cost’s query preview.

Costed from the tile directories, so it fetches no tile. It does read the dictionary far enough to resolve the bound terms (that is a real, unavoidable cost of the dump too) and to learn whether the graph exists.

Source

pub fn dump_iter( &self, graph: Option<&str>, ) -> impl Iterator<Item = (String, String, String)>

The pull twin of Rete::dump_each: every triple of a graph (None = default) as a lazy iterator of resolved terms.

Same constant-memory scan — one triple decoded and resolved per next(), never a Vec of the whole graph — but the caller drives it, so it can stop early or be suspended and resumed across a foreign-function boundary. That is what the wasm/JS client’s batched quad cursor needs: a callback cannot be paused mid-scan to hand control back to JavaScript, an iterator can.

The dictionary is NOT prefetched: terms fault in as they are resolved, so stopping after a handful of triples costs a handful of dictionary reads rather than the whole dictionary — which on a lazy remote open of a large file would be gigabytes before the first triple, defeating the point of a pull API. dump_each does not prefetch it whole either; it batches the same per-term faults. Peak memory is O(faulted dictionary chunks + faulted index tiles), not O(triples).

Source

pub fn query_iter( &self, graph: Option<&str>, s: Option<&str>, p: Option<&str>, o: Option<&str>, ) -> impl Iterator<Item = (String, String, String)>

dump_iter generalized to a triple pattern: the lazy pull form of query_in_graph. A bound term unknown to the dictionary, or an unknown graph IRI, yields nothing.

Same constant-memory guarantee — one triple decoded and resolved per next(), never a Vec of every match — and the same absence of a dictionary prefetch, so stopping after a handful of matches costs a handful of dictionary reads. query_in_graph is the eager twin; prefer this one wherever the consumer can stop early (ASK, LIMIT, an RDF4J getStatements that is about to be sliced).

Source

pub fn query_batch( &self, graph: Option<&str>, s: Option<&str>, p: Option<&str>, o: Option<&str>, cursor: u64, max_quads: usize, ) -> (Vec<(String, String, String)>, u64, bool)

One bounded, resumable slice of a pattern’s matches inside one graph (None = default): (triples, next_cursor, done). Start at cursor = 0 and feed the returned cursor back until done.

This is query_in_graph made streamable across a boundary that cannot hold a Rust borrow — the wasm/JVM one. The engine keeps no iterator alive between calls; the entire resume state is the opaque u64 (see GraphIndex::scan_batch), so a client’s cursor survives being suspended, and a cursor that is abandoned leaks nothing but a u64.

The dictionary is faulted per batch, not whole: taking ten matches off the front of a 9.8-billion-quad file reads ten matches’ worth of terms. That is the difference between LIMIT 1 answering in bounded time and it reading a 30 GB dictionary first.

max_quads is a floor (batches end on an a-group boundary), and every call either returns at least one row or reports done.

Source

pub fn dump_batch( &self, graph: Option<&str>, cursor: u32, max_quads: usize, ) -> (Vec<(String, String, String)>, u32, bool)

Resolve one bounded slice of a graph’s triples. Returns (triples, next_cursor, done); start with cursor = 0 and feed the returned cursor back until done.

This is the PULL half of the dump, and it exists because Rete::dump_each is push-based: it drives the scan itself and hands each triple to a callback. A client that wants to pull (a Python generator, a JS async iterator) has to hold the scan’s stack between calls — which means a thread (unavailable on Pyodide and in wasm) or a self-referential struct holding an iterator that borrows the Rete stored next to it, whose soundness rests on drop order plus an unstated promise that this crate’s lazily-faulted caches stay write-once. Neither is worth it for a dump.

Instead the scan is made RESUMABLE: SPO tiles are ordered by subject id, so “every triple of subject sid, for ascending sid” visits exactly the tiles a full scan visits, in the same order, and the whole resume state collapses to one u32. No thread, no unsafe, no borrow held across calls — and unlike (offset, limit) it is O(n) overall rather than O(n²/limit), because nothing is ever re-scanned.

max_quads is the FLOOR of a batch, not a hard cut: a batch always ends on a subject boundary, so no subject is split across two calls. Work per call is bounded by a probe budget too — a sparse named graph whose few subjects are spread over a huge id space returns early with done = false rather than grinding through absent ids, and the caller just asks again.

Source

pub fn named_graph_count(&self) -> usize

How many named graphs this dataset has. On a lazy ranged open this reads only the section’s leading count varint — never the graphs.

Source

pub fn named_graph_name_at(&self, i: usize) -> Option<&str>

The i-th named graph’s IRI (stored order). On a lazy ranged open this walks entry HEADERS up to i — it never decodes a graph’s index.

Source

pub fn named_graph_at(&self, i: usize) -> Option<(&str, &GraphIndex)>

The i-th named graph as (iri, index). On a lazy ranged open the index is fetched and decoded on first access and memoised; check index_incomplete after evaluating.

Source

pub fn graph_names(&self) -> Vec<&str>

IRIs of the named graphs in this dataset (the default graph is unnamed). On a lazy ranged open this walks the whole directory (headers only) — prefer named_graph_count when the number is all that’s needed.

Source

pub fn graph_index(&self, iri: &str) -> Option<&GraphIndex>

The permutation index of a named graph, or None if absent.

Source

pub fn match_ids( &self, pattern: (Option<u32>, Option<u32>, Option<u32>), ) -> Vec<(u32, u32, u32)>

Match a triple pattern in dictionary-ID space (subject/predicate/object IDs), returning integer triples — the fast path used by the BGP engine.

Source

pub fn predicate_pairs(&self, predicate: &str) -> Vec<(u32, u32)>

All (subject_node, object_node) pairs for a predicate, as unified node IDs — no term resolution. The fast path for graph traversal.

Source

pub fn open_ranged<R>(reader: &R) -> Result<Rete, FileError>
where R: RangeReader,

Open via a RangeReader, fetching only the header and the named section ranges — never a linear scan of the whole resource. A full query open touches at most 4 ranges (header, dictionary, index, pyramid-meta).

Source

pub fn open_ranged_lazy<R>(reader: R) -> Result<Rete, FileError>
where R: RangeReader + Send + Sync + 'static,

Open via an owned RangeReader with lazy tile faulting (tiled v0.2 files): fetches the header, dictionary, pyramid meta, named graphs, and each permutation’s tile directory — but no default-graph tile payloads. Tiles fault in (one range request each) the first time a scan touches them, so a selective SPARQL query fetches O(touched tiles) bytes instead of the whole index.

Failure contract: scans are infallible by design, so a failed tile fetch yields an empty tile and sets a sticky flag — after evaluating, callers MUST check index_incomplete and surface an error instead of the (possibly partial) results.

Source

pub fn index_incomplete(&self) -> bool

Did any lazy fetch (index tile or dictionary chunk) fail since this Rete was opened? When true, query results may be silently incomplete — callers using Rete::open_ranged_lazy must check this after evaluating and turn it into an error.

Source

pub fn reset_load_failures(&self)

Forget recorded lazy-fetch failures — the start-of-evaluation reset for a RESIDENT session (a browser worker holding one Rete across many queries): it makes index_incomplete a per-query verdict instead of a per-open one, so a single transient network failure no longer fails every subsequent query on the session. Sound because failed tiles/chunks are never cached — the next evaluation simply retries the fetch.

Source

pub fn set_memory_budget(&self, budget: Option<u64>) -> MemoryBudget

Bound peak resident memory to roughly budget bytes by capping the evictable dictionary chunk cache and the default-graph index tile cache (bounded-export phase 2). None or u64::MAX means unlimited — nothing is evicted and residency is identical to a plain open, which is what --in-memory and every query/wasm path use.

Sound at any cap: the dictionary and the default-graph index both fault their chunks/tiles through a loader, so an evicted body is simply re-faulted on its next touch; the export dump reads each needed body while holding it (the cache Arc handout invariant), so eviction never drops bytes in use. Named-graph indexes are not capped here — a small resident graph has no loader and could never reload an evicted tile — so they stay unlimited and are released whole per slot instead (see release_named_graph).

Returns the split actually applied, for RETE_OPEN_DEBUG logging.

Source

pub fn dict_cache_stats(&self) -> CacheStats

The dictionary chunk cache’s observability counters (decoded chunks and bytes, hits/misses, evictions, resident bytes) — for the bounded-export work-accounting tests and RETE_OPEN_DEBUG.

Source

pub fn index_cache_stats(&self) -> CacheStats

The default-graph index tile cache’s observability counters.

Source

pub fn dict_cache_cap(&self) -> u64

The current dictionary chunk-cache byte cap (u64::MAX = unlimited).

Source

pub fn release_named_graph(&mut self, iri: &str)

Release a named graph’s decoded index after its slot has been exported, dropping its GraphIndex (and, with it, that graph’s decoded tiles and its own tile cache) so a many-graph dump does not accumulate one resident index per graph — the next linear floor on many-graph files once the dictionary is bounded (plan §6). A no-op for the default graph, an unknown IRI, a resident (in-memory/eager) open, or a graph never opened.

Takes &mut self: the export loop finishes one slot’s dump (which only borrows &self) before releasing, so no outstanding borrow of the index exists. Correct because a released index is simply re-opened (headers + tile directory) on a later access — the same lazy contract as a chunk that was never faulted.

Source

pub fn query_with_provenance( &self, s: Option<&str>, p: Option<&str>, o: Option<&str>, ) -> Vec<TripleProvenance>

Evaluate a triple pattern and include the file/index provenance for every matched result. A bound term that is unknown to the dictionary yields no matches.

Source

pub fn query( &self, s: Option<&str>, p: Option<&str>, o: Option<&str>, ) -> Vec<(String, String, String)>

Evaluate a triple pattern given as optional term strings, returning matching triples resolved back to terms. A bound term that is unknown to the dictionary yields no matches.

Source

pub fn query_in_graph( &self, graph: Option<&str>, s: Option<&str>, p: Option<&str>, o: Option<&str>, ) -> Vec<(String, String, String)>

Match a triple pattern within a single graph — None is the default graph, Some(iri) a named graph — resolving matches to canonical terms. This is Rete::query (default-graph only) generalized to any graph: the graph-scoped primitive a quad-aware consumer (e.g. an RDF4J Sail’s getStatements) needs. An unknown graph IRI, or a bound term absent from the shared dictionary, yields an empty result. All graphs share one dictionary, so the pattern resolves once against that ID space.

Source

pub fn query_quads( &self, s: Option<&str>, p: Option<&str>, o: Option<&str>, ) -> Vec<((String, String, String), Option<String>)>

Match a triple pattern across the default graph and every named graph, tagging each match with its graph (None = default). The quad-level companion to Rete::query; default-graph matches come first, then each named graph in stored order.

Source

pub fn query_ranged<R>( reader: &R, s: Option<&str>, p: Option<&str>, o: Option<&str>, ) -> Result<Vec<(String, String, String)>, FileError>
where R: RangeReader,

Evaluate one triple pattern through a RangeReader by fetching only the header, the dictionary, and — for a tiled (v0.2) file — the selected permutation section’s tile directory plus the tile(s) the bound leading id routes to; an unbound leading id fetches the section’s tile body in one request. v0.1 files fetch the whole selected section. Unknown bound terms return an empty result before touching the index.

Source

pub fn route_pattern_ranged<R>( reader: &R, s: Option<&str>, p: Option<&str>, o: Option<&str>, ) -> Result<bool, FileError>
where R: RangeReader,

Route one triple pattern to its permutation section without fetching any payload bytes. Returns false when a bound term is unknown and the index was skipped.

Auto Trait Implementations§

§

impl !Freeze for Rete

§

impl !RefUnwindSafe for Rete

§

impl !UnwindSafe for Rete

§

impl Send for Rete

§

impl Sync for Rete

§

impl Unpin for Rete

§

impl UnsafeUnpin for Rete

Blanket Implementations§

Source§

impl<T> Any for T
where T: 'static + ?Sized,

Source§

fn type_id(&self) -> TypeId

Gets the TypeId of self. Read more
Source§

impl<T> Borrow<T> for T
where T: ?Sized,

Source§

fn borrow(&self) -> &T

Immutably borrows from an owned value. Read more
Source§

impl<T> BorrowMut<T> for T
where T: ?Sized,

Source§

fn borrow_mut(&mut self) -> &mut T

Mutably borrows from an owned value. Read more
Source§

impl<ST, DT> CastableFrom<ST, Initialized, Initialized> for DT
where ST: ?Sized, DT: ?Sized,

Source§

impl<ST, DT> CastableFrom<ST, Uninit, Uninit> for DT
where ST: ?Sized, DT: ?Sized,

Source§

impl<T> From<T> for T

Source§

fn from(t: T) -> T

Returns the argument unchanged.

Source§

impl<T, U> Into<U> for T
where U: From<T>,

Source§

fn into(self) -> U

Calls U::from(self).

That is, this conversion is whatever the implementation of From<T> for U chooses to do.

Source§

impl<T> Read<Exclusive, BecauseExclusive> for T
where T: ?Sized,

Source§

impl<T> Same for T

Source§

type Output = T

Should always be Self
Source§

impl<T, U> TryFrom<U> for T
where U: Into<T>,

Source§

type Error = !

The type returned in the event of a conversion error.
Source§

fn try_from(value: U) -> Result<T, !>

Performs the conversion.
Source§

impl<T, U> TryInto<U> for T
where U: TryFrom<T>,

Source§

type Error = <U as TryFrom<T>>::Error

The type returned in the event of a conversion error.
Source§

fn try_into(self) -> Result<U, <U as TryFrom<T>>::Error>

Performs the conversion.
Source§

impl<V, T> VZip<V> for T
where V: MultiLane<T>,

Source§

fn vzip(self) -> V