Skip to main content

Crate rete_graph

Crate rete_graph 

Source
Expand description

rete-graph — the Rust library for Rete graph files, under the name it already has on PyPI and npm.

This crate is a facade: every item is re-exported from rete_core, which is where the engine actually lives and where the implementation documentation belongs. Depending on either gets you the same code; this one exists so that

pip install rete-graph      # Python
npm  install rete-graph     # JavaScript
cargo add    rete-graph     # Rust

all name the same thing. Before it existed, a Rust user following any of the project’s other install instructions had to know that the crate was called something else.

§What this is NOT

It does not pull in the whole workspace, because the three published crates are not interchangeable:

  • rete-core is the library — what a Rust program depends on, and what this crate re-exports.
  • rete-cli is a binary. You install it (cargo install rete-cli); depending on it from a library would drag an executable into your build for nothing.
  • rete-wasm is the browser binding, meaningful only on wasm32 targets.

PyPI’s and npm’s rete-graph are likewise the library for their language, not the CLI, so re-exporting the library is the faithful mapping.

§Features

compression (default), parallel and wasm-js are forwarded verbatim to rete-core, so this crate can be configured exactly like it.

§Example

use rete_graph::Rete;

let bytes = std::fs::read("graph.rete")?;
let graph = Rete::open(&bytes)?;
println!("{} quads", graph.dump(None).len());

Modules§

format
Stable file-format, reader, and in-memory build API.
geo3
A 3D extension of the GeoSPARQL built-ins (geo3).
ingest
Ingestion: parse RDF text (N-Triples / N-Quads / Turtle) and assemble a complete .rete file image. Shared by the CLI’s build/validate commands and the wasm bindings (the playground’s in-browser builder).
query
Stable SPARQL, graph-pattern, federation, and result API.
range
Stable byte-range and lazy-open API for local or remote .rete files.
reasoning
Stable RDFS/OWL reasoning and schema-coherence API.
validation
Stable integrity and SHACL validation API.

Structs§

BlockCacheReader
ByteRange
A byte range in the .rete file image.
CharSet
A characteristic set — one entity “shape”: the sorted set of predicates a subject carries, plus how many subjects share exactly that shape. Computed at build (top-N shapes by subject count), range-fetched with the pyramid. Useful for data profiling (the distinct entity shapes) and for estimating the selectivity of a subject star (subjects matching ALL of several predicates).
ClassNode
One node of the shipped subClassOf DAG: a class, all its direct parents, and its depth (0 = root). The hierarchy is non-exclusive — a class may have several parents (multiple inheritance), so this is a directed acyclic graph, not a tree. The first parent (parents are sorted) is the canonical one used for the depth/rollup spanning tree; the rest preserve the cross-links.
ClassRelation
A class-to-class relation (the lateral, non-is-a connection): subjects of s_class related by predicate to objects of o_class, with the instance count. (literal) / (untyped) are the object-class sentinels.
CommunityDescriptor
A per-community refinement descriptor (Phase 4): what a client sees when it zooms into one community, without fetching that community’s triples. Carries the dominant class, the local type histogram, and optional spatial/temporal extents. (The physical per-community triple tiles are still future work; this descriptor index ships in the index-free pyramid-meta and is ready to attach to those tiles when they exist.)
CommunityPartial
One community’s contribution to a community-split evaluation: how many member subjects it holds and how many solution rows it produced.
CountingReader
Wraps a reader and tallies how many ranges were requested and how many bytes were returned — the metric that matters for a range-streamed format. Atomically counted, so it stays Sync (a lazily-faulting remote index holds its reader behind a shared loader).
DataGraph
A validation data graph held fully in memory as a sorted triple vector. Backs the shapes graph and the eager data path; the lazy data path uses ReteGraph.
Dendrogram
A hierarchy of community partitions (a dendrogram). levels[0] partitions the base nodes; levels[k] partitions level k-1’s communities. Round 0 is the finest grouping; the last round is the coarsest (pyramid level 0).
DictSection
A parsed, read-only dictionary section (parses its metadata on construction).
DictSectionBuilder
Build a dictionary section from terms (any order; sorted + deduped here).
Dictionary
A read-only dictionary mapping terms ↔ role-specific IDs.
DictionaryBuilder
Builds a Dictionary from observed (subject, predicate, object) terms.
Graph
A weighted undirected graph over nodes 0..n. Parallel edges are summed; self-loops are kept (they affect modularity but not community moves).
GraphIndex
The six tiled permutation sections, queryable by triple pattern.
GraphIndexBuilder
Build a GraphIndex from canonical (s, p, o) integer triples.
GroupDirectory
A byte-offset directory of a block’s a-groups: one entry per group, sorted by leading id (the storage order). Built once per block with TripleBlock::group_directory; TripleBlock::scan_from then binary-searches it to jump a probe straight to its group.
GroupSpec
GROUP BY specification: grouping variables and result aggregates.
Header
Decoded file header. All multi-byte fields are little-endian on disk. The *_offset / *_len fields are a convenience view over the section directory (populated from it on parse, emitted back to it on serialize).
Inconsistency
One detected incoherent point (a logical contradiction in the graph).
LabelEntry
One entry of the label index: the display label of a subject (its rdfs:label / skos:prefLabel / … literal) paired with that subject’s id. Computed at build (the most-connected labeled subjects, bounded), the table is sorted by the label’s lowercased form so a prefix query is a binary search — autocomplete and CONTAINS/prefix lookup without scanning the literals. Range-fetched with the pyramid; the label string is stored inline so the search is self-contained on the remote-lazy path (no dictionary fault).
LayoutSegment
One labelled byte region of a .rete file image (see Rete::file_layout). kind is a stable machine tag: header, metadata, dictionary, directory, tile, pyramid, named-graphs.
LevelLinks
The class-relation graph rolled up to depth — the lateral connections at one semantic-zoom level. Coarse levels show relations between abstract classes (Agent → Agent); finer levels resolve them (Person → Organisation). This is what makes the pyramid a leveled graph, not just a leveled histogram.
LevelRollup
A type histogram rolled up to depth — the global class distribution at one semantic-zoom level. Coarse levels (small depth) hold abstract ancestor classes; fine levels (large depth) resolve to leaves. round is the dendrogram round this level is aligned with (informational).
Partition
A community partition with dense IDs 0..count.
PredStat
Per-predicate cardinality statistics for the cost-based query planner — computed at build, range-fetched with the pyramid (index-free). count is the triple total; the distinct-subject/object counts and max multiplicities let the planner estimate a bound subject’s/object’s selectivity (functional iff max_objects_per_subject == 1, inverse-functional iff max_subjects_per_object == 1).
PyramidMeta
Decoded pyramid metadata.
Reasoning
The result of reasoning over a base graph.
Rete
A read-only, in-memory view over a .rete file image.
ReteGraph
A SHACL data-graph view backed directly by a .rete file’s index: every lookup is a routed pattern query, so over a lazy (Rete::open_ranged_lazy) open a validation faults only the tiles holding the shapes’ target nodes — not the whole graph. Views the default graph (named-graph validation uses the eager DataGraph).
RoutedTriplePattern
A single triple pattern that can be answered by the range-routed permutation reader. None means the position is a variable/wildcard; Some(term) means the query pins that term.
Section
One parsed/encoded section-directory entry: a typed (offset, length) into the file, plus 16 bits of per-section flags (reserved).
Select
A lowered SELECT query: solution modifiers plus the evaluation plan tree.
ShaclShapes
Parsed SHACL shapes graph.
SliceReader
A RangeReader over an in-memory byte slice (tests, embedded files).
SummaryView
A lightweight, overview-only view of a file: the pyramid summary graph plus just enough dictionary to label predicates. Fetched via ranges without touching the (large) triple index — the “load the coarse graph first” path from SPEC.md §7.2.
SuperEdge
An aggregated relation between two communities in the summary graph.
TextIndex
A parsed text index: the token table (always resident) plus a posting source.
TextIndexBuilder
Accumulates token → subjects and serializes the section.
Tile
One tile: a community’s triples and their encoded SPO block.
TripleBlock
A parsed triple block.
TripleBlockBuilder
Accumulates triples and serializes a block.
TriplePattern
A triple pattern (subject, predicate, object) of PatternTerms.
TripleProvenance
Why a triple-pattern result is present in the file.
ValidationReport
ValidationResult
ZoneMap
Per-block summary statistics enabling block-skipping.

Enums§

Agg
A supported aggregate function.
FExpr
A small boolean/comparison expression for FILTER (a subset of SPARQL exprs).
GraphTarget
The target of a GRAPH block.
HeaderError
IndexPermutation
Which stored permutation to scan. The full six orders (a SPARQL engine’s classic set, as in QLever): together they sort the triples on every prefix of (s, p, o) columns, so for any bound-component prefix and any free component there is a permutation that routes on the prefix and yields the stream sorted on that free component — the precondition for a merge join.
Op
Comparison operators supported in FILTER.
PathAst
A lowered property path expression, evaluated as a binary relation over the graph’s nodes.
PatternTerm
A term in a pattern: a named variable or a constant term token.
Plan
A SPARQL graph-pattern evaluation plan (the supported algebra subset).
PyramidAlgo
Which algorithm forms the pyramid’s communities. The format and every downstream consumer key on opaque community IDs, so the choice only changes how the partition is computed — all modes must be deterministic to keep the file content-hash reproducible (see docs/BENCHMARK.md).
QueryOutput
The result of evaluating any SPARQL query form.
Rep
Path repetition operator.
SectionKind
A top-level file section, addressed by SectionKind in the header directory.
Severity
ShaclError
SparqlError
SummaryQueryShape
Query shapes that can be answered exactly from crate::range::SummaryView predicate totals, without opening the triple index.

Constants§

CODEC_NONE
No compression.
CODEC_ZSTD
zstd compression (per section).
CURRENT_FORMAT_VERSION
Current format generation written by this crate.
DEFAULT_BLOCK
Default block size: 64 KiB — large enough to swallow a dictionary chunk or an index tile in one fetch, small enough to keep over-fetch modest.
DEFAULT_CACHE_CAP
Default cap on resident cached bytes: 256 MiB. Large enough that a working set (the tiles + dictionary chunks a query family touches) stays warm; small enough that a full sweep of a multi-GB file leaves plenty of the 32-bit wasm address space for the decompressed structures built on top of these bytes. At the auto-tuned 128–512 KiB block sizes this is 512–2048 resident blocks, so the eviction scan is trivial.
DEFAULT_TILE_BUDGET
Default per-tile byte budget T (SPEC.md §7.1).
ENGINE_VERSION
The rete-core version this facade re-exports.
HEADER_LEN
Fixed header size in bytes.
MAGIC
Magic bytes at offset 0: ASCII RETE.
MIN_STABLE_READ_VERSION
Oldest stable format generation accepted by this reader.
RDF_TYPE
rdf:type — the predicate that assigns a class to a resource.
REASON_RULESET
Version tag of this reasoner’s rule set. Stamped into a baked coherence card so a coherent: true can never be misread as a guarantee from a different set of rules. Bump this whenever materialize/detect_inconsistencies changes (a rule added/removed/altered), so rete reason --verify-card rejects a stale stamp.
VERSION
The engine version, as published on crates.io.

Traits§

GraphView
A read-only view of the data graph for SHACL validation. The validator only ever asks targeted questions — a focus node’s values, the subjects of a predicate, the instances of a class — so this surface is small enough to back two ways: an in-memory triple set (DataGraph, eager) or a .rete file’s index directly (ReteGraph), which routes each lookup as a range read so a remote validation faults only the tiles holding the shapes’ targets, not the whole graph. The class/instance helpers are derived from the primitives, so a backend only implements the six lookups.
RangeReader
Something that can serve arbitrary byte ranges of a .rete resource.
ServiceClient
Executes one SPARQL query against a remote endpoint (the SPARQL Protocol) and returns its solutions. Implementations own transport, auth, and timeouts; they typically POST the query with Accept: application/sparql-results+json and feed the body through parse_sparql_json_results. Errors are strings, surfaced verbatim as the query error (or, under SERVICE SILENT, swallowed per the spec) — name the endpoint in the message, the engine adds no prefix.

Functions§

auto_block
Pick a BlockCacheReader block size from the file length: bigger files get bigger blocks so a remote query makes far fewer (but larger) round trips — 128 KiB ≤ 10 MB, 256 KiB ≤ 100 MB, 512 KiB above. The over-fetch is modest next to the round-trip latency it saves on a high-latency link (S3/CDN). Shared by the CLI and the wasm client so both size identically; the file length is known for free from the opening HEAD / Content-Range.
batch_reach_serial
Per-seed transitive reach, serial loop (reference). Results are returned in seed order. The parallel sibling crate::parallel::batch_reach_parallel produces an identical result.
build_adjacency
Forward adjacency in unified node space for one predicate: node -> [succ]. Built once and shared (read-only) across all seeds. For reverse reachability (“who reaches the seed?”), build the map yourself from Rete::predicate_pairs swapping (s, o) -> (o, s).
build_dendrogram
Build the full dendrogram by repeated Louvain + coarsening, stopping when a round no longer compresses the graph (or only one node remains).
build_pyramid_meta
Build the encoded pyramid-meta section for a graph: cluster, pick a round sized to budget, then emit the summary (quotient) graph. Returns (encoded_meta, pyramid_levels).
build_pyramid_meta_algo
Like build_pyramid_meta_with, but selects the community PyramidAlgo. PyramidAlgo::Types partitions by rdf:type — the deterministic, parallelizable alternative to Louvain (one linear pass, no modularity) that still emits the full summary + query_stats; it falls back to Louvain when the graph has no usable typing. Everything downstream of the dendrogram (round choice, summary, schema pyramid, planner stats) is shared across algorithms.
build_pyramid_meta_with
Like build_pyramid_meta, but type_override forces the schema-pyramid’s type predicate (e.g. wdt:P31) instead of auto-detection. Uses the default PyramidAlgo::Louvain community algorithm — byte-identical to before.
build_schema_pyramid
Compute the schema pyramid for a graph at the materialized round, auto-picking the type predicate. Empty when the graph has no usable typing.
choose_round_for_budget
The coarsest round whose every tile fits budget_bytes, else round 0. Coarser rounds mean fewer, larger tiles; we want the fewest tiles that still respect the per-tile budget (PMTiles-style).
eval_bgp
Evaluate a BGP against the file’s default graph, returning all solutions resolved to terms (the public convenience API).
eval_query
Evaluate any supported SPARQL query form (SELECT / ASK / CONSTRUCT).
eval_query_reasoned
Like eval_query, but with OWL 2 QL entailment on: the lowered plan is rewritten by the internal QL lowering pass so the answer includes ontology-entailed solutions (Stage 1a: rdfs:subClassOf), computed over the raw data with no materialization. Opt-in — a plain eval_query is byte-identical to before.
eval_select_communities
Evaluate a SELECT per pyramid community, then merge: each community’s subjects are pushed into the plan as a VALUES binding, the partial rows are concatenated, and the solution modifiers (GROUP BY / ORDER BY / LIMIT / DISTINCT) run once on the union — so the rows are identical to eval_query’s answer. Sound only for subject-star queries over the default graph (every triple pattern sharing one subject variable; FILTERs allowed); anything else returns SparqlError::Unsupported rather than a possibly-wrong split answer. round picks the dendrogram granularity (None = the build’s tile-budget round). Also returns each community’s subject and row counts for display.
eval_sparql
Parse and evaluate a SELECT against a file, applying the plan then projection, DISTINCT, OFFSET, and LIMIT. Returns (projected_vars, solutions).
eval_sparql_reasoned
Like eval_sparql, but with OWL 2 QL entailment on (see eval_query_reasoned).
louvain_one_level
One level of Louvain local-moving modularity optimization.
parse_select
Parse a SPARQL SELECT query and lower it to a Select.
parse_sparql_json_results
Parse a SPARQL 1.1 Query Results JSON document (application/sparql-results+json) into bindings of variable → N-Triples term token. An ASK document (no results) yields no bindings. Unknown per-binding fields are ignored; xsd:string datatypes are dropped (a simple literal — matching how plain literals are tokenized everywhere else in the engine).
project_graph
Project RDF integer triples onto the undirected node graph the pyramid clusters: one node per distinct term (via the dictionary’s unified node space), one unit-weight edge per (subject, object) pair (predicates are ignored for clustering; parallel edges accumulate weight).
push_json_string
Append v to out as a JSON string literal (RFC 8259 escaping): the mandatory escapes (", \, and C0 controls, the common ones in short form), every other char — including all UTF-8 — passed through. Matches what serde_json emits for a string by default.
query_predicates
Collect the concrete predicate IRIs a query constrains on — i.e. every IRI that appears in the predicate position of a triple pattern, or as a plain predicate inside a property path. Variable predicates (?p) and the special a (rdf:type) keyword are normalized to their IRI tokens (<…>).
reach_one
Transitive reach of seed over the adjacency (excludes the seed itself). Plain BFS; deterministic BTreeSet result. This is the single shared BFS used by both the serial and parallel batch drivers.
read_metadata_ranged
Fetch only the metadata section (the opaque Dataset Card blob) via a RangeReader: read the 128-byte header, then the metadata byte range — nothing else. This is the index-free CARD tier of the exploration model: a remote/S3 client learns the dataset’s self-description in two small range requests, never touching the dictionary, index, or pyramid. Returns None when the file carries no metadata.
read_schema_coherence_ranged
Dictionary-free Tier-0 coherence read. Fetch only the header and the pyramid-meta range (2 small range reads) and run schema_coherence over the schema pyramid — never touching the dictionary (which a literal-heavy file makes large) or the triple index. Ok(None) if the file ships no pyramid.
read_schema_summary_ranged
The schema summary (per-class histogram + class relations at the finest level) read over a RangeReader from the schema pyramid alone — the index-free, range-readable source for a Schema view of a remote graph. Returns (classes, relations) with classes = [(class_iri, count)] and relations = [(s_class, predicate, o_class, count)]; None when the file has no schema pyramid. Like read_schema_coherence_ranged, it reads only the trailing schema block, so it stays flat at any graph size.
reason
Forward-chain the supported RDFS/OWL rules to a fixpoint, then scan for inconsistencies over the closed graph. inferred excludes triples already present in base_triples.
results_envelope_json
Serialize out as { "kind": …, … }. extra is a raw JSON fragment of additional object members appended before the closing brace (e.g. ,"remote":{…}); pass "" for none. CONSTRUCT is rendered as a triples array — the text formats (Turtle / JSON-LD) are handled by the caller, which owns those serializers.
routed_triple_pattern
Classify queries whose graph access is exactly one default-graph triple pattern. Solution modifiers (projection, LIMIT, aggregate wrappers) do not change the underlying range access, but named graphs, FROM, joins, filters, paths, and other algebra need the full SPARQL evaluator.
schema_classes
Class populations: the number of resources of each rdf:type class in the default graph, descending by count. The instance-count companion to schema_summary.
schema_coherence
Compute T-Box coherence points from the schema-pyramid fields alone — no dictionary, no index, no instance data. Shared by SummaryView::tbox_coherence and the dictionary-free read_schema_coherence_ranged. Emits subclass-cycle and unsatisfiable-class (a class whose ancestor closure — over all parents, folded through owl:equivalentClass — contains both ends of a disjoint pair).
schema_summary
An ontology-aware coarse graph: instead of structural communities, group entities by their rdf:type class and aggregate relations between classes. Returns (subject_class, predicate, object_class, count) over the default graph. Entities with no type are (untyped); literals are (literal). rdf:type triples themselves define the classes and are not counted as relations. This is the dataset’s effective schema with instance volumes.
sparql_json_ask
Serialize an ASK result as a SPARQL 1.1 Query Results JSON document.
sparql_json_results
Serialize a SELECT result as a SPARQL 1.1 Query Results JSON document. vars fixes the head order (pass the projection); solutions’ bindings are term tokens exactly as the engine returns them.
summarize
Aggregate triples into the quotient (summary) graph at round.
summary_query_shape
Classify SPARQL queries that can be answered exactly from the pyramid summary’s per-predicate totals. This is intentionally conservative: anything with constants, repeated variables, filters, joins, paths, named graphs, ORDER BY, OFFSET/LIMIT, or non-summary-safe aggregates still requires the index. The only accepted DISTINCT shape is a predicate list over one fully unbound triple pattern.
tile_by_community
Partition triples into per-community tiles at round, ordered by community.
tokenize
Split text into index/query tokens: Unicode-alphanumeric runs, lowercased, length ≥ MIN_TOKEN_LEN. The build and query sides MUST use this same function so a query word matches how it was indexed.
validate_shacl
Validate data against shapes. data is any GraphView — an in-memory DataGraph (eager) or a ReteGraph that routes lookups as range reads (lazy / remote, fetching only the shapes’ targets).
verify
Recompute the content hash from a file image and check it against the header — detects corruption or truncation of the payload sections.
write_dataset
Serialize a full RDF dataset: the default-graph index plus zero or more named graphs (iri, index), all sharing one dictionary.
write_dataset_with_metadata
Serialize a dataset with an opaque metadata payload occupying the file’s metadata section (the application layer defines its meaning — the CLI stores a JSON Dataset Card there). The section sits immediately after the header and before the dictionary, so metadata_offset stays at HEADER_LEN and every downstream section shifts by metadata.len(). The payload is folded into the content_hash, so verify covers it and it is tamper-evident.
write_file
Serialize a complete .rete file image from a dictionary, index, and an (optionally empty) encoded pyramid-meta section. pyramid_levels records the number of dendrogram rounds the pyramid spans (0 if no pyramid).

Type Aliases§

Binding
A solution: variable name → bound term (the public, resolved form).
CommunitySelect
The outcome of a community-split SELECT: the projected variables, the merged solution rows, and each community’s contribution.
NodeId
A dictionary id in the unified node space — the single id space that covers every term that ever appears as a subject or an object. This is the id reachability, the community pyramid, and the graph index work in.
ObjectId
A dictionary id in the object id space (terms seen in object position). Map to a NodeId with Dictionary::object_node.
Pattern
A triple pattern: None is an unbound variable, Some(id) a bound term.
PredicateId
A dictionary id in the predicate id space. Predicates have their own dense id space and are never part of the unified node space.
SubjectId
A dictionary id in the subject id space (terms seen in subject position). Map to a NodeId with Dictionary::subject_node.
TermToken
The textual form of an RDF term as stored in the dictionary and emitted in N-Triples: an IRI (<…>), a blank node (_:…), or a literal ("…", optionally with an @lang or ^^<datatype> suffix). An alias for str; it names intent at API boundaries that take a term rather than arbitrary text.
TermTriple
A resolved triple as terms.
Triple
A triple of dictionary IDs in some permutation’s component order.