Skip to main content

Module catalog

Module catalog 

Source
Expand description

DataFusion TableProvider and CatalogProvider implementations.

Each GraphForge graph directory (project/) maps to a GraphCatalog which presents its Parquet files as DataFusion tables under the address graph.graph.<table_name>:

Table nameFileSchema
topology_nodestopology/nodes.parquetTOPOLOGY_NODES_SCHEMA
edges_TYPENAMEtopology/edges/TYPENAME.parquetTYPED_EDGE_SCHEMA
edges__exploratorytopology/edges/_exploratory.parquetEXPLORATORY_EDGE_SCHEMA
properties_ENTITYproperties/ENTITY.parquetproperty_schema(entity, defs)

§Scan implementation

Scans are implemented via DataFusion’s MemTable: the Parquet file is read into memory at query time and wrapped in a MemTable which handles projection and filter application. This is correct and simple for M12; lower-level pushdown can be added in a later milestone.

Structs§

EdgePropertyTable
TableProvider for edge_properties/REL_TYPE.parquet (#784).
GraphCatalog
DataFusion CatalogProvider for a GraphForge project directory.
PropertyTable
TableProvider for properties/ENTITY_TYPE.parquet.
TopologyNodeTable
TableProvider for topology/nodes.parquet.
TypedEdgeTable
TableProvider for topology/edges/TYPENAME.parquet.
UnionEdgeTable
TableProvider over the union of every relation’s edge file (#823) — the scan source for an untyped single-hop pattern ((a)-[]->(b)) in a typed project, where the _exploratory table does not exist. Materializes [read_edges_union] into a MemTable; the schema is always EXPLORATORY_EDGE_SCHEMA (each row tagged with its source relation), so it is a drop-in for the exploratory edge scan the untyped lowering already uses.

Functions§

list_edge_property_stems
Stems (relation names) of every edge_properties/<stem>.parquet under dir, sorted so schema unions built from them are deterministic (#1023). Empty when the directory is absent — a project with no persisted edge properties.
list_property_stems
Stems (entity type names, or _untyped) of every properties/<stem>.parquet under dir, sorted for deterministic schema unions (#1024). Empty when the directory is absent.
read_edge_properties
Edge analogue of read_properties: read edge_properties/<stem>.parquet (keyed by edge_uuid), discovering its dynamic schema from the file. Returns an empty Vec when the file is absent.
read_edges
Read all edge rows for relation rel_name from the project at dir.
read_edges_filtered
Like read_edges but returns only rows whose edge_id is in edge_ids — the traversal’s lazy edge-record read (#830): on an adjacency Hit, only the traversed edges’ records are needed, not the whole file.
read_nodes
Read all node rows from topology/nodes.parquet in the project at dir.
read_nodes_filtered
Like read_nodes but returns only rows whose node_id is in node_ids — the traversal’s lazy node-record read (#838): on an adjacency Hit only the reached destination nodes’ records are needed to project the destination columns, not the whole node table. Canonical dense files use exact physical row selection; legacy, gapped, or noncanonical files retain conservative row-group pruning plus a membership predicate.
read_properties
Read properties/<stem>.parquet for the project at dir, discovering its (dynamic) schema from the file. Returns an empty Vec when the file is absent — so a caller decoding rows sees zero pre-existing property rows.