Expand description
DataFusion TableProvider and CatalogProvider implementations.
Each GraphForge graph directory (project/) maps to a GraphCatalog which
presents its Parquet files as DataFusion tables under the address
graph.graph.<table_name>:
| Table name | File | Schema |
|---|---|---|
topology_nodes | topology/nodes.parquet | TOPOLOGY_NODES_SCHEMA |
edges_TYPENAME | topology/edges/TYPENAME.parquet | TYPED_EDGE_SCHEMA |
edges__exploratory | topology/edges/_exploratory.parquet | EXPLORATORY_EDGE_SCHEMA |
properties_ENTITY | properties/ENTITY.parquet | property_schema(entity, defs) |
§Scan implementation
Scans are implemented via DataFusion’s MemTable: the Parquet file is read
into memory at query time and wrapped in a MemTable which handles projection
and filter application. This is correct and simple for M12; lower-level
pushdown can be added in a later milestone.
Structs§
- Edge
Property Table TableProviderforedge_properties/REL_TYPE.parquet(#784).- Graph
Catalog - DataFusion
CatalogProviderfor a GraphForge project directory. - Property
Table TableProviderforproperties/ENTITY_TYPE.parquet.- Topology
Node Table TableProviderfortopology/nodes.parquet.- Typed
Edge Table TableProviderfortopology/edges/TYPENAME.parquet.- Union
Edge Table TableProviderover the union of every relation’s edge file (#823) — the scan source for an untyped single-hop pattern ((a)-[]->(b)) in a typed project, where the_exploratorytable does not exist. Materializes [read_edges_union] into aMemTable; the schema is alwaysEXPLORATORY_EDGE_SCHEMA(each row tagged with its source relation), so it is a drop-in for the exploratory edge scan the untyped lowering already uses.
Functions§
- list_
edge_ property_ stems - Stems (relation names) of every
edge_properties/<stem>.parquetunderdir, sorted so schema unions built from them are deterministic (#1023). Empty when the directory is absent — a project with no persisted edge properties. - list_
property_ stems - Stems (entity type names, or
_untyped) of everyproperties/<stem>.parquetunderdir, sorted for deterministic schema unions (#1024). Empty when the directory is absent. - read_
edge_ properties - Edge analogue of
read_properties: readedge_properties/<stem>.parquet(keyed byedge_uuid), discovering its dynamic schema from the file. Returns an emptyVecwhen the file is absent. - read_
edges - Read all edge rows for relation
rel_namefrom the project atdir. - read_
edges_ filtered - Like
read_edgesbut returns only rows whoseedge_idis inedge_ids— the traversal’s lazy edge-record read (#830): on an adjacency Hit, only the traversed edges’ records are needed, not the whole file. - read_
nodes - Read all node rows from
topology/nodes.parquetin the project atdir. - read_
nodes_ filtered - Like
read_nodesbut returns only rows whosenode_idis innode_ids— the traversal’s lazy node-record read (#838): on an adjacency Hit only the reached destination nodes’ records are needed to project the destination columns, not the whole node table. Canonical dense files use exact physical row selection; legacy, gapped, or noncanonical files retain conservative row-group pruning plus a membership predicate. - read_
properties - Read
properties/<stem>.parquetfor the project atdir, discovering its (dynamic) schema from the file. Returns an emptyVecwhen the file is absent — so a caller decoding rows sees zero pre-existing property rows.