Expand description
Process-global I/O counters for storage reads and staged rewrite commits.
The read counters cover read_edges,
read_edges_filtered, and
read_nodes.
§Why
These prove the M15 T1 criterion (#767): with the adjacency index present,
variable-length traversal must not scan the full edge file — it issues only
an edge_id-filtered read whose materialized row count is proportional to
the traversed neighborhood, independent of the total edge count. A scan
over the index path can then assert edge_full_reads == 0, while the
scan-build baseline shows edge_full_rows >= total_edges.
§Semantics
- A full read is one decode of a whole file:
read_edges/read_nodes, plus theread_edges_filteredfallback (a requested id set covering more than half the file reads it whole, then trims in memory — it is a full scan and is counted as one, so it cannot hide behind the filtered API). - A filtered read is the predicate-pushdown path of
read_edges_filtered; its row count is the rows actually materialized after row-group and row-filter pruning. rowscount rows actually returned to the caller. A missing or empty file still counts as one read of zero rows (the reader was invoked).- A rewrite commit is one successful, non-empty
RewriteBatchcommit, regardless of how many files were staged in that batch.
§Caveats
Counters are process-global and aggregate across threads and queries
(read_nodes is also used by the writer/mutator). They are advisory
instrumentation, not per-query state. A test that asserts on them must
reset immediately before the measured operation and keep that section
single-threaded — the counters cannot attribute concurrent work. Each
increment is one relaxed atomic add, negligible against a Parquet decode.
Structs§
- IoSnapshot
- A point-in-time copy of the process-global I/O counters. Difference two
snapshots — or
resetthensnapshot— to attribute work to a region.