Expand description
Aggregate statistics over every recorded run.
These tables are a by-product of running the graph, not a benchmark. The seat assignment rotates, the task distribution is whatever the operator happened to ask for, and a model that draws harder tasks looks worse. Read them as “relative performance on my workload”, which is the only claim the data supports.
Structs§
- Agent
Stats - Implementation record for one agent.
- E2eStats
- What real-machine verification caught that static review did not.
- Node
Duration - Per-node duration breakdown, aggregated across every loaded run.
- Reviewer
Stats - Review record for one agent.
- Stats
- Everything, aggregated.
- Totals
- Run-level counters.
Functions§
- collect
- Aggregate
states. - load_
all - Load every run on disk, skipping any that cannot be read.
- node_
durations - Derive a per-node duration breakdown from every loaded run’s events.