The rows a sampled run read, kept beside its results. Keyed by acquisition (dataset,
view, scope, method, size, seed); other plan settings are the report’s, so a run
changing only those re-cuts these rows instead of reading the source. Every column
and each row’s position are kept.
A text column read as a date or datetime for one study: grain and time roles see
the parsed value, other checks the stored text. Unreadable values count as
unparsed, never as missing.
Which time puts an interval in a window under a time-window grain: the grain’s
column, or the interval’s start or end (by end, counted on the day it finished).
The intervals measured when none are chosen: start role to end role, the pairs
whose order the roles state. Others are chosen in Setup; a role in no interval
measures nothing, as Setup says.
How nearly unique a column must be before its repeats are worth naming: an id
repeating twice in a million rows is a finding, a category repeating is not.
The formats Setup offers for reading text as time, unambiguous first. Named, not
inferred: every row reads the same way, and an unreadable value is counted.
Read every sample of audio once and add clipping, zero runs and DC offset per
channel to results; only for a full run over every frame (the caller decides).
The data quality of lf by plan, cutting kept instead of reading when it
serves the plan. Names each stage to watch and stops between stages (or inside a
streamed read) once cancelled. A sampled run’s rows come back even if it stopped:
the read is paid for, and the next run can cut it.
The rows the duplicate check counted: rows equal to another in every keys column,
copies together, most first, in one pass. Each group’s key is its rows, repeated
by count rather than looked up again.
Where a run reading a new sample gets segment totals. Seeded runs of one file
(may_read_blocks) and the head see too few rows; any other sample is one streamed
pass that counts the grain’s key as it goes.
Whether a sampled run of plan counts its segments’ rows: partitions and time
windows are counted for exact totals; files and row chunks are known without it.
Observations from crate::formats::audio::SignalReports: a channel with runs at full
scale, runs of exact zeros, or a mean 1% of full scale or more from zero.
Whether windows of width fine nest exactly in coarse: hours in a day, days in a
week (Monday start) and in a month; weeks not in months. Windows are cut on the
zoneless stored clock (UTC for zoned; see time_window_start), so no day has
23 or 25 hours and finer counts sum to the coarser.