Which time windows a study expects rows in, as stated in Setup: every window of
the grain, or Monday to Friday’s only, from one time and before another. Unset,
no window is called a gap: a quiet weekend is not a defect unless someone says so.
Reads named columns of named files at the type each file wrote, which is the only
way back to the values a type conflict hides. Given to a run that already reads
every value, so the extra read is one column of the few files that disagree.
A run’s line to the screen: its stages as it enters them, the rows its reads
have seen, and a stop the run checks between stages and its reads check between
batches.
Columns that are null the same number of times, and how many rows are null in all
of them at once. When the two counts agree, the columns go missing together: one
fact about some rows, not one per column.
A text column read as a date or datetime for one study. Grain and time roles see
the parsed value; every other check sees the text as stored, so a column’s own
findings keep their physical meaning. A value the format does not read is counted
as unparsed, never folded into the column’s missing values.
How a full scan of a remote source gets its rows, as Setup says before Run: one
fetch into a local copy that every pass reads, a copy fetched earlier, or a pass
over the source for each check.
Which time puts an interval in a window, when the grain is time windows: the
grain’s own column, or the interval’s start or end. By the end, an interval is
counted on the day it finished rather than the day it began.
The intervals a run measures when none are chosen, start role to end role: the
pairs whose order the roles themselves state. Any other start and end is a
choice under Intervals in Setup; a role in no interval measures nothing, and
Setup says so before a run.
The share of non-null text values that must parse before a text column is said
to hold numbers or dates. Below it the column is text that happens to contain a
few numbers, which is not a finding.
The formats Setup offers for reading text as time, the unambiguous ones first.
Named formats rather than inference: a run reads every row the same way, and a
value the format does not read is counted, not guessed at.
Read every sample of audio once and add what a recording’s quality turns on to
results: clipping, runs of exact zeros, and DC offset, per channel. Only for a
full run whose scope is every frame of the file, which the caller decides: the
read is of the file, not of the view.
compute_data_quality, cutting kept instead of reading when it serves the
plan, and returning the sample a sampled run read so the next run can do the same.
The rows the duplicate check counted: every row equal to another in every one of
keys, copies together, most copies first. One pass, grouping as the check did.
Where a run of plan that reads a new sample gets its segment totals.
may_read_blocks is whether the sample may be seeded runs of one file, which see
too few rows to count; the head sees too few as well. Every other sample is one
streamed pass over the scope, which counts the grain’s key as it goes.
Whether the pass that samples plan’s rows also counts its segments: an
equal-per-value sample by the column the grain splits by counts every value as it
streams.
Every column’s measures in segment index: beside the segment it is compared
with and largest move first, or on its own worst first. A measure that is zero
on both sides says nothing and is left out.
Whether a sampled run of plan counts its segments’ rows: partitions and time
windows are counted for exact totals; files and row chunks are known without it.
Whether windows of width fine nest exactly in windows of coarse: every hour in
one day, every day in one week (weeks start on Monday) and one month. Windows are
cut on the stored clock with no time zone (UTC for a zoned column; see
[time_window_start]), where no day has 23 or 25 hours, so a sum of the finer
counts is the coarser count. A week does not nest in a month.