Skip to main content

Module data_quality

Module data_quality 

Source

Structs§

CategoryVariantGroup
ColumnQualityProfile
CopyRead
A local copy of a remote source that a full scan’s passes read.
DataQualityPlan
DataQualityResults
DuplicateExample
One group of identical rows: how many there are, and the row, a value a column.
ExpectedWindows
Which time windows a study expects rows in, as stated in Setup: every window of the grain, or Monday to Friday’s only, from one time and before another. Unset, no window is called a gap: a quiet weekend is not a defect unless someone says so.
FindingExamples
A few of the values behind one column’s finding, from the rows the run kept.
IdentityProfile
ObservedReads
What a run’s reads of the source were seen to do, counted as the rows went by. Bytes and requests are not counted: a Polars scan does not report them.
QualityConflictScan
Reads named columns of named files at the type each file wrote, which is the only way back to the values a type conflict hides. Given to a run that already reads every value, so the extra read is one column of the few files that disagree.
QualityFileEvidence
One file behind a drift observation: what it holds, and what that costs the column.
QualityObservation
QualityPhase
A stage, whether it reads the source or works on rows already read, and whether a cancel stops it partway.
QualitySample
The rows a sampled run read, kept beside its results: an acquisition.
QualitySourceContext
QualityWatch
A run’s line to the screen: its stages as it enters them, the rows its reads have seen, and a stop the run checks between stages and its reads check between batches.
SegmentChange
One column’s measure in a segment, and in the segment it is compared with.
SegmentQualityProfile
SharedNulls
Columns that are null the same number of times, and how many rows are null in all of them at once. When the two counts agree, the columns go missing together: one fact about some rows, not one per column.
TemporalLatencyProfile
TemporalRoleAssignment
TimeInterpretation
A text column read as a date or datetime for one study. Grain and time roles see the parsed value; every other check sees the text as stored, so a column’s own findings keep their physical meaning. A value the format does not read is counted as unparsed, never folded into the column’s missing values.
UnsampledSegment
A segment the scope has rows in and a sample drew none of.

Enums§

CopyPlan
How a full scan of a remote source gets its rows, as Setup says before Run: one fetch into a local copy that every pass reads, a copy fetched earlier, or a pass over the source for each check.
IntervalClock
Which time puts an interval in a window, when the grain is time windows: the grain’s own column, or the interval’s start or end. By the end, an interval is counted on the day it finished rather than the day it began.
IntervalFact
What an interval’s detail counts, each with the rows behind it.
NoCopy
Why a remote full scan reads the source in each pass instead of a local copy.
ObservationKind
QualityComparison
QualityCompute
QualityGrain
QualityMetric
QualityPage
QualityPrecision
QualityScope
QualitySetup
What an empty page is missing, which Enter opens in Setup.
QualityStage
What a Data Quality run is doing now. The worker names each stage as it enters it, and the progress view shows the latest.
SegmentCount
Where a run’s exact segment totals come from, as Setup says before Run.
TemporalRole
TextReading
What the values of a text column parse as, most specific first.
TimeKind
Whether text read as time is a date or a date with a time of day.

Constants§

INTERVAL_PAIRS
The intervals a run measures when none are chosen, start role to end role: the pairs whose order the roles themselves state. Any other start and end is a choice under Intervals in Setup; a role in no interval measures nothing, and Setup says so before a run.
KEY_LIKE_UNIQUENESS
How nearly unique a column’s values must be before its repeats are worth naming.
MAX_FINDING_EXAMPLES
Values kept per finding from the rows a run read, and groups of duplicate rows: enough to recognize the problem in the detail, which opens the rest.
QUALITY_SOURCE_FILE_COLUMN
QUALITY_WINDOW_WIDTHS
Window widths offered for time-window grain, in the order the plan cycles them.
TEXT_READING_SHARE
The share of non-null text values that must parse before a text column is said to hold numbers or dates. Below it the column is text that happens to contain a few numbers, which is not a finding.
TIME_FORMATS
The formats Setup offers for reading text as time, the unambiguous ones first. Named formats rather than inference: a run reads every row the same way, and a value the format does not read is counted, not guessed at.

Functions§

add_signal_observations
Read every sample of audio once and add what a recording’s quality turns on to results: clipping, runs of exact zeros, and DC offset, per channel. Only for a full run whose scope is every frame of the file, which the caller decides: the read is of the file, not of the view.
apply_quality_scope
beyond_noise
Whether rates a of n_a rows and b of n_b rows differ by more than two samples of those sizes would by chance (a two-proportion z-test).
compute_data_quality
compute_data_quality_kept
compute_data_quality, cutting kept instead of reading when it serves the plan, and returning the sample a sampled run read so the next run can do the same.
compute_data_quality_watched
compute_data_quality_kept, naming each stage to watch as it enters it and stopping between stages, or inside a streamed read, once watch is cancelled.
duplicate_rows
The rows the duplicate check counted: every row equal to another in every one of keys, copies together, most copies first. One pass, grouping as the check did.
fresh_segment_count
Where a run of plan that reads a new sample gets its segment totals. may_read_blocks is whether the sample may be seeded runs of one file, which see too few rows to count; the head sees too few as well. Every other sample is one streamed pass over the scope, which counts the grain’s key as it goes.
interval_label
event to received.
interval_passes
How many groupings a run’s intervals take: one per distinct grain. A full run reads the scope once for each.
page_setup
The plan setting a result page needs before it has anything to show, if any. Intervals need time roles, and ask only when there are dates to assign.
prepare_source_quality_scan
Prepare the loaded source in the worker, keeping only a provenance index and replacing binary payloads before any value collection.
sampler_counts_segments
Whether the pass that samples plan’s rows also counts its segments: an equal-per-value sample by the column the grain splits by counts every value as it streams.
segment_changes
Every column’s measures in segment index: beside the segment it is compared with and largest move first, or on its own worst first. A measure that is zero on both sides says nothing and is left out.
segment_order
The order Segments lists its rows in: as they fall, or the clearest change first (ties, and segments with no clear change, keep their order).
segments_need_count
Whether a sampled run of plan counts its segments’ rows: partitions and time windows are counted for exact totals; files and row chunks are known without it.
shows_trend
Whether the Trends page can draw a column’s measure across segments: that needs segments in an order, and more than one of them.
signal_observations
Observations from crate::audio::SignalReports: a channel with runs at full scale, runs of exact zeros, or a mean 1% of full scale or more from zero.
text_reading
The one typed reading a text column’s values support, with how many parse.
unparsed_text
The rows of a parseable-text column its reading does not parse: non-null text that stops a cast. None when the column has no reading.
window_cadence
A window width as a cadence: 1d is daily.
window_nests
Whether windows of width fine nest exactly in windows of coarse: every hour in one day, every day in one week (weeks start on Monday) and one month. Windows are cut on the stored clock with no time zone (UTC for a zoned column; see [time_window_start]), where no day has 23 or 25 hours, so a sum of the finer counts is the coarser count. A week does not nest in a month.