Skip to main content

Module data_quality

Module data_quality 

Source

Structs§

CategoryVariantGroup
ColumnQualityProfile
CopyRead
A local copy of a remote source that a full scan’s passes read.
DataQualityPlan
DataQualityResults
DuplicateExample
One group of identical rows: how many there are, and the row, a value a column.
ExpectedWindows
Which windows a study expects rows in, from Setup: every window of the grain or weekdays only, within a range. Unset, no window is a gap.
FindingExamples
A few of the values behind one column’s finding, from the rows the run kept.
IdentityProfile
ObservedReads
What a run’s reads of the source were seen to do, counted as the rows went by. Bytes and requests are not counted: a Polars scan does not report them.
QualityConflictScan
Reads named columns of named files at each file’s own type, the only way back to conflict-hidden values: one column of the few disagreeing files.
QualityFileEvidence
One file behind a drift observation: what it holds, and what that costs the column.
QualityObservation
QualityPhase
A stage, whether it reads the source or works on rows already read, and whether a cancel stops it partway.
QualitySample
The rows a sampled run read, kept beside its results. Keyed by acquisition (dataset, view, scope, method, size, seed); other plan settings are the report’s, so a run changing only those re-cuts these rows instead of reading the source. Every column and each row’s position are kept.
QualitySourceContext
QualityWatch
A run’s line to the screen: stages as entered, rows its reads have seen, and a stop checked between stages and between read batches.
SegmentChange
One column’s measure in a segment, and in the segment it is compared with.
SegmentQualityProfile
SharedNulls
Columns null the same number of times, and the rows null in all at once: when equal they go missing together, one fact rather than one per column.
TemporalLatencyProfile
TemporalRoleAssignment
TimeInterpretation
A text column read as a date or datetime for one study: grain and time roles see the parsed value, other checks the stored text. Unreadable values count as unparsed, never as missing.
UnsampledSegment
A segment the scope has rows in and a sample drew none of.

Enums§

CopyPlan
How a remote full scan gets its rows, as Setup says: one fetch into a local copy all passes read, an earlier copy, or a source pass per check.
IntervalClock
Which time puts an interval in a window under a time-window grain: the grain’s column, or the interval’s start or end (by end, counted on the day it finished).
IntervalFact
What an interval’s detail counts, each with the rows behind it.
NoCopy
Why a remote full scan reads the source in each pass instead of a local copy.
ObservationKind
QualityComparison
QualityCompute
QualityGrain
QualityMetric
QualityPage
QualityPrecision
QualityScope
QualitySetup
What an empty page is missing, which Enter opens in Setup.
QualityStage
What a Data Quality run is doing now. The worker names each stage as it enters it, and the progress view shows the latest.
SegmentCount
Where a run’s exact segment totals come from, as Setup says before Run.
TemporalRole
TextReading
What the values of a text column parse as, most specific first.
TimeKind
Whether text read as time is a date or a date with a time of day.

Constants§

INTERVAL_PAIRS
The intervals measured when none are chosen: start role to end role, the pairs whose order the roles state. Others are chosen in Setup; a role in no interval measures nothing, as Setup says.
KEY_LIKE_UNIQUENESS
How nearly unique a column must be before its repeats are worth naming: an id repeating twice in a million rows is a finding, a category repeating is not.
MAX_FINDING_EXAMPLES
Values kept per finding from the rows a run read, and groups of duplicate rows: enough to recognize the problem in the detail, which opens the rest.
QUALITY_SOURCE_FILE_COLUMN
QUALITY_WINDOW_WIDTHS
Window widths offered for time-window grain, in the order the plan cycles them.
TEXT_READING_SHARE
The share of non-null text that must parse before a column is said to hold numbers or dates.
TIME_FORMATS
The formats Setup offers for reading text as time, unambiguous first. Named, not inferred: every row reads the same way, and an unreadable value is counted.

Functions§

add_signal_observations
Read every sample of audio once and add clipping, zero runs and DC offset per channel to results; only for a full run over every frame (the caller decides).
apply_quality_scope
beyond_noise
Whether rates a of n_a rows and b of n_b rows differ by more than two samples of those sizes would by chance (a two-proportion z-test).
compute_data_quality_watched
The data quality of lf by plan, cutting kept instead of reading when it serves the plan. Names each stage to watch and stops between stages (or inside a streamed read) once cancelled. A sampled run’s rows come back even if it stopped: the read is paid for, and the next run can cut it.
duplicate_rows
The rows the duplicate check counted: rows equal to another in every keys column, copies together, most first, in one pass. Each group’s key is its rows, repeated by count rather than looked up again.
fresh_segment_count
Where a run reading a new sample gets segment totals. Seeded runs of one file (may_read_blocks) and the head see too few rows; any other sample is one streamed pass that counts the grain’s key as it goes.
interval_label
event to received.
interval_passes
How many groupings a run’s intervals take: one per distinct grain. A full run reads the scope once for each.
page_setup
The plan setting a result page needs before it has anything to show, if any. Intervals need time roles, and ask only when there are dates to assign.
prepare_source_quality_scan
Prepare the loaded source in the worker, keeping only a provenance index and replacing binary payloads before any value collection.
sampler_counts_segments
Whether sampling plan’s rows also counts its segments: an equal-per-value sample by the grain’s column counts every value as it streams.
segment_changes
Every column’s measures in segment index, beside its comparison segment largest move first, or alone worst first; zero on both sides is left out.
segment_order
The order Segments lists its rows in: as they fall, or the clearest change first (ties, and segments with no clear change, keep their order).
segments_need_count
Whether a sampled run of plan counts its segments’ rows: partitions and time windows are counted for exact totals; files and row chunks are known without it.
shows_trend
Whether the Trends page can draw a column’s measure across segments: that needs segments in an order, and more than one of them.
signal_observations
Observations from crate::formats::audio::SignalReports: a channel with runs at full scale, runs of exact zeros, or a mean 1% of full scale or more from zero.
text_reading
The one typed reading a text column supports, with how many parse: the most specific of overlapping readings. Whole only when every number is.
unparsed_text
The rows of a parseable-text column its reading does not parse: non-null text that stops a cast. None when the column has no reading.
window_cadence
A window width as a cadence: 1d is daily.
window_nests
Whether windows of width fine nest exactly in coarse: hours in a day, days in a week (Monday start) and in a month; weeks not in months. Windows are cut on the zoneless stored clock (UTC for zoned; see time_window_start), so no day has 23 or 25 hours and finer counts sum to the coarser.