Skip to main content

Module sampling

Module sampling 

Source
Expand description

The one sampler every analysis tool reads through: which rows (a scope), how they are picked (a method), how many, and the seed. Describe, Distribution, Correlation and Data Quality all take their rows from read, so a sample means the same thing whichever tool shows it.

Structs§

AnalysisRows
The rows an analysis reads, and how many the table has.
PerValue
What an equal-per-value sample learned beside its rows, from the same pass.
ReadWatch
A read’s line to the screen: a stop flag set on cancel and a count of rows seen, both shared with the UI. Streamed reads check stop between batches, block reads between blocks; a single collect runs to its end.
Sample
Which rows an analysis reads and how it picks them.
SampleSource
Where a tool’s rows come from before scoping: the table as shown, or the loaded source with its footer facts. Built on the UI thread, cut on the worker (a source scan can read its schema).

Enums§

Counted
What a streamed pass counted beside its sample.
SampleMethod
How the rows of a scope are picked.
SizeError
Why a typed sample size cannot be read.

Constants§

CANCELLED
What a read that was stopped says. Its work is dropped, never shown as a result.
DEFAULT_SAMPLE_ROWS
The default sample size, before [analysis] sample_rows says otherwise.
MAX_COUNTED_KEYS
Distinct keys a pass counts before it gives up counting. Past this the grain is finer than a report can show, and the map would grow with the table.

Functions§

analysis_rows
The rows an analysis works on: all when the table has at most sample_rows (or it is None), else a seeded sample spread across the table:
count_rows
Count a frame’s rows.
no_rows_error
A chosen scope matching nothing is an error (a value not in the data, a range past its end), not an empty sample; the table as shown may simply be empty.
parse_size
A typed sample size: 50000, 50,000, 50_000, 50k, 2m, 2.5M. A size of no rows is refused; past usize it saturates and the caller clamps.
read
Read the rows sample asks for from a frame already scoped. known_total spares a count; without one, first-rows reports the rows read as the total (sample_size: None) rather than count.
slices_reach_into_the_scan
Whether a slice of this plan is read by one file’s scan skipping ahead: a single Parquet or IPC file (seeking by row group), stubbed columns allowed. Not a filter or CSV (reads everything before), nor many files (each slice opens every earlier footer; on 135 S3 files fifty slices beat streaming 37M rows). Asked of the optimized plan (SLICE: Positive in the SCAN, or a SLICE[ node); if Polars changes its plan text, this says no and streaming takes over: slower, never wrong.
view_scope_rows
How many rows a view scope holds, from the view’s row count; None for a source scope, whose size only a read can tell.

Type Aliases§

HeldJudge
What ReadWatch::hold asks of the bytes a sampler holds and the rows they are.