pub fn analysis_rows(
lf: &LazyFrame,
sample_rows: Option<usize>,
known_total: Option<usize>,
seed: u64,
polars_streaming: bool,
) -> Result<AnalysisRows>Expand description
Read the rows an analysis works on: all of them when the table has no more than
sample_rows (or sample_rows is None), and otherwise a seeded sample of that
many, spread across the whole table rather than taken from its head.
Two ways to spread it, chosen by what the plan can do cheaply:
- A plan whose slices reach into a single Parquet or IPC scan reads
[
SAMPLE_BLOCKS] short runs at seeded places across the table. Each run is a row group or two, so a sample of a 400-million-row hive table reads a few dozen row groups, not the table.known_totalsaves the count; the footers give it cheaply otherwise. - Anything else — a filter, a query, a union of files, a CSV — is read once as a stream, keeping the rows whose seeded rank is lowest. That is a uniform sample in bounded memory, and the same pass counts the rows, so a filtered view is read once rather than counted and then read.