Describe for a Date, Datetime, Time or Duration column: each statistic a value of the
column’s own type, written as the table writes it; None for a null.
The most values Spearman’s ρ ranks: eight bytes each, so 512 MiB beside the rows
read. A sample’s 100,000 rows rank up to 671 columns; a read of every row of a
large table can be past it, and the matrix then has Pearson’s r only.
Read the rows an analysis works on: all of them when the table has no more than
sample_rows (or sample_rows is None), and otherwise a seeded sample of that
many, spread across the whole table rather than taken from its head.
Computes describe statistics from a LazyFrame without materializing all rows.
When sampling is disabled, runs a single aggregation collect (like Polars describe) for similar performance.
When sampling is enabled, samples then runs describe on the sample.
Describe statistics for a frame. With sample_size, a table with more rows than
that is described from a sample (see analysis_rows); without it, every row is
aggregated in one streaming pass, never held. known_total saves a count.
Computes describe statistics in a single aggregation pass over the DataFrame.
Uses one collect() with aggregated expressions for all columns (count, null_count, mean, std, min, percentiles, max).
Whether a query over lf may use the streaming engine: asked for, and possible.
Polars 0.55’s streaming engine cannot run an anonymous scan (a SQLite table): it
stops at a todo!.
Whether a slice of this plan is read by the scan of one file, skipping what comes
before it: true of a single Parquet or IPC file, which seeks by row group, with or
without columns stubbed above it. Not of a filter or a CSV, whose slice reads
everything ahead of it, nor of a scan of many files, where each slice opens the
footer of every file before it — measured on 135 files in S3, fifty slices took
longer than streaming all 37 million rows once.