Skip to main content

Module statistics

Module statistics 

Source

Structs§

AnalysisContext
AnalysisResults
AnalysisRows
The rows an analysis reads, and how many the table has.
CategoricalStatistics
ColumnStatistics
ColumnStats
ComputeOptions
CorrelationMatrix
CorrelationPair
DistributionAnalysis
DistributionCharacteristics
DistributionInfo
NumericStatistics
OutlierAnalysis
OutlierRow
PercentileBreakdown
TemporalStatistics
Describe for a Date, Datetime, Time or Duration column: each statistic a value of the column’s own type, written as the table writes it; None for a null.

Enums§

CorrelationMethod
Which coefficient the correlation matrix shows.
DistributionType
IqrPosition
OutlierMethod

Constants§

RANK_VALUES
The most values Spearman’s ρ ranks: eight bytes each, so 512 MiB beside the rows read. A sample’s 100,000 rows rank up to 671 columns; a read of every row of a large table can be past it, and the matrix then has Pearson’s r only.
SAMPLING_THRESHOLD
Default sampling threshold: datasets >= this size are sampled. Used as fallback when sample_size is None. App uses config value.

Functions§

analysis_results_from_describe
Builds describe-only AnalysisResults from a list of column statistics.
analysis_rows
Read the rows an analysis works on: all of them when the table has no more than sample_rows (or sample_rows is None), and otherwise a seeded sample of that many, spread across the whole table rather than taken from its head.
calculate_theoretical_probability_in_interval
Calculates the probability that a value falls in [lower, upper] for the given distribution.
collect_lazy
Collects a LazyFrame into a DataFrame.
compute_correlation_matrix
Computes pairwise Pearson correlation matrix for all numeric columns.
compute_correlation_pair
Computes correlation statistics for a pair of columns.
compute_describe_from_lazy
Computes describe statistics from a LazyFrame without materializing all rows. When sampling is disabled, runs a single aggregation collect (like Polars describe) for similar performance. When sampling is enabled, samples then runs describe on the sample. Describe statistics for a frame. With sample_size, a table with more rows than that is described from a sample (see analysis_rows); without it, every row is aggregated in one streaming pass, never held. known_total saves a count.
compute_describe_single_aggregation
Computes describe statistics in a single aggregation pass over the DataFrame. Uses one collect() with aggregated expressions for all columns (count, null_count, mean, std, min, percentiles, max).
compute_statistics
Computes statistics for a LazyFrame with default options.
compute_statistics_for_sample
compute_statistics_with_options over the rows a crate::sampling::Sample picks from lf, which is already cut to the sample’s scope.
compute_statistics_with_options
Computes comprehensive statistics for a LazyFrame.
count_rows
Count a frame’s rows.
may_stream
Whether a query over lf may use the streaming engine: asked for, and possible. Polars 0.55’s streaming engine cannot run an anonymous scan (a SQLite table): it stops at a todo!.
slices_reach_into_the_scan
Whether a slice of this plan is read by the scan of one file, skipping what comes before it: true of a single Parquet or IPC file, which seeks by row group, with or without columns stubbed above it. Not of a filter or a CSV, whose slice reads everything ahead of it, nor of a scan of many files, where each slice opens the footer of every file before it — measured on 135 files in S3, fifty slices took longer than streaming all 37 million rows once.