Skip to main content

Module distribution_fit

Module distribution_fit 

Source
Expand description

Fitting a column to each candidate distribution and testing each fit. One parameter set per family, estimated once, serves the test, Q-Q plot and histogram overlay. The test is Kolmogorov-Smirnov calibrated by parametric bootstrap (the statistic ranked among those of samples drawn from the fit and refitted), so it is a valid p-value with estimated parameters and for discrete families. A p-value is not the probability of the model, and the largest is not the best: among unrejected families, the lowest AIC wins.

Structs§

FitTest
A fit and its test.
Rng
A small seeded generator (SplitMix64): the same seed draws the same samples on every machine, which a test of a test needs.

Enums§

FitOutcome
How a family fared against a column.
Fitted
A distribution with its parameters, as fitted to a column.

Constants§

FAMILIES
Every family tested, in the order a tie is listed.
REJECT_P
Below this p-value a family is rejected. At 0.05 the true family of a column is rejected one time in twenty by design, and a verdict that flips between samples of the same data says more about the sample.
REPLICATES
Simulated samples per test. The smallest p-value it can give is 1/200.
TEST_VALUES
The most values a test is run on. The bootstrap costs this times REPLICATES fits per family; past a few hundred values, a KS test rejects any real-world data anyway, which says more about the sample size than about the fit.

Functions§

beta_inc
The regularized incomplete beta I_x(a, b) (Numerical Recipes, betai).
gamma_p
The regularized lower incomplete gamma P(a, x).
gamma_q
Its complement Q(a, x) = 1 - P(a, x), computed directly: in the upper tail P is within 1e-16 of 1, and 1 - P would be noise.
ks_statistic
The two-sided KS statistic of sorted values against fitted: the largest CDF gap on either side of each step. Ties are one step, compared at and just below the value, so discrete fits are measured at their jumps.
listing_order
The families as a list is read: tested ones by p-value, highest first, then the ones that do not apply, in FAMILIES order.
ln_gamma
ln Γ(x) for x > 0, by the Lanczos approximation (g = 7, nine terms), accurate to about 1e-15.
normal_cdf
The standard normal CDF, through erfc(x) = Q(1/2, x²): accurate in the tails, where 1 - erf loses everything.
normal_quantile
The standard normal quantile. Acklam’s rational approximation, then one Halley step on the exact CDF: good to about 1e-15. Exactly 0 at the median, -inf and inf at 0 and 1, NaN outside them.
qq_quantiles
Theoretical quantiles for a Q-Q plot of n sorted values: the fitted quantile at each plotting position i / (n + 1).
select
The family the values are consistent with, or Unknown when all are rejected: the lowest AIC among the unrejected, preferring a count distribution for counts (density and probability are not on one scale).
t_cdf
Student’s t CDF with df degrees of freedom.
test_all
Test every family, in parallel: each is independent, and each is a few hundred thousand CDF evaluations.
test_family
Test one family against values, all of them as read.