Skip to main content

Module distribution_fit

Module distribution_fit 

Source
Expand description

Fitting a column’s values to each candidate distribution, and testing each fit.

One set of parameters per family, estimated once, is what the test, the Q-Q plot and the histogram overlay all use. The test is Kolmogorov-Smirnov, calibrated by a parametric bootstrap: the statistic of the values against their fit is ranked among the statistics of samples drawn from that fit and refitted the same way. That is what makes it a p-value when the parameters come from the data being tested, which the textbook KS table assumes they do not, and it holds for the discrete families, where the table does not apply at all.

A p-value says how surprising the values would be if they came from the fitted distribution. It is not the probability that they did, and the largest one is not the best model: among the families a test does not reject, the one chosen is the one with the lowest AIC.

Structs§

FitTest
A fit and its test.
Rng
A small seeded generator (SplitMix64): the same seed draws the same samples on every machine, which a test of a test needs.

Enums§

FitOutcome
How a family fared against a column.
Fitted
A distribution with its parameters, as fitted to a column.

Constants§

FAMILIES
Every family tested, in the order a tie is listed.
REJECT_P
Below this p-value a family is rejected. At 0.05 the true family of a column is rejected one time in twenty by design, and a verdict that flips between samples of the same data says more about the sample.
REPLICATES
Simulated samples per test. The smallest p-value it can give is 1/200.
TEST_VALUES
The most values a test is run on. The bootstrap costs this times REPLICATES fits per family; past a few hundred values, a KS test rejects any real-world data anyway, which says more about the sample size than about the fit.

Functions§

beta_inc
The regularized incomplete beta I_x(a, b) (Numerical Recipes, betai).
gamma_p
The regularized lower incomplete gamma P(a, x).
gamma_q
Its complement Q(a, x) = 1 - P(a, x), computed directly: in the upper tail P is within 1e-16 of 1, and 1 - P would be noise.
ks_statistic
The two-sided Kolmogorov-Smirnov statistic of sorted values against fitted: the largest gap between the empirical and fitted CDFs, on either side of each step. Tied values are one step, compared at the value and just below it, so a discrete fit is measured where its jumps are.
listing_order
The families as a list is read: tested ones by p-value, highest first, then the ones that do not apply, in FAMILIES order.
ln_gamma
ln Γ(x) for x > 0, by the Lanczos approximation (g = 7, nine terms), accurate to about 1e-15.
normal_cdf
The standard normal CDF, through erfc(x) = Q(1/2, x²): accurate in the tails, where 1 - erf loses everything.
normal_quantile
The standard normal quantile. Acklam’s rational approximation, then one Halley step on the exact CDF: good to about 1e-15. Exactly 0 at the median, -inf and inf at 0 and 1, NaN outside them.
qq_quantiles
Theoretical quantiles for a Q-Q plot of n sorted values: the fitted quantile at each plotting position i / (n + 1).
select
The family the values are consistent with, or Unknown when every test rejects them. Among those not rejected, the lowest AIC; counts are described by a count distribution when one is not rejected, since a density and a probability are not on one scale.
t_cdf
Student’s t CDF with df degrees of freedom.
test_all
Test every family, in parallel: each is independent, and each is a few hundred thousand CDF evaluations.
test_family
Test one family against values, all of them as read.