Expand description
Fitting a column’s values to each candidate distribution, and testing each fit.
One set of parameters per family, estimated once, is what the test, the Q-Q plot and the histogram overlay all use. The test is Kolmogorov-Smirnov, calibrated by a parametric bootstrap: the statistic of the values against their fit is ranked among the statistics of samples drawn from that fit and refitted the same way. That is what makes it a p-value when the parameters come from the data being tested, which the textbook KS table assumes they do not, and it holds for the discrete families, where the table does not apply at all.
A p-value says how surprising the values would be if they came from the fitted distribution. It is not the probability that they did, and the largest one is not the best model: among the families a test does not reject, the one chosen is the one with the lowest AIC.
Structs§
- FitTest
- A fit and its test.
- Rng
- A small seeded generator (SplitMix64): the same seed draws the same samples on every machine, which a test of a test needs.
Enums§
- FitOutcome
- How a family fared against a column.
- Fitted
- A distribution with its parameters, as fitted to a column.
Constants§
- FAMILIES
- Every family tested, in the order a tie is listed.
- REJECT_
P - Below this p-value a family is rejected. At 0.05 the true family of a column is rejected one time in twenty by design, and a verdict that flips between samples of the same data says more about the sample.
- REPLICATES
- Simulated samples per test. The smallest p-value it can give is 1/200.
- TEST_
VALUES - The most values a test is run on. The bootstrap costs this times
REPLICATESfits per family; past a few hundred values, a KS test rejects any real-world data anyway, which says more about the sample size than about the fit.
Functions§
- beta_
inc - The regularized incomplete beta
I_x(a, b)(Numerical Recipes,betai). - gamma_p
- The regularized lower incomplete gamma
P(a, x). - gamma_q
- Its complement
Q(a, x) = 1 - P(a, x), computed directly: in the upper tailPis within 1e-16 of 1, and1 - Pwould be noise. - ks_
statistic - The two-sided Kolmogorov-Smirnov statistic of sorted
valuesagainstfitted: the largest gap between the empirical and fitted CDFs, on either side of each step. Tied values are one step, compared at the value and just below it, so a discrete fit is measured where its jumps are. - listing_
order - The families as a list is read: tested ones by p-value, highest first, then the
ones that do not apply, in
FAMILIESorder. - ln_
gamma ln Γ(x)forx > 0, by the Lanczos approximation (g = 7, nine terms), accurate to about 1e-15.- normal_
cdf - The standard normal CDF, through
erfc(x) = Q(1/2, x²): accurate in the tails, where1 - erfloses everything. - normal_
quantile - The standard normal quantile. Acklam’s rational approximation, then one Halley step
on the exact CDF: good to about 1e-15. Exactly 0 at the median,
-infandinfat 0 and 1, NaN outside them. - qq_
quantiles - Theoretical quantiles for a Q-Q plot of
nsorted values: the fitted quantile at each plotting positioni / (n + 1). - select
- The family the values are consistent with, or
Unknownwhen every test rejects them. Among those not rejected, the lowest AIC; counts are described by a count distribution when one is not rejected, since a density and a probability are not on one scale. - t_cdf
- Student’s t CDF with
dfdegrees of freedom. - test_
all - Test every family, in parallel: each is independent, and each is a few hundred thousand CDF evaluations.
- test_
family - Test one family against
values, all of them as read.