A Lean 4 library for reusable, formalized statistics — 1,368 theorems, checked end to end.

Abstract. statlib builds the foundational layer that statistical and machine-learning formalization needs: stochastic-order asymptotics, uniform integrability, empirical-process tools, high-dimensional concentration, matrix analysis, conformal prediction, and nonparametric approximation on top of Mathlib. Every entry is checked end to end by the Lean compiler — no gaps, no hand-waving.

375
Lean files
269,391
Lines of Lean
1,368
Theorems proved
672
Supporting lemmas
1

What's inside

statlib is built around three main boards, each a tree of topic areas formalized against Mathlib, with supporting packages in the full library tree. Today the work centers on these boards:

1.1

Statistical foundations

Statlib.StatFoundation

The probability and statistics core the rest of the library builds on: stochastic-order asymptotics (big-O/little-o in probability, Slutsky), uniform integrability and integral convergence, conformal prediction, sufficiency and statistical inference, convergence and limit theorems, empirical-process tools, sub-Gaussian and sub-exponential variables, finite-dimensional Gaussian functional inequalities, Lipschitz concentration, and concentration inequalities.

Statistical inference & sufficiency

5 thm · 5 lem

A measure-theoretic foundation for sufficient, ancillary, and complete statistics, culminating in the Lehmann-Scheffé and Basu theorems of classical estimation theory.

Convergence & limit theorems

32 thm · 51 lem

Formalizes the central limit theorems (IID, Lindeberg-Feller, multivariate), the Lévy continuity and Cramér-Wold tools behind weak convergence, and uniform laws of large numbers including Glivenko-Cantelli.

Stochastic-order asymptotics

82 thm · 4 lem

Big-O and little-o in probability, rate bounds, Slutsky-type product theorems, and convergence bridges connecting stochastic orders to classical limit theorems.

Uniform integrability

25 thm · 0 lem

Uniformly integrable families, Vitali convergence, integral-convergence bridges, and conditions ensuring L¹ convergence from convergence in probability.

Empirical processes

9 thm · 6 lem

Rademacher complexity, contraction and symmetrization tools, Dudley entropy, bounded-difference controls, and quantitative Glivenko-Cantelli statements for uniform deviation arguments.

Tail behavior of random variables

94 thm · 70 lem

Concentration and tail-bound theory for sub-Gaussian and sub-exponential random variables, plus Gaussian functional inequalities culminating in dimension-free Lipschitz concentration.

Concentration inequalities

37 thm · 2 lem

Tail and moment bounds controlling how sums and functions of independent (or martingale) random variables deviate from their means.

Conformal prediction

13 thm · 0 lem

Conformal inference foundations: exchangeability, nonconformity scores, conformal quantile regression, and finite-sample coverage guarantees.

Roadmap. Grow shared vocabulary as concepts get reused across areas, broaden the estimation, testing, empirical-process, and asymptotic foundations, and expand stochastic-order and conformal-inference toolchains.

1.2

High-dimensional statistics

Statlib.HighDim

Statistics in the high-dimensional regime: matrix concentration (Hanson–Wright, matrix Bernstein), Wedin sin-theta perturbation, covariance estimation, L1 quadratic-process analysis, L1 RSE bounds from covariance, high-dimensional geometry, regression, and debiased LASSO inference.

Roadmap. Strengthen the operator-convexity, matrix-concentration, covariance, RIP, spectral-perturbation, debiasing, and regression theorem chains; extend L1-process results to broader design matrix families.

1.3

Nonparametric statistics

Statlib.Nonparametric

Approximation and risk vocabulary for nonparametric statistics: finite sieves, Holder classes, high-order tensor-product B-spline and wavelet approximation, RKHS and neural-network rates, conformal quantile regression, and oracle interfaces.

Roadmap. Fill in the remaining nonparametric approximation chains while keeping the sieve and risk interfaces reusable across estimators; extend conformal-prediction results to broader nonparametric settings.

2

Selected results

Matrix analysis

Wedin sin-theta theorem for singular subspaces

max{UUU^U^F,VVV^V^F}2rEop/δ\max\{\|UU^\top-\hat U\hat U^\top\|_F,\|VV^\top-\hat V\hat V^\top\|_F\}\le \sqrt{2r}\,\|E\|_{\mathrm{op}}/\delta

For Ahat = A + E with a separated top-r singular spectrum and small operator-norm perturbation, the left and right singular subspace projectors move by at most a Wedin sin-theta bound.

wedin_sin_theta
High-dimensional regression

Debiased LASSO standard Wald interval coverage

P ⁣(βjb^j±cs^j/an)1α\mathbb{P}\!\left(\beta_j\in \hat b_j \pm c\,\hat s_j/a_n\right)\to 1-\alpha

An iid-score debiased LASSO theorem: row-approximation error, L1 consistency, studentization, and Gaussian critical-value calibration imply asymptotic standard Wald confidence-interval coverage.

tendsto_measure_debiasedLasso_standardWaldCI_coverage_iidScoreSum_real
Nonparametric approximation

Holder-smooth ReLU approximation at the -2s/d rate

fHr+β([0,1]d)    gNL,W: supx[0,1]dg(x)f(x)c(LW2)2(r+β)/df\in\mathcal{H}^{r+\beta}([0,1]^d) \implies \exists g\in\mathcal{N}_{L,W}:\ \sup_{x\in[0,1]^d}|g(x)-f(x)|\le c\,(LW^2)^{-2(r+\beta)/d}

An explicit fixed-width construction is converted into the architecture-scale ReLU approximation rate (L W^2)^(-2(r+beta)/d) for high-dimensional Holder-smooth functions.

holderSmoothBall_unitCube_LW2_rate_from_fixed_width_M_rate
Nonparametric approximation

High-order multivariate B-spline Holder rate

f0Hr+β([0,1]d)    EmK(f0)MmK2(r+β)/df_0\in\mathcal{H}^{r+\beta}([0,1]^d) \implies \mathcal{E}_{m_K}(f_0) \le M\,m_K^{-2(r+\beta)/d}

A positive-degree tensor-product B-spline system on the high-dimensional unit cube achieves the uniform squared-error sieve rate m^(-2(r+beta)/d) over the trace Holder-smooth ball.

unit_cube_bspline_high_order_holder_smooth_uniform_sieve_approximation_rate
High-dimensional concentration

Rectangular matrix Bernstein inequality

P ⁣(kXkt)(p+q)exp ⁣(t2/2σ2+Rt/3)\mathbb{P}\!\left(\Big\|\sum_k X_k\Big\| \ge t\right) \le (p+q)\,\exp\!\left(\frac{-t^2/2}{\sigma^2 + Rt/3}\right)

The matrix Bernstein bound extended to sums of independent centered rectangular p×q random matrices via Hermitian dilation, with variance the max of the two one-sided second-moment norms.

matrix_bernstein_rect
Tail behavior of random variables

Gaussian concentration for Lipschitz functions

P ⁣(f(X)Ef(X)t)2exp ⁣(t22L2)\mathbb{P}\!\left(|f(X)-\mathbb{E}f(X)|\ge t\right) \le 2\exp\!\left(-\frac{t^2}{2L^2}\right)

An L-Lipschitz function of a vector with independent standard Gaussian coordinates has a dimension-free two-sided Gaussian tail around its mean.

gaussian_lipschitz_concentration
Empirical processes

Dudley entropy integral

EsuptTXtC0diam(T)logN(T,ρ,ε)dε\mathbb{E}\sup_{t\in T} X_t \le C\int_0^{\operatorname{diam}(T)}\sqrt{\log N(T,\rho,\varepsilon)}\,d\varepsilon

A chaining-style entropy integral control for finite sub-Gaussian processes.

dudley_entropy_integral

Built to be reused — and extended

statlib aims to fill a foundation gap: reusable, machine-checked infrastructure for statistics and machine learning, built only on Mathlib. Read how it's designed, or help extend it.