Files
tensor/stats/doc.go
T

91 lines
4.8 KiB
Go
Raw Permalink Normal View History

2026-09-03 10:00:00 +02:00
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: MIT
// Package stats is the library's statistics: probability distributions,
// descriptive summaries, classical inference, and the models that fit a
// response to a design.
//
// # Distributions
//
// Every univariate law the package carries has a cumulative
// distribution function and a quantile, and the laws the library draws
// from have a matched generator on the house generator
// (tensor.Generator): the normal, exponential, gamma, chi-square,
// Student t, Poisson and binomial families, the Weibull, lognormal and
// Pareto laws, the negative binomial, and the noncentral chi-square, F
// and t distributions. The multivariate normal has a log density and
// draws through the Cholesky factor of its covariance, and the
// Dirichlet has a density, a mean, an interior mode and draws.
//
// Everything rests on three foundations: GammaLower, GammaUpper and
// BetaIncomplete, the regularised incomplete gamma and beta functions.
// A quantile inverts its CDF by a bracketed Newton iteration through
// the distribution's density, falling back to bisection whenever the
// derivative step would leave the bracket, rather than by a closed
// form; the convergence is unconditional on every monotone CDF, and
// every iteration that fails to converge inside its budget is an
// error, never a silently truncated value. A draw is deterministic
// for a given generator state.
//
// # Descriptives and inference
//
// Median, Std, Var and VarSample are the summaries, with the robust
// pair MedianAbsoluteDeviation and TrimmedMean, the sampling
// Quantile, the Histogram family, and the windowed RollingMean,
// RollingSum, RollingMin and RollingMax.
//
// The inference entries are classical frequentist tooling built on the
// package's own distribution functions, so a p-value travels no further
// than the incomplete gamma and beta above: the association matrices
// CovarianceMatrix and CorrelationMatrix, the group comparisons
// WelchTTest, ANOVAOneWay and MannWhitneyU, the goodness-of-fit test
// ChiSquareGoodnessOfFit, the two-sample KolmogorovSmirnovTest, the
// contingency table tests FisherExactTest, ChiSquareIndependence and
// McNemarTest with Cramér's V, the percentile BootstrapCI, the rank
// correlations SpearmanRho and KendallTau beside Pearson, and the
// multiple-testing corrections Bonferroni, Holm and BenjaminiHochberg.
//
// # Models
//
// LinearRegression is ordinary least squares with the full classical
// inference (standard errors, t-tests, R², adjusted R² and the model F
// test), WeightedLinearRegression its weighted counterpart, and
// LogisticRegression and PoissonRegression the generalised linear
// models, fitted by Newton-Raphson on the exact likelihood with Wald
// inference from the inverse Fisher information. LinearMixedModel adds
// the grouped random effects, estimating their covariance and the
// residual variance by residual maximum likelihood.
//
// The regularised family is Lasso, ElasticNet and LassoPath, fitted by
// coordinate descent on the standardised design; the robust family is
// HuberRegression (with HuberRegressionTuned for the tuning constant)
// and TheilSenRegression; QuantileRegression fits the tau-th
// conditional quantile by the Frisch-Newton interior-point method.
// PCA rotates an observation cloud onto its principal components and
// carries the whitening transforms between the two representations,
// KMeans and GaussianMixture (with GaussianMixtureBIC) partition or
// model it, HierarchicalClustering records the full merge tree of the
// agglomerative construction for cutting afterwards,
// FitHiddenMarkovModel fits a hidden Markov model over discrete
// sequences with Forward, Smooth and Viterbi answering the filtered
// and smoothed posteriors and the most likely path of a fitted or
// hand-built model, and GaussianProcessRegression conditions the prior
// a Kernel defines on the observations, with MarginalLogLikelihood
// exposed as the objective of a hyperparameter fit.
//
// # Contracts
//
// Every entry point returns a value with an error, and every error
// carries the library's "tensor: " prefix and names the entry point
// that raised it. The estimation entries refuse complex input and,
// by name, any non-finite observation: a single NaN would otherwise
// spread silently through a whole result. Arrays are built through the
// root package's constructors (tensor.FromFloats and its siblings), and
// a regression design carries n rows and p columns with the intercept
// included by the caller as a constant column when one is wanted.
//
// The package depends only on the library's own internal packages, so a
// caller that fits the hyperparameters of a Gaussian process drives
// MarginalLogLikelihood from outside with the house minimiser.
package stats