On this page

StatsLib

Statistical analysis domain for Free Pascal covering descriptive statistics, hypothesis testing, correlation, normality tests, robust methods, bootstrap, and non-parametric tests.

Depends on: MathBase

Learning routes

Beginner route

Copy and run the descriptive-statistics quick start. It uses a TDoubleArray, prints a checked correlation and seeded bootstrap interval, and returns newly allocated result arrays where the operation requires them.

Common tasks and algorithm choice

Task Start with Contract or failure guidance
One in-memory sample TStatsKit.Describe Descriptive methods
Incremental or mergeable sample TOnlineStatistics Streaming contract
Mean/group comparison TInferenceKit test selected by design Test selection
Distribution fit EstimateNormal/matching estimate Distribution selection
Regression with diagnostics FitOLS or FitLogistic Regression selection

Advanced route

Run example 19 for streaming statistics in a typed data pipeline and example 21 for inference. Both consume the shared double-real containers directly; the transition from TStatsKit to fitted/streaming records needs no glue copy.

Units

Unit File Class
StatsLib.Stats StatsLib.Stats.pas TStatsKit
StatsLib.Streaming StatsLib.Streaming.pas TOnlineStatistics
StatsLib.Inference StatsLib.Inference.pas TInferenceKit

---

Streaming statistics in 1.8

TOnlineStatistics calculates count, weighted mean, extrema, and population or sample variance in constant retained memory. Independent records can be merged and choose explicit reject/ignore handling for non-finite observations. The applied numerics guide documents its bounded state, ownership, and error contracts.

Inference, distributions, and regression in 1.8

StatsLib.Inference is the stable typed inference surface. It is separate from the older TStatsKit compatibility methods so every new result can carry its assumptions and diagnostics.

Distribution selection

Model Construct Operations
Normal TNormalDistribution.Create(Mean, StandardDeviation) PDF, LogPDF, CDF, Survival, Quantile, Sample
Exponential TExponentialDistribution.Create(Rate) PDF, LogPDF, CDF, Survival, Quantile, Sample
Binomial TBinomialDistribution.Create(Trials, Probability) PMF, LogPMF, CDF, Survival, Quantile, Sample

Sampling takes var TLocalRandom; it is reproducible and never uses global RandSeed. CDF and Survival are paired to avoid cancellation in ordinary tails. Quantiles are monotone and accept the documented probability endpoints, returning mathematical infinities where appropriate.

EstimateNormal, EstimateExponential, EstimateGamma, and EstimateBinomial return TDistributionEstimate. Inspect Status and Identifiable before using StandardErrors; a returned parameter value does not imply that uncertainty was identifiable.

Test selection

Question API Effect/interval
One mean versus a constant OneSampleT standardized mean and confidence interval
Matched observations PairedT paired standardized difference and interval
Two independent means, unequal variance WelchT standardized difference and Welch degrees of freedom
Three or more independent means OneWayANOVA eta-squared
Categorical association ChiSquareContingency Cramer's V
Two independent ordinal/continuous samples MannWhitneyU rank-biserial effect
Family-wise p-value control AdjustBonferroni adjusted p-values in input order
False-discovery-rate control AdjustBenjaminiHochberg monotone BH adjusted p-values in input order

All test records include the statistic, p-value, degrees of freedom where applicable, and an effect size. T-test records also include a confidence interval. These are two-sided tests. The Mann-Whitney p-value uses the tie-corrected normal approximation in this unit.

Regression selection and failure contracts

FitOLS(Design, Response) uses the shared compact SVD. It returns coefficients, standard errors, fitted values, residuals, R-squared, adjusted R-squared, numerical rank, degrees of freedom, and Status. A rank-deficient design is reported rather than hidden; covariance-style standard errors require full rank and positive residual degrees of freedom.

FitLogistic(Design, Labels) fits binary labels using Newton/IRLS steps. Status, Iterations, and Identifiable distinguish convergence from complete/quasi separation or singular information. Probabilities and the best finite coefficient estimate remain inspectable on a non-identifiable result.

Every new inference API rejects non-finite data, inconsistent shapes, invalid probabilities, and insufficient sample sizes with EInferenceError. The TCountMatrix alias represents contingency rows. The implementation is a portable double-precision baseline; survival analysis, factor analysis, robust covariance families, multinomial/count GLMs, and certified exact small-sample tables remain outside 1.8.

Core Types

Exception


EStatsError = class(Exception);

Raised on empty input arrays, insufficient data for a given calculation, or invalid parameter values.

TDescriptiveStats Record

A single call to TStatsKit.Describe(Data) populates all fields:

Field Type Description
N Integer Number of observations
Mean Double Arithmetic mean
Median Double Middle value of sorted data
Mode Double Most frequent value (first mode if multimodal)
Q1 Double First quartile (25th percentile)
Q3 Double Third quartile (75th percentile)
Min Double Minimum value
Max Double Maximum value
Range Double Max − Min
IQR Double Q3 − Q1
Variance Double Sample variance (n − 1 denominator)
StdDev Double Population standard deviation (n denominator)
Skewness Double Distribution asymmetry; positive = right tail
Kurtosis Double Sample excess kurtosis; 0 for normal
SEM Double Standard error of the mean
CV Double Coefficient of variation (%)

function TDescriptiveStats.ToString: string;     // Vertical summary

function TDescriptiveStats.ToStringWide: string; // Table-like summary

---

TStatsKit — Static Methods

All methods are class ... static — no instance required. Input data is always TDoubleArray.

Descriptive Statistics

Method Min N Description
Mean(Data) 1 Arithmetic mean
Median(Data) 1 Middle value; average of two middles for even N
Mode(Data) 1 Most frequent value
Range(Data) 1 Max − Min
Describe(Data) 4 Full TDescriptiveStats record; inherited constraints (such as non-zero spread and mean) also apply

var Data: TDoubleArray;

    Stats: TDescriptiveStats;

begin

  Data := TDoubleArray.Create(1, 2, 3, 4, 5);

  Stats := TStatsKit.Describe(Data);

  Writeln(Stats.ToString);

end.

Variance and Standard Deviation

Method Denominator Min N
Variance(Data) n − 1 2
SampleVariance(Data) n − 1 2 (same calculation as Variance)
StandardDeviation(Data) n (population) 2
SampleStandardDeviation(Data) n − 1 2

Distribution Measures and Aggregates


class function Skewness(const Data: TDoubleArray): Double;            // N >= 3

class function Kurtosis(const Data: TDoubleArray): Double;            // Sample excess kurtosis; N >= 4

class function StandardErrorOfMean(const Data: TDoubleArray): Double; // SampleStdDev / sqrt(N); N >= 2

class function CoefficientOfVariation(const Data: TDoubleArray): Double; // PopulationStdDev / abs(Mean) * 100

class function GeometricMean(const Data: TDoubleArray): Double;       // Values must be > 0

class function HarmonicMean(const Data: TDoubleArray): Double;        // Values must be non-zero

class function Sum(const Data: TDoubleArray): Double;

class function SumOfSquares(const Data: TDoubleArray): Double;

Percentiles and Quartiles


class function Percentile(const Data: TDoubleArray; const P: Double): Double; // 0 <= P <= 100; method R-7

class function Quartile1(const Data: TDoubleArray): Double;   // Percentile(Data, 25)

class function Quartile3(const Data: TDoubleArray): Double;   // Percentile(Data, 75)

class function InterquartileRange(const Data: TDoubleArray): Double; // Q3 − Q1

class function Quantile(const Data: TDoubleArray; const Q: Double): Double; // 0 <= Q <= 1

Percentile uses linear interpolation (Excel/R default, method R-7).

Correlation and Covariance


class function PearsonCorrelation(const X, Y: TDoubleArray): Double;  // r ∈ [−1, 1]

class function SpearmanCorrelation(const X, Y: TDoubleArray): Double; // ρ ∈ [−1, 1]

class function Covariance(const X, Y: TDoubleArray): Double;          // Sample covariance

Standardisation


class function ZScore(const Value, AMean, StdDev: Double): Double;

class procedure Standardize(var Data: TDoubleArray);

Standardize modifies Data in place using the population standard deviation. It raises EStatsError for constant data. StatsLib does not expose a min-max normalisation routine.

Hypothesis Testing


class function TTest(const X, Y: TDoubleArray; out TPValue: Double): Double;

TTest is an independent two-sample, equal-variance (pooled) t-test. It returns the t-statistic and writes the two-sided p-value to TPValue; each group needs at least two observations. There are no one-sample or paired t-test entry points.

Normality Tests


class function KolmogorovSmirnovTest(const Data: TDoubleArray; out KSPValue: Double): Double;

class function IsNormal(const Data: TDoubleArray; const Alpha: Double = 0.05): Boolean;

class function ShapiroWilkTest(const Data: TDoubleArray; out WPValue: Double): Double;

KolmogorovSmirnovTest requires at least five values, returns the one-sample D statistic, and writes an asymptotic two-sided p-value to KSPValue. Parameters of the comparison normal distribution are estimated from the sample, so this is an approximate diagnostic rather than an exact Lilliefors test. Its empirical CDF steps are evaluated explicitly as Double values so FPC targets do not reduce them to single precision through untyped-literal evaluation. ShapiroWilkTest accepts 3 through 5000 values, returns W, and writes a Royston-approximation p-value. P-values are clamped to [0, 1]; use specialist software when a regulatory or publication workflow requires independently certified inference.

Non-Parametric Tests

Method Description
SignTest(X, Y) Proportion of non-tied pairs for which X[i] > Y[i]
WilcoxonSignedRank(Data1, Data2) Wilcoxon signed-rank W statistic
MannWhitneyU(Data1, Data2, out PValue) Smaller Mann-Whitney U; exact two-sided p-value for untied samples with total N <= 50, otherwise tie-corrected normal approximation with continuity correction
KendallTau(X, Y) Kendall's τ correlation

Robust Statistics


class function MedianAbsoluteDeviation(const Data: TDoubleArray): Double;  // MAD

class function RobustStandardDeviation(const Data: TDoubleArray): Double;  // 1.4826 * MAD

class function HuberM(const Data: TDoubleArray; K: Double = 1.5): Double;

class function TrimmedMean(const Data: TDoubleArray; Percent: Double): Double;

class function WinsorizedMean(const Data: TDoubleArray; Percent: Double): Double;

MAD and the Huber M-estimator are resistant to outliers.

Effect Size


class function CohensD(const Data1, Data2: TDoubleArray): Double;   // Cohen's d

class function HedgesG(const Data1, Data2: TDoubleArray): Double;   // Hedges' g (bias-corrected)

CohensD uses the pooled sample variance. Interpretation: 0.2 = small, 0.5 = medium, 0.8 = large.

Bootstrap Methods


class function BootstrapMean(const Data: TDoubleArray;

  Iterations: Integer): TDoubleArray; overload;

class function BootstrapMean(const Data: TDoubleArray;

  Iterations: Integer; Seed: LongWord): TDoubleArray; overload;



class function BootstrapConfidenceInterval(const Data: TDoubleArray;

  Alpha: Double = 0.05; Iterations: Integer = 1000): TDoublePair; overload;

class function BootstrapConfidenceInterval(const Data: TDoubleArray;

  Alpha: Double; Iterations: Integer; Seed: LongWord): TDoublePair; overload;



class function RandomSample(const Data: TDoubleArray): TDoubleArray;

BootstrapMean returns one resampled mean per iteration. TDoublePair (from MathBase.SharedTypes) holds the lower and upper confidence bounds.

The compatibility overloads use caller-managed global random state and never call Randomize. The seeded overloads use the shared explicit-state generator and unbiased bounded-index sampling; they are reproducible and do not change the process-wide RandSeed.

---

Quick Start


uses StatsLib.Stats, MathBase.SharedTypes;



var

  Data1, Data2: TDoubleArray;

  Stats: TDescriptiveStats;

  CI: TDoublePair;

  r: Double;

begin

  Data1 := TDoubleArray.Create(2.1, 3.5, 2.8, 4.2, 3.0, 3.7, 2.5, 4.1);

  Data2 := TDoubleArray.Create(3.0, 4.1, 3.5, 5.0, 3.8, 4.5, 3.2, 5.1);



  Stats := TStatsKit.Describe(Data1);

  Writeln(Stats.ToStringWide);



  r := TStatsKit.PearsonCorrelation(Data1, Data2);

  Writeln('r = ', r:0:4);



  CI := TStatsKit.BootstrapConfidenceInterval(Data1, 0.05, 2000, 2026);

  Writeln('95% CI: [', CI.Lower:0:4, ', ', CI.Upper:0:4, ']');

end.

Expected output contains:


r = 0.9880

95% CI: [2.7747, 3.7375]

---

Design Notes

  • Variance and SampleVariance both use the n − 1 (Bessel-corrected) denominator.
  • StandardDeviation uses the n (population) denominator; use SampleStandardDeviation for the n − 1 version.
  • Skewness uses the population standard deviation; Kurtosis uses the sample standard deviation.
  • Percentile uses method R-7 (linear interpolation), matching Excel's PERCENTILE and R's default.
  • Sum, SumOfSquares, and Sort accept empty arrays; the sums return zero and sorting is a no-op. Sort is a stable O(n log n) merge sort, modifies its argument in place, and orders NaNs after finite values.
  • Normality-test p-values are approximations: K-S uses the asymptotic distribution despite estimating normal parameters from the sample, and Shapiro-Wilk uses Royston's approximation.
  • Prefer the seeded bootstrap overloads for reproducible tests and analyses. Use Randomize once in the application only when intentionally using the global-RNG overloads.
  • EStatsError is raised for empty arrays, arrays too small for a given statistic, or out-of-range parameters.

Common mistakes

  • Sample versus population convention. In TDescriptiveStats, Variance uses the sample denominator (n−1) while StdDev uses the population denominator (n); the two are not the same convention.
  • Two-sided tests by default. The hypothesis tests report two-sided p-values; express the question so a one-sided statement is not read into a two-sided result.
  • Inference diagnostics. FitOLS covariance-style standard errors require full rank and positive residual degrees of freedom, and distribution or logistic estimates need Status/Identifiable inspection before their uncertainty is used.
  • Streaming policy. TOnlineStatistics keeps constant retained state and applies an explicit reject/ignore policy for non-finite observations.