On this page

MLLib Reference

MLLib.MachineLearning — machine learning primitives for Free Pascal.

MLLib.Analysis adds typed-dense PCA, seeded k-means++, deterministic validation splits, fitted standardization, binary LDA, hierarchical clustering, reproducible decision forests, and an exact low-dimensional TKDTree. Rows are observations and columns are features. See the applied numerics guide for stable 1.8 contracts and explicitly open data-science families.

Typed analysis additions in 1.8

Use TAnalysisKit.FitStandardization only on training rows, then apply the returned TStandardizationModel to training, validation, and test matrices with TransformStandardized. Constant columns receive a safe unit scale. The model owns its means and scales; transforms return new dense matrices. This explicit fit/transform split prevents validation leakage.

HierarchicalCluster(Data, Linkage) supports hlSingle, hlComplete, and hlAverage. It returns the two merged cluster IDs, merge distances, and cluster sizes for each of N-1 deterministic merges. CutHierarchy converts that tree to exactly the requested number of flat clusters. The portable baseline stores pairwise distances and is intended for in-memory data; its worst-case merge search is cubic in the row count.

Choose a forest entry point by target:

Target Train Predict OOB metric
Non-negative integer class FitClassificationForest PredictForestClasses accuracy
Finite real value FitRegressionForest PredictForestValues R-squared

Forests use a caller-selected seed, bootstrap rows, and a deterministic feature-subset/split rule. TDecisionForest.FeatureImportances contains normalized total impurity reduction and OOBScore uses only rows that received at least one out-of-bag prediction. For a small tree count, not every row is guaranteed an OOB prediction. Impurity importance can favour continuous or high-cardinality inputs and is not a causal measure. Models own all tree nodes and do not retain the training matrix.

The inspectable tree layout uses TForestTask, TDecisionTreeNode, TDecisionTreeNodes, TDecisionTreeModel, and TDecisionTrees; nodes record their split feature/threshold, child indexes, leaf flag, and prediction.

All typed-analysis additions reject nil/empty, non-finite, or shape-inconsistent input with EAnalysisError. They are dense, serial, in-memory baselines. Multiclass LDA, gradient boosting, survival forests, factor analysis, and distributed training are not part of 1.8.

---

Quick Start


uses MLLib.MachineLearning;



// Preprocess: scale features to [0,1]

Xscaled := TMLKit.Normalise(X);



// Train a linear regression model

model := TMLKit.LinearRegression(X, Y);

pred  := TMLKit.LinearPredict(model, Xtest);



// Classify with K-Nearest Neighbours

labels := TMLKit.KNearestNeighbours(TrainX, TrainY, TestX, 5);



// Cluster with K-Means

clusters := TMLKit.KMeans(X, 3);



// Reduce dimensions with PCA

pca       := TMLKit.PCA(X, 2);

Xreduced  := TMLKit.PCATransform(pca, X);



// Evaluate

WriteLn(TMLKit.Accuracy(YTrue, YPred));

WriteLn(TMLKit.RMSE(YTrue, YPred));

All methods are class static — no Create/Free needed.

---

Key Types


{ 2-D matrix: array of row vectors }

TDoubleMatrix = array of TDoubleArray;



{ Fitted linear/logistic model }

TLinearModel = record

  Coefficients: TDoubleArray;   { one per feature }

  Intercept:    Double;

  RSquared:     Double;         { training R² (not meaningful for logistic) }

end;



{ K-Means result }

TKMeansResult = record

  Labels:    TIntegerArray;   { cluster index per sample }

  Centroids: TDoubleMatrix;   { K × NFeatures }

  Inertia:   Double;          { within-cluster sum of squared distances }

  Iters:     Integer;

end;



{ PCA result }

TPCAResult = record

  Components:        TDoubleMatrix;   { NComponents × NFeatures }

  ExplainedVariance: TDoubleArray;

  ExplainedRatio:    TDoubleArray;    { fraction of total variance }

  Mean:              TDoubleArray;    { training mean, needed to transform new data }

  Iterations:         TIntegerArray;   { power iterations used per component }

end;



{ DBSCAN result }

TDBSCANResult = record

  Labels:    TIntegerArray;   { -1 = noise }

  NClusters: Integer;

end;



{ Confusion matrix }

TConfusionMatrix = record

  Counts:   TDoubleMatrix;    { NClasses × NClasses; [true][predicted] }

  NClasses: Integer;

end;

---

Typical ML Workflow


Raw data

  ↓ TrainTestSplit (hold out 20%)

  ↓ Normalise / Standardise

  ↓ Train model (LinearRegression, KNN, NaiveBayes, ...)

  ↓ Predict on test set

  ↓ Evaluate (Accuracy, RMSE, R2Score, ConfusionMatrix, ...)

---

Preprocessing

Normalise


Xscaled := TMLKit.Normalise(X);

Scales each feature (column) to [0, 1]. Formula: (x − min) / (max − min). Columns where max = min are set to 0.

When to use: distance- and gradient-based models such as KNN and logistic regression benefit when feature ranges differ substantially.

Standardise


Xscaled := TMLKit.Standardise(X);

Centres each feature at zero mean and scales to unit variance. Formula: (x − μ) / σ. Columns with σ = 0 are set to 0.

When to use: PCA when feature scales should contribute comparably, logistic regression, and models that use distances or dot products. PCA centres input itself; standardisation is optional and changes covariance PCA into a scale-normalised analysis.

TrainTestSplit


TMLKit.TrainTestSplit(X, Y, TestFraction, Seed,

  TrainX, TrainY, TestX, TestY);

Shuffles rows with a deterministic LCG (seed-reproducible) then splits. TestFraction = 0.2 → 20% test, 80% train.

OneHotEncode


M := TMLKit.OneHotEncode(Labels, NClasses);

Converts integer labels [0..NClasses-1] to a binary indicator matrix. Row i has a 1 at column Labels[i] and 0 elsewhere.

---

Regression

LinearRegression (OLS)


model := TMLKit.LinearRegression(X, Y);

// model.Coefficients  — one coefficient per feature

// model.Intercept     — bias term

// model.RSquared      — training R²

Uses centered Householder QR rather than forming X'X, avoiding the condition- number squaring of normal equations. The intercept is fitted internally — do not add a bias column to X. At least as many samples as features are required; rank-deficient designs raise EMLError.

RidgeRegression


model := TMLKit.RidgeRegression(X, Y, Lambda);

L2-regularised OLS: β = (X'X + λI)⁻¹ X'y. The intercept is not regularised.

  • Lambda = 0 → same as OLS
  • Lambda large → coefficients shrink toward zero

When to use: correlated features, N < NFeatures.

PolynomialFeatures


Xpoly := TMLKit.PolynomialFeatures(X1D, Degree);

// Xpoly[i] = [1, x_i, x_i², ..., x_i^Degree]



// LinearRegression fits an intercept, so omit PolynomialFeatures' bias column:

Xpoly := TMLKit.PolynomialFeatures(X1D, Degree, False);

model := TMLKit.LinearRegression(Xpoly, Y);

Expands a 1-D feature vector to a polynomial design matrix. The two-argument compatibility overload includes a bias column. Pass IncludeBias = False when combining it with LinearRegression, which fits its own intercept.

LinearPredict


YHat := TMLKit.LinearPredict(model, Xnew);

Applies a fitted LinearRegression or RidgeRegression model to new data.

---

Classification

KNearestNeighbours


labels := TMLKit.KNearestNeighbours(TrainX, TrainY, TestX, K);

For each test point, finds the K closest training points by Euclidean distance and returns the majority class label.

Tips:

  • Standardise features first (distance is not scale-invariant)
  • K=1 overfits; K=sqrt(N) is a common starting point
  • Use odd K to avoid ties in binary classification

NaiveBayes (Gaussian)


labels := TMLKit.NaiveBayes(TrainX, TrainY, TestX);

Estimates a Gaussian distribution per feature per class, then classifies by maximum log-posterior. Very fast and works well with many features.

When to use: text classification, spam filtering, or when features are (approximately) independent.

LogisticRegression


model := TMLKit.LogisticRegression(TrainX, TrainY);

model := TMLKit.LogisticRegression(TrainX, TrainY, LR, MaxIter, Tol);

labels := TMLKit.LogisticPredict(model, Xnew);

Binary (0/1) logistic regression trained with gradient descent. Sigmoid(Intercept + Dot(Coefficients, x)) ≥ 0.5 → class 1.

Parameters: LR learning rate (0.1), MaxIter (1000), Tol gradient norm (1e-5).

When to use: binary classification. LogisticPredict returns labels only; there is no public probability method, so callers needing scores must evaluate 1 / (1 + Exp(-(Intercept + dot(Coefficients, x)))) themselves. Probability calibration is not assessed by this implementation.

---

Clustering

KMeans


R := TMLKit.KMeans(X, K);

R := TMLKit.KMeans(X, K, MaxIter, Seed);

// R.Labels    — cluster index 0..K-1 per sample

// R.Centroids — K centroid coordinates

// R.Inertia   — total within-cluster squared distance

Lloyd's algorithm: alternate between assigning each point to its nearest centroid and recomputing centroids.

Choosing K: 1. Run KMeans for K = 1, 2, ..., 10 2. Plot Inertia vs K 3. Pick the "elbow" where improvement flattens

Tips:

  • Run multiple times with different seeds; pick lowest Inertia
  • Standardise features first

DBSCAN


R := TMLKit.DBSCAN(X, Eps, MinPts);

// R.Labels    — cluster index or -1 (noise) per sample

// R.NClusters — number of clusters found

Finds arbitrarily-shaped clusters without specifying K. Points in low-density regions are labelled noise (-1).

Choosing parameters:

  • Eps: plot sorted k-distances (k = MinPts); pick the knee
  • MinPts: rule of thumb = 2 × NFeatures; larger = more conservative

---

Dimensionality Reduction

PCA


R   := TMLKit.PCA(X, NComponents);

Xr  := TMLKit.PCATransform(R, Xnew);

// R.Components[k]       — k-th principal component (unit vector in feature space)

// R.ExplainedVariance   — eigenvalue of each component

// R.ExplainedRatio[k]   — fraction of total variance explained by component k

// R.Mean                — training mean (subtracted before projection)

// R.Iterations[k]       — iterations used for component k

Extracts the NComponents directions of maximum variance using power iteration + deflation (no LAPACK needed). Each component is re-orthogonalised, and convergence is sign-invariant. If power iteration exhausts its iteration limit, EMLError is raised instead of returning an unconverged component. Rank-deficient covariance matrices receive a deterministic orthogonal completion with zero explained variance; a completely zero-variance dataset is rejected.

Steps: 1. Standardise X before calling PCA 2. Check ExplainedRatio to decide how many components to keep 3. Use PCATransform to project both training and test data

When to use: visualisation (NComponents=2), noise reduction, or to decorrelate features before logistic regression.

---

Model Evaluation

Function Formula Returns
Accuracy correct / total 0..1
Precision TP / (TP+FP) 0..1
Recall TP / (TP+FN) 0..1
F1Score 2·P·R / (P+R) 0..1
MSE mean (y−ŷ)² ≥ 0
RMSE √MSE ≥ 0, same units as y
| `MAE` | mean |y−ŷ| | ≥ 0 |
| `R2Score` | 1 − SS_res/SS_tot | ≤ 1 |

// Classification

WriteLn(TMLKit.Accuracy(YTrue, YPred));

WriteLn(TMLKit.F1Score(YTrue, YPred, ClassLabel := 1));

CM := TMLKit.BuildConfusionMatrix(YTrue, YPred, NClasses);



// Regression

WriteLn(TMLKit.RMSE(YTrue, YPred));

WriteLn(TMLKit.R2Score(YTrue, YPred));

For a constant YTrue (zero total sum of squares), R2Score returns 1.0 regardless of predictions. Handle that degenerate case separately if your application needs a different convention.

---

Error Handling

EMLError is raised for:

  • Empty, ragged, or zero-feature matrices
  • NaN or infinite feature/target values
  • Feature-count or sample/target length mismatches
  • K > number of training samples (KNN, KMeans)
  • NComponents > NFeatures (PCA)
  • Negative or out-of-range labels, and non-binary logistic-regression labels
  • too few samples or a rank-deficient design in LinearRegression
  • a zero-variance PCA dataset or exhausted PCA power iteration
  • Invalid fractions, iteration counts, tolerances, learning rates, radii, or regularisation parameters

PCATransform is an exception to the general empty-matrix rule: it returns nil for empty input.

---

Dependencies

  • MathBase.SharedTypes — TDoubleArray, TIntegerArray

No other external libraries required.