On this page
MLLib Reference
MLLib.MachineLearning — machine learning primitives for Free Pascal.
MLLib.Analysis adds typed-dense PCA, seeded k-means++, deterministic validation splits, fitted standardization, binary LDA, hierarchical clustering, reproducible decision forests, and an exact low-dimensional TKDTree. Rows are observations and columns are features. See the applied numerics guide for stable 1.8 contracts and explicitly open data-science families.
Learning routes
Beginner route
Copy and run the double-real regression quick start. It creates a small row-major TDoubleMatrix, allocates a fitted model, and prints y = 1.0 + 2.0*x.
Common tasks and algorithm choice
| Task | Start with | Contract or failure guidance |
|---|---|---|
| Scale features | fitted standardisation for train/test work | Preprocessing |
| Continuous prediction | LinearRegression / typed FitOLS |
Regression |
| Classification | logistic, KNN, or forest by data/diagnostics | Classification |
| Unlabelled grouping | seeded k-means++ or hierarchical clustering | Clustering |
| Lower-dimensional representation | typed PCA | Dimensionality reduction |
Advanced route
Run example 19 for typed PCA and seeded clustering or example 21 for fitted standardisation and forests. Nested beginner matrices convert through the explicit typed factories described by the advanced guide; fitted train/test paths themselves need no undocumented glue.
Typed analysis additions in 1.8
Use TAnalysisKit.FitStandardization only on training rows, then apply the returned TStandardizationModel to training, validation, and test matrices with TransformStandardized. Constant columns receive a safe unit scale. The model owns its means and scales; transforms return new dense matrices. This explicit fit/transform split prevents validation leakage.
HierarchicalCluster(Data, Linkage) supports hlSingle, hlComplete, and hlAverage. It returns the two merged cluster IDs, merge distances, and cluster sizes for each of N-1 deterministic merges. CutHierarchy converts that tree to exactly the requested number of flat clusters. The portable baseline stores pairwise distances and is intended for in-memory data; its worst-case merge search is cubic in the row count.
Choose a forest entry point by target:
| Target | Train | Predict | OOB metric |
|---|---|---|---|
| Non-negative integer class | FitClassificationForest |
PredictForestClasses |
accuracy |
| Finite real value | FitRegressionForest |
PredictForestValues |
R-squared |
Forests use a caller-selected seed, bootstrap rows, and a deterministic feature-subset/split rule. TDecisionForest.FeatureImportances contains normalized total impurity reduction and OOBScore uses only rows that received at least one out-of-bag prediction. For a small tree count, not every row is guaranteed an OOB prediction. Impurity importance can favour continuous or high-cardinality inputs and is not a causal measure. Models own all tree nodes and do not retain the training matrix.
The inspectable tree layout uses TForestTask, TDecisionTreeNode, TDecisionTreeNodes, TDecisionTreeModel, and TDecisionTrees; nodes record their split feature/threshold, child indexes, leaf flag, and prediction.
All typed-analysis additions reject nil/empty, non-finite, or shape-inconsistent input with EAnalysisError. They are dense, serial, in-memory baselines. Multiclass LDA, gradient boosting, survival forests, factor analysis, and distributed training are not part of 1.8.
---
Quick Start
uses MathBase.SharedTypes, MLLib.MachineLearning;
var
X: TDoubleMatrix;
Y: TDoubleArray;
Model: TLinearModel;
begin
X := TDoubleMatrix.Create(
TDoubleArray.Create(0),
TDoubleArray.Create(1),
TDoubleArray.Create(2));
Y := TDoubleArray.Create(1, 3, 5);
Model := TMLKit.LinearRegression(X, Y);
WriteLn('y = ', Model.Intercept:0:1, ' + ',
Model.Coefficients[0]:0:1, '*x');
end.
Expected output:
y = 1.0 + 2.0*x
All methods are class static — no Create/Free needed.
---
Key Types
{ 2-D matrix: array of row vectors }
TDoubleMatrix = array of TDoubleArray;
{ Fitted linear/logistic model }
TLinearModel = record
Coefficients: TDoubleArray; { one per feature }
Intercept: Double;
RSquared: Double; { training R² (not meaningful for logistic) }
end;
{ K-Means result }
TKMeansResult = record
Labels: TIntegerArray; { cluster index per sample }
Centroids: TDoubleMatrix; { K × NFeatures }
Inertia: Double; { within-cluster sum of squared distances }
Iters: Integer;
end;
{ PCA result }
TPCAResult = record
Components: TDoubleMatrix; { NComponents × NFeatures }
ExplainedVariance: TDoubleArray;
ExplainedRatio: TDoubleArray; { fraction of total variance }
Mean: TDoubleArray; { training mean, needed to transform new data }
Iterations: TIntegerArray; { power iterations used per component }
end;
{ DBSCAN result }
TDBSCANResult = record
Labels: TIntegerArray; { -1 = noise }
NClusters: Integer;
end;
{ Confusion matrix }
TConfusionMatrix = record
Counts: TDoubleMatrix; { NClasses × NClasses; [true][predicted] }
NClasses: Integer;
end;
---
Typical ML Workflow
Raw data
↓ TrainTestSplit (hold out 20%)
↓ Normalise / Standardise
↓ Train model (LinearRegression, KNN, NaiveBayes, ...)
↓ Predict on test set
↓ Evaluate (Accuracy, RMSE, R2Score, ConfusionMatrix, ...)
---
Preprocessing
Normalise
Xscaled := TMLKit.Normalise(X);
Scales each feature (column) to [0, 1]. Formula: (x − min) / (max − min). Columns where max = min are set to 0.
When to use: distance- and gradient-based models such as KNN and logistic regression benefit when feature ranges differ substantially.
Standardise
Xscaled := TMLKit.Standardise(X);
Centres each feature at zero mean and scales to unit variance. Formula: (x − μ) / σ. Columns with σ = 0 are set to 0.
When to use: PCA when feature scales should contribute comparably, logistic regression, and models that use distances or dot products. PCA centres input itself; standardisation is optional and changes covariance PCA into a scale-normalised analysis.
TrainTestSplit
TMLKit.TrainTestSplit(X, Y, TestFraction, Seed,
TrainX, TrainY, TestX, TestY);
Shuffles rows with a deterministic LCG (seed-reproducible) then splits. TestFraction = 0.2 → 20% test, 80% train.
OneHotEncode
M := TMLKit.OneHotEncode(Labels, NClasses);
Converts integer labels [0..NClasses-1] to a binary indicator matrix. Row i has a 1 at column Labels[i] and 0 elsewhere.
---
Regression
LinearRegression (OLS)
model := TMLKit.LinearRegression(X, Y);
// model.Coefficients — one coefficient per feature
// model.Intercept — bias term
// model.RSquared — training R²
Uses centered Householder QR rather than forming X'X, avoiding the condition- number squaring of normal equations. The intercept is fitted internally — do not add a bias column to X. At least as many samples as features are required; rank-deficient designs raise EMLError.
RidgeRegression
model := TMLKit.RidgeRegression(X, Y, Lambda);
L2-regularised OLS: β = (X'X + λI)⁻¹ X'y. The intercept is not regularised.
Lambda = 0→ same as OLSLambdalarge → coefficients shrink toward zero
When to use: correlated features, N < NFeatures.
PolynomialFeatures
Xpoly := TMLKit.PolynomialFeatures(X1D, Degree);
// Xpoly[i] = [1, x_i, x_i², ..., x_i^Degree]
// LinearRegression fits an intercept, so omit PolynomialFeatures' bias column:
Xpoly := TMLKit.PolynomialFeatures(X1D, Degree, False);
model := TMLKit.LinearRegression(Xpoly, Y);
Expands a 1-D feature vector to a polynomial design matrix. The two-argument compatibility overload includes a bias column. Pass IncludeBias = False when combining it with LinearRegression, which fits its own intercept.
LinearPredict
YHat := TMLKit.LinearPredict(model, Xnew);
Applies a fitted LinearRegression or RidgeRegression model to new data.
---
Classification
KNearestNeighbours
labels := TMLKit.KNearestNeighbours(TrainX, TrainY, TestX, K);
For each test point, finds the K closest training points by Euclidean distance and returns the majority class label.
Tips:
- Standardise features first (distance is not scale-invariant)
- K=1 overfits; K=sqrt(N) is a common starting point
- Use odd K to avoid ties in binary classification
NaiveBayes (Gaussian)
labels := TMLKit.NaiveBayes(TrainX, TrainY, TestX);
Estimates a Gaussian distribution per feature per class, then classifies by maximum log-posterior. Very fast and works well with many features.
When to use: text classification, spam filtering, or when features are (approximately) independent.
LogisticRegression
model := TMLKit.LogisticRegression(TrainX, TrainY);
model := TMLKit.LogisticRegression(TrainX, TrainY, LR, MaxIter, Tol);
labels := TMLKit.LogisticPredict(model, Xnew);
Binary (0/1) logistic regression trained with gradient descent. Sigmoid(Intercept + Dot(Coefficients, x)) ≥ 0.5 → class 1.
Parameters: LR learning rate (0.1), MaxIter (1000), Tol gradient norm (1e-5).
When to use: binary classification. LogisticPredict returns labels only; there is no public probability method, so callers needing scores must evaluate 1 / (1 + Exp(-(Intercept + dot(Coefficients, x)))) themselves. Probability calibration is not assessed by this implementation.
---
Clustering
KMeans
R := TMLKit.KMeans(X, K);
R := TMLKit.KMeans(X, K, MaxIter, Seed);
// R.Labels — cluster index 0..K-1 per sample
// R.Centroids — K centroid coordinates
// R.Inertia — total within-cluster squared distance
Lloyd's algorithm: alternate between assigning each point to its nearest centroid and recomputing centroids.
Choosing K: 1. Run KMeans for K = 1, 2, ..., 10 2. Plot Inertia vs K 3. Pick the "elbow" where improvement flattens
Tips:
- Run multiple times with different seeds; pick lowest Inertia
- Standardise features first
DBSCAN
R := TMLKit.DBSCAN(X, Eps, MinPts);
// R.Labels — cluster index or -1 (noise) per sample
// R.NClusters — number of clusters found
Finds arbitrarily-shaped clusters without specifying K. Points in low-density regions are labelled noise (-1).
Choosing parameters:
Eps: plot sorted k-distances (k = MinPts); pick the kneeMinPts: rule of thumb = 2 × NFeatures; larger = more conservative
---
Dimensionality Reduction
PCA
R := TMLKit.PCA(X, NComponents);
Xr := TMLKit.PCATransform(R, Xnew);
// R.Components[k] — k-th principal component (unit vector in feature space)
// R.ExplainedVariance — eigenvalue of each component
// R.ExplainedRatio[k] — fraction of total variance explained by component k
// R.Mean — training mean (subtracted before projection)
// R.Iterations[k] — iterations used for component k
Extracts the NComponents directions of maximum variance using power iteration + deflation (no LAPACK needed). Each component is re-orthogonalised, and convergence is sign-invariant. If power iteration exhausts its iteration limit, EMLError is raised instead of returning an unconverged component. Rank-deficient covariance matrices receive a deterministic orthogonal completion with zero explained variance; a completely zero-variance dataset is rejected.
Steps: 1. Standardise X before calling PCA 2. Check ExplainedRatio to decide how many components to keep 3. Use PCATransform to project both training and test data
When to use: visualisation (NComponents=2), noise reduction, or to decorrelate features before logistic regression.
---
Model Evaluation
| Function | Formula | Returns |
|---|---|---|
Accuracy |
correct / total | 0..1 |
Precision |
TP / (TP+FP) | 0..1 |
Recall |
TP / (TP+FN) | 0..1 |
F1Score |
2·P·R / (P+R) | 0..1 |
MSE |
mean (y−ŷ)² | ≥ 0 |
RMSE |
√MSE | ≥ 0, same units as y |
| `MAE` | mean |y−ŷ| | ≥ 0 |
| `R2Score` | 1 − SS_res/SS_tot | ≤ 1 |
// Classification
WriteLn(TMLKit.Accuracy(YTrue, YPred));
WriteLn(TMLKit.F1Score(YTrue, YPred, ClassLabel := 1));
CM := TMLKit.BuildConfusionMatrix(YTrue, YPred, NClasses);
// Regression
WriteLn(TMLKit.RMSE(YTrue, YPred));
WriteLn(TMLKit.R2Score(YTrue, YPred));
For a constant YTrue (zero total sum of squares), R2Score returns 1.0 regardless of predictions. Handle that degenerate case separately if your application needs a different convention.
---
Error Handling
EMLError is raised for:
- Empty, ragged, or zero-feature matrices
- NaN or infinite feature/target values
- Feature-count or sample/target length mismatches
- K > number of training samples (KNN, KMeans)
- NComponents > NFeatures (PCA)
- Negative or out-of-range labels, and non-binary logistic-regression labels
- too few samples or a rank-deficient design in
LinearRegression - a zero-variance PCA dataset or exhausted PCA power iteration
- Invalid fractions, iteration counts, tolerances, learning rates, radii, or regularisation parameters
PCATransform is an exception to the general empty-matrix rule: it returns nil for empty input.
---
Dependencies
MathBase.SharedTypes—TDoubleArray,TIntegerArray
No other external libraries required.
Common mistakes
- Fit standardisation on training data only.
FitStandardizationmust run on training rows; apply the returned model to training, validation, and test sets to avoid leakage. - Importance is not causality. Forest
FeatureImportancesis impurity based and can favour continuous or high-cardinality inputs; it is not a causal measure. - OOB needs enough trees.
OOBScoreuses only rows with an out-of-bag prediction; for a small tree count not every row is guaranteed one. - Dense in-memory baselines. The typed-analysis APIs are dense, serial, in-memory paths; multiclass LDA, boosting, and distributed training are outside 1.8.