Boosting Libraries: XGBoost, LightGBM, CatBoost
If your data is a table, a gradient-boosted tree ensemble is the model to beat. It is a strong baseline, not a universal winner. Rankings depend on dataset size, feature types, pretrained representations, tuning budget, hardware, and metric. Compare complete pipelines under the same evaluation protocol.
This page explains why the algorithm works, then what the three major implementations actually do differently.
Gradient boosting from first principles
Boosting builds an additive model, one weak learner at a time, where each new learner fits the errors of the ensemble so far.
The insight that turns this into a general algorithm: fit \(h_m\) to the negative gradient of the loss with respect to the current predictions.
For squared error, \(r_{im} = y_i - F_{m-1}(x_i)\) — literally the residual. For log-loss, it is \(y_i - p_i\). This is gradient descent in function space: each tree is a step in the direction that most reduces the loss, and \(\eta\) is the learning rate.
Ctrl/Cmd + wheel to zoom · drag to pan · double-click to fit · ⛶ full size
XGBoost's second-order objective
XGBoost's original contribution was to Taylor-expand the loss to second order and add explicit regularisation:
with \(g_i\) the first derivative, \(h_i\) the second, and \(T\) the number of leaves. Minimising over leaf weights gives a closed form:
and the gain of a candidate split is the improvement in \(\mathcal{L}^\star\):
Three things fall out of this formula and are worth reading off it directly:
- \(\lambda\) shrinks leaf values toward zero — it is L2 regularisation on the leaf weights.
- \(\gamma\) is a minimum gain threshold. A split whose improvement is below \(\gamma\) is not made, which is pre-pruning by cost.
- Second-order information adapts the step size per leaf. \(H\) in the denominator shrinks a fixed-gradient update when curvature is larger. For logistic loss, \(h=p(1-p)\) is largest near 0.5, not near confident 0/1 predictions. This does not eliminate learning-rate tuning or protect against tiny Hessians.
Boosting vs bagging
| Bagging (Random Forest) | Boosting (GBM family) | |
|---|---|---|
| Trees are | independent, parallel | sequential, each fits the last one's errors |
| Base learners | deep, low bias, high variance | shallow, high bias, low variance |
| Reduces | variance | bias (and variance, via shrinkage) |
| Overfits with more trees? | no, it plateaus | yes — needs early stopping |
| Parallelism | across trees | within a tree (split finding) only |
| Tuning sensitivity | low | moderate to high |
| Typical accuracy on tabular | good | usually better |
That "overfits with more trees" row is the practical difference. A random forest with 5,000 trees is fine; a booster with 5,000 trees and no early stopping is usually worse than the same model stopped at 400.
The three libraries
XGBoost — the exact-ish one
The original scalable implementation. Its defining features:
- Second-order objective with explicit \(\lambda\), \(\alpha\), and \(\gamma\) regularisation.
- Level-wise (depth-wise) tree growth by default: grow every node at a depth
before descending. Balanced trees, controlled by
max_depth. - Sparsity-aware split finding: missing values get a learned default direction at each split rather than being imputed.
- Weighted quantile sketch for approximate split finding on large data.
- Mature GPU support (
device="cuda",tree_method="hist"), distributed training, and integrations everywhere.
import xgboost as xgb
clf = xgb.XGBClassifier(
n_estimators=5000, learning_rate=0.03,
max_depth=6, min_child_weight=5,
subsample=0.8, colsample_bytree=0.8,
reg_lambda=1.0, reg_alpha=0.0, gamma=0.0,
tree_method="hist", device="cuda",
eval_metric="aucpr", early_stopping_rounds=100,
enable_categorical=True,
)
clf.fit(X_train, y_train, eval_set=[(X_val, y_val)], verbose=100)
print(clf.best_iteration, clf.best_score)LightGBM — the fast one
Microsoft's implementation, built around two ideas plus a different growth strategy.
- Leaf-wise growth: always split the leaf with the highest gain, anywhere in
the tree. This reaches a lower loss for the same number of leaves, but produces
deep, unbalanced trees that overfit small datasets. Control with
num_leaves,min_data_in_leaf, and the supportedmax_depthbound. - GOSS (Gradient-based One-Side Sampling): keep all large-gradient examples and randomly sample the small-gradient ones, reweighting to stay unbiased. This is an optional sampling strategy, not active in the displayed ordinary bagging configuration; subsampling adds estimation variance, not identical information.
- EFB (Exclusive Feature Bundling): bundle mutually exclusive sparse features (they are rarely non-zero simultaneously) into single features, shrinking the effective feature count on one-hot-heavy data.
- Histogram binning of continuous features into 255 buckets, which turns split finding from a sort into a histogram scan.
import lightgbm as lgb
clf = lgb.LGBMClassifier(
n_estimators=5000, learning_rate=0.03,
num_leaves=63, min_child_samples=40,
subsample=0.8, subsample_freq=1, colsample_bytree=0.8,
reg_lambda=1.0, max_bin=255,
objective="binary", metric="average_precision",
)
clf.fit(X_train, y_train, eval_set=[(X_val, y_val)],
categorical_feature=cat_cols,
callbacks=[lgb.early_stopping(100), lgb.log_evaluation(100)])The num_leaves trap. A depth-\(d\) level-wise tree has \(2^d\) leaves. Setting
num_leaves=1024 with the intent of "depth 10" gives LightGBM licence to build
a wildly unbalanced tree that memorises the training set. The usual guidance is
num_leaves < 2^max_depth, and in practice 31–127 covers most problems.
CatBoost — the categorical one
Yandex's implementation, designed around two specific biases in the standard algorithm.
- Ordered target statistics. Naive target encoding uses a category's mean target computed from all rows including the current one, which leaks. CatBoost processes examples in a random permutation and computes each row's encoding using preceding rows in that permutation. This reduces self-target leakage but is not a chronological out-of-time estimate or a universal unbiasedness claim.
- Ordered boosting. The same leak exists in gradient estimation: the residual for a training row is computed from a model that was fit on that row. CatBoost maintains models trained on prefixes of a permutation to remove this "prediction shift".
- Oblivious (symmetric) trees. Every node at a given depth uses the same split. That is a strong regulariser and makes inference extremely fast — a tree becomes a bit-index into a lookup table. Whether this gives the lowest CPU latency depends on model size, batch size, implementation and hardware.
- Native categorical and text features, plus automatic combinations of categorical features.
from catboost import CatBoostClassifier, Pool
train_pool = Pool(X_train, y_train, cat_features=cat_cols, text_features=text_cols)
val_pool = Pool(X_val, y_val, cat_features=cat_cols, text_features=text_cols)
clf = CatBoostClassifier(
iterations=5000, learning_rate=0.03, depth=6,
l2_leaf_reg=3.0, loss_function="Logloss", eval_metric="PRAUC",
early_stopping_rounds=100, task_type="GPU", verbose=200,
)
clf.fit(train_pool, eval_set=val_pool)Side by side
| XGBoost | LightGBM | CatBoost | |
|---|---|---|---|
| Tree growth | level-wise (leaf-wise available) | leaf-wise | oblivious/symmetric |
| Main complexity knob | max_depth |
num_leaves |
depth |
| Speed on large data | benchmark on target hardware | benchmark on target hardware | benchmark on target hardware |
| Small-data robustness | tune regularization | constrain leaf growth | test ordered/symmetric-tree choices |
| Categorical handling | one-hot or enable_categorical |
integer codes, native | ordered target statistics |
| Missing values | learned default direction | native | native |
| Default hyperparameters | need tuning | need tuning | often good as-is |
| Inference latency | model/batch-dependent | model/batch-dependent | symmetric trees can help |
| Text features | no | no | yes |
| GPU training | mature | mature | mature |
| Overfitting protection | \(\gamma\), \(\lambda\), min_child_weight |
min_data_in_leaf, num_leaves |
ordered boosting, symmetric trees |
A reasonable default policy: LightGBM when rows exceed ~1M or you are iterating quickly; CatBoost when categoricals dominate or data is small; XGBoost when you want the most predictable, best-documented behaviour or you are matching an existing production model. On most problems all three land within noise of each other after tuning; this must be measured, and averaging does not guarantee improvement. The row-count suggestions are pilot choices, not cutoffs.
Hyperparameters that matter, in order
| Rank | Parameter | Effect | Sensible range |
|---|---|---|---|
| 1 | n_estimators + early stopping |
the number of trees; let early stopping choose it | 5000 cap, early_stopping_rounds=100 |
| 2 | learning_rate |
step size; lower needs more trees | 0.01–0.1 |
| 3 | max_depth / num_leaves |
model capacity | depth 3–10; leaves 15–255 |
| 4 | min_child_weight / min_data_in_leaf |
minimum evidence per leaf | 1–100; raise on noisy data |
| 5 | subsample (row sampling) |
stochastic boosting, decorrelates trees | 0.6–1.0 |
| 6 | colsample_bytree (feature sampling) |
decorrelates trees, speeds training | 0.4–1.0 |
| 7 | reg_lambda (L2) |
shrinks leaf values | 0–10, log scale |
| 8 | reg_alpha (L1) |
sparsifies leaf values | 0–10, log scale |
| 9 | gamma / min_split_gain |
minimum gain to split | 0–5 |
| 10 | scale_pos_weight |
class imbalance | \(n_{neg}/n_{pos}\) |
The learning-rate/trees trade-off is the core one. Halving the learning rate
roughly doubles the number of trees needed and usually improves final accuracy a
little. A practical workflow: tune structure at learning_rate=0.1 for speed,
then drop to 0.02–0.03 with early stopping for the final model.
A tuning recipe that converges
import optuna
import numpy as np
import lightgbm as lgb
from sklearn.model_selection import StratifiedKFold
from sklearn.metrics import average_precision_score
def objective(trial):
params = dict(
learning_rate = 0.03,
num_leaves = trial.suggest_int("num_leaves", 15, 255, log=True),
min_child_samples= trial.suggest_int("min_child_samples", 5, 300, log=True),
subsample = trial.suggest_float("subsample", 0.5, 1.0),
subsample_freq = 1,
colsample_bytree = trial.suggest_float("colsample_bytree", 0.3, 1.0),
reg_lambda = trial.suggest_float("reg_lambda", 1e-3, 30.0, log=True),
reg_alpha = trial.suggest_float("reg_alpha", 1e-3, 30.0, log=True),
n_estimators = 5000,
)
scores = []
for fold, (tr, va) in enumerate(StratifiedKFold(5, shuffle=True, random_state=0).split(X, y)):
m = lgb.LGBMClassifier(**params, metric="average_precision", random_state=0)
m.fit(X.iloc[tr], y[tr], eval_set=[(X.iloc[va], y[va])],
callbacks=[lgb.early_stopping(100, verbose=False)])
scores.append(average_precision_score(y[va], m.predict_proba(X.iloc[va])[:, 1]))
trial.report(float(np.mean(scores)), step=fold)
if trial.should_prune():
raise optuna.TrialPruned()
return np.mean(scores)
study = optuna.create_study(direction="maximize",
pruner=optuna.pruners.MedianPruner())
study.optimize(objective, n_trials=100)This optional fragment assumes a numeric/categorical-compatible DataFrame X
and NumPy labels y; LightGBM and Optuna are additional dependencies. Pruning
reports completed folds only, with no iteration callback sharing step numbers.
These validation folds also select stopping iterations, so this remains a
development objective. Evaluate the whole search on outer folds or an untouched
test set, and predeclare the final refit duration. Fixed learning rate narrows the
search; it does not itself make comparisons statistically valid.
Categorical features
Native numeric missing-value support does not remove schema preparation. CatBoost categorical inputs need consistent strings or integers; convert missing categories to an explicit string sentinel rather than passing floating NaN. Preserve train/validation vocabulary metadata for categorical XGBoost/LightGBM columns, and test unseen values, column names and ordering. A random ordered-statistics permutation cannot substitute for a temporal evaluation cutoff. See the CatBoost FAQ and LightGBM parameters.
| Approach | Cardinality | Note |
|---|---|---|
| One-hot | low (< ~15) | explodes dimension; poor for trees at high cardinality |
| Ordinal/label encoding | any | imposes a false order, but trees can partly recover with enough splits |
| Native categorical (LightGBM/XGBoost) | medium | finds an optimal partition of levels per split via a sorted-by-gradient heuristic |
| Target encoding | high | powerful and dangerous — must be cross-fitted |
| CatBoost ordered statistics | high | target encoding done correctly by construction |
| Hashing | very high | fixed width, collisions, no fit needed |
| Learned embeddings | very high | requires a neural model |
Why one-hot is bad for trees at high cardinality: each one-hot column can only produce a "this level vs everything else" split, so isolating a group of \(k\) levels needs \(k\) levels of depth. Native categorical support instead sorts levels by their gradient statistics and finds the best contiguous partition in one split — a far better use of tree capacity.
Target encoding leakage is the single most common way to get a wonderful CV score and a broken model. If a category appears once, its target mean is its target. Always cross-fit (encode each fold using only the other folds), and add smoothing toward the global mean:
Imbalanced data
XGBClassifier(scale_pos_weight=neg/pos) # reweights the positive gradient
LGBMClassifier(is_unbalance=True) # or class_weight="balanced"
CatBoostClassifier(auto_class_weights="Balanced")Class weighting changes what the model optimises, but it also decalibrates the
probabilities — a reweighted model no longer predicts the true base rate. If
you need calibrated probabilities, train unweighted, then tune the decision
threshold on a validation set, and check a reliability diagram. Optimising
aucpr/average_precision as the early-stopping metric matters more than the
weighting in most cases.
Interpretation
import shap
explainer = shap.TreeExplainer(model)
sv = explainer(X_val)
shap.summary_plot(sv, X_val) # global: which features, which direction
shap.plots.waterfall(sv[0]) # local: this prediction, explained
shap.plots.scatter(sv[:, "age"], color=sv) # dependence-colored scatter, Explanation APITreeSHAP is exact and fast for tree ensembles — polynomial rather than exponential in features — which is why SHAP became the default explanation tool in tabular ML. Its values are additive: each prediction decomposes as the base value plus one contribution per feature.
Built-in importances are much weaker and should be read with care:
| Importance type | What it measures | Bias |
|---|---|---|
weight / split |
number of times a feature is split on | favours high-cardinality features |
gain |
total loss reduction from its splits | better, still training-set-based |
cover |
number of examples affected | rarely what you want |
| permutation | held-out degradation when shuffled | honest, but misleading under correlation |
| SHAP | per-example additive attribution | the best default; still not causal |
None of these are causal. A feature can be important because it is a proxy for something else, or because it leaks the target. High importance on an unexpected feature is a signal to investigate leakage, not to celebrate.
Why trees still beat deep learning on tabular data
Reasons axis-aligned trees can be strong baselines, to verify on your data:
- Coordinate-specific structure. Axis-aligned splits exploit individually meaningful columns. Neural training procedures are not universally rotationally invariant; architecture, initialization, regularization, and optimizer matter.
- Uninformative features. Split selection can ignore many weak columns, but finite-sample chance correlations can still select noise. Compare against shuffled features and held-out performance rather than assuming immunity.
- Threshold-like targets. Trees represent piecewise-constant partitions naturally. Neural networks can also approximate sharp changes; generalization depends on data, capacity, and training, not a blanket inability.
- Less need for numerical scaling. Exact axis-aligned training partitions are invariant to strictly monotone transforms, with histogram/rounding caveats. Schema, categorical missingness, leakage and feature availability still matter.
- Different tuning needs. Boosters can be strong with modest search, but no universal percentage gap or unusable-MLP claim applies across datasets.
Deep tabular models (TabNet, FT-Transformer, SAINT, TabPFN) close the gap in specific regimes — very large datasets, heavy multi-modality, transfer across related tables, or very small datasets in TabPFN's case — but the default answer for a table with 10k–10M rows remains a boosted ensemble.
Production notes
model.save_model("model.json") # XGBoost: JSON/UBJ, version-portable
model.booster_.save_model("model.txt") # LightGBM: plain text
model.save_model("model.cbm") # CatBoost: native binary- Save the native format, not a pickle. Native formats survive library upgrades; pickles frequently do not.
- Pin the library version with the artefact anyway; split thresholds and histogram binning can change between versions.
- Record
best_iterationand prediction semantics. XGBoost's sklearn wrapper automatically uses the best iteration after early stopping. Native Booster prediction uses its own defaults; specifyiteration_range=(0, best_iteration+1)when required.ntree_limitis a legacy API, not the current prescription. - Freeze the feature order and names. All three libraries index by position internally; a reordered frame silently produces garbage.
- For latency, benchmark symmetric trees and alternatives; optionally compile
the ensemble with Treelite/
lleavesfor a large single-row speedup. - Monitor feature drift, not just prediction drift. Trees extrapolate terribly — a feature moving outside its training range gets clipped to the outermost leaf, which fails silently rather than loudly.
Self-check
Worked Newton split and CPU fit
At initial binary probabilities 0.5, labels \((0,0,1,1)\) give gradients \((0.5,0.5,-0.5,-0.5)\) and Hessians all 0.25. A split separating the labels has \(G_L=1,G_R=-1,H_L=H_R=0.5\). With \(\lambda=1,\gamma=0.1\), gain is \(\tfrac12(1/1.5+1/1.5)-0.1\approx0.567\) and leaf values are \(-2/3,+2/3\) before learning-rate shrinkage. Raising \(\gamma\) above \(2/3\) rejects this split.
import numpy as np
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from sklearn.metrics import roc_auc_score
from xgboost import XGBClassifier
g = np.array([.5, .5, -.5, -.5])
h = np.full(4, .25)
gain = .5*(g[:2].sum()**2/(h[:2].sum()+1)+g[2:].sum()**2/(h[2:].sum()+1))-.1
assert np.isclose(gain, 17/30)
X, y = make_classification(n_samples=240, n_features=6, n_informative=4, random_state=8)
train, valid, yt, yv = train_test_split(X, y, test_size=.25, stratify=y, random_state=2)
model = XGBClassifier(n_estimators=60, max_depth=3, learning_rate=.1,
tree_method="hist", device="cpu", n_jobs=1, random_state=0,
eval_metric="logloss", early_stopping_rounds=5)
model.fit(train, yt, eval_set=[(valid, yv)], verbose=False)
automatic = model.predict_proba(valid)
explicit = model.predict_proba(valid, iteration_range=(0, model.best_iteration+1))
np.testing.assert_allclose(automatic, explicit)
assert roc_auc_score(yv, automatic[:, 1]) > .75
print("gain", gain, "best iteration", model.best_iteration, "wrapper parity passed")This validation set selects stopping and is not an untouched final quality estimate. The fixture tests CPU XGBoost, not GPU kernels or the optional LightGBM/CatBoost snippets. For custom objectives, state whether inputs are raw margins or transformed probabilities and supply compatible gradient/Hessian shapes. Count, quantile, multiclass and ranking losses have different output/group contracts; do not reuse a binary metric blindly. The prediction guide documents best-iteration behavior, and the SHAP scatter API documents the corrected plotting call.
- Write the gradient boosting update, and say what \(h_m\) is fit to.
- Read \(\gamma\) and \(\lambda\) off the XGBoost gain formula and say what each does.
- Why does
num_leaves=1024overfit LightGBM even with a small learning rate? - Explain the leak that CatBoost's ordered target statistics prevent.
- Your booster's validation score improves for 300 trees then degrades. What is happening, and what is the fix?
- When is one-hot encoding actively harmful for a tree model?
- Give three reasons boosted trees beat MLPs on tabular data.
Where to go next
- Scikit-learn —
HistGradientBoosting, pipelines, and leak-free cross-validation. - Pandas — building the table these models consume.
- Machine Learning notes — trees, ensembles, and the bias–variance reasoning behind boosting.