hessboost
A faithful, fast, pure-Rust reimplementation of XGBoost gradient boosting with no C/C++ dependency and no FFI.
hessboost re-implements XGBoost's algorithms from scratch in idiomatic
Rust. It includes the regularized second-order boosting objective, exact,
histogram, and approximate tree construction, the full objective and metric
catalog, monotone and interaction constraints, categorical splits, DART and
gblinear boosters, TreeSHAP, and numeric-tree XGBoost-format model interop with
multi-core (rayon) acceleration.
Objective, metric, and parameter names mirror XGBoost, so configurations transfer directly.
Built with AI. The implementation was generated with Claude (Anthropic's AI coding assistant) under human direction and review. It is AI-generated code: it is covered by unit, property, and doc tests plus CI-checked XGBoost 3.4.1 parity, but it may still contain bugs, subtle numerical errors, or wrong edge-case behavior. Review and validate it for your own use case. It is provided as-is, without warranty (see LICENSE). Issue reports and fixes are welcome.
Using AI coding agents? See
AGENTS.mdfor a task-oriented guide.
Quick start
use *;
Examples
Runnable, self-contained examples live in
examples/. Run any with
cargo run --release --example <name>:
| Example | Shows |
|---|---|
binary_classification |
binary:logistic, watched eval set, early stopping, AUC |
multiclass |
multi:softprob, per-class probabilities, predict_class |
ranking |
LambdaMART rank:ndcg over query groups |
shap |
predict_contribs and predict_interactions (TreeSHAP) |
model_io |
native binary / JSON and XGBoost-format model save & load |
custom_objective |
custom loss and custom eval-metric hooks |
constraints |
monotone + interaction constraints and categorical features |
train_regression |
end-to-end regression with feature importance |
Feature status
Implemented & tested
- Boosters:
gbtree,dart(tree dropout), andgblinear(linear model via coordinate descent). - Trees:
tree_method = exact | hist | approx(approx uses hessian-weighted per-round binning),grow_policy = depthwise | lossguide, histogram binning with the parent−child subtraction trick, sparsity-aware missing-value handling, row/column subsampling (bytree/bylevel/bynode). - Regularization:
lambda,alpha,gamma,min_child_weight,max_delta_step,max_depth,max_leaves,max_bin. - Objectives:
reg:squarederror(aliasreg:linear),reg:logistic,reg:pseudohubererror(huber_slope),binary:logistic,multi:softmax,multi:softprob,count:poisson,reg:gamma,reg:tweedie(tweedie_variance_power), learning-to-rank (rank:pairwise,rank:ndcg,rank:map, LambdaMART withlambdarank_num_pair_per_sample), and a user custom-objective hook. Intercepts are estimated per output exactly as XGBoost 3.4.1 does (label mean, class log-frequencies, or a Newton step). - Metrics:
rmse,mae,logloss,error,auc,aucpr,mlogloss,merror,poisson/gamma/tweedie-nloglik,ndcg,map(with@k), and a custom-metric hook. Default metrics follow XGBoost (ndcg@k/map@kfor ranking,tweedie-nloglik@rhofor Tweedie), exceptreg:pseudohubererror, which reportsmaebecause XGBoost'smpheis not implemented. - Constraints: monotone constraints and interaction constraints,
supported in both the
histandexactbuilders. - Modeling: native categorical splits (hist and exact), per-instance
base_margin(warm-start), TreeSHAP contributions (predict_contribs) and interaction values (predict_interactions), early stopping, feature importance (weight / gain / cover / totals), leaf-index and margin prediction. - Ecosystem: libsvm & CSV loaders, native binary + JSON model I/O,
XGBoost-format JSON model import/export for numeric
gbtreeensembles, k-fold cross-validation, multi-core histogram construction, and runtime-detected SIMD kernels: AArch64 NEON for objective and metric kernels, prediction transforms, multiclass operations, and histogram split evaluation. x86-64 AVX2/FMA for objective gradients, prediction transforms, and histogram split evaluation.
Not implemented: the UBJSON (binary) XGBoost model format, a GPU backend, distributed or external-memory training, and Python/CLI/C-ABI wrappers.
Performance
AArch64 builds use runtime-detected NEON kernels for objective gradients,
probability transforms, and metric reductions. x86-64 builds use AVX2+FMA for
exponential/sigmoid transforms, logistic and short-softmax gradients, and SSE2
for quantile bin search. Scalar fallbacks cover other CPUs, short inputs, and
values outside the approximation ranges. Split search is scalar and follows
XGBoost's f32 gain arithmetic exactly. Histogram training parallelizes data
preparation and independent nodes, scales histogram tasks to node size, and
reuses training-row partitions when that reduces prediction work. Leaves at
max_depth skip histograms and split searches.
Compared with XGBoost
Measured on Apple M3 Max against XGBoost 3.4.1,
the latest stable PyPI release checked on 2026-09-14 UTC. Both engines use the
same dense f32 data and CPU hist parameters: 100 boosting rounds, depth 6,
256 bins, eta=0.1, and lambda=1. Times include fresh training-matrix
preparation and training, and report the median of six fits after warmup.
| Workload | Threads | hessboost | XGBoost 3.4.1 |
|---|---|---|---|
| Regression, 100k × 30 | 1 | 0.458 s | 1.054 s |
| Regression, 100k × 30 | 4 | 0.201 s | 0.366 s |
| Regression, 100k × 30 | 16 | 0.264 s | 0.362 s |
| Regression, 50k × 128 | 1 | 1.153 s | 3.200 s |
| Regression, 50k × 128 | 4 | 0.452 s | 0.994 s |
| Regression, 50k × 128 | 16 | 0.428 s | 0.669 s |
| Binary, 100k × 30 | 1 | 0.459 s | 1.044 s |
| Binary, 100k × 30 | 4 | 0.197 s | 0.366 s |
| Binary, 100k × 30 | 16 | 0.256 s | 0.361 s |
| 4-class, 50k × 30 | 1 | 1.094 s | 2.506 s |
| 4-class, 50k × 30 | 4 | 0.523 s | 0.981 s |
| 4-class, 50k × 30 | 16 | 0.746 s | 1.219 s |
hessboost has lower median fit time in all 12 configurations in this run. Single-thread speedups are 2.28× to 2.77×, four-thread speedups are 1.82× to 2.20×, and sixteen-thread speedups are 1.37× to 1.63×.
See Performance for held-out quality, workload definitions, kernel benchmarks, and reproduction commands.
Testing & parity
Numerical parity with XGBoost 3.4.1 is checked in CI by a fixture harness
(scripts/gen_fixtures.py, tests/parity.rs,
scripts/check_exports.py). Each case is checked three ways: train parity
(same data and parameters, compare predictions), import parity
(from_xgboost_json on the XGBoost model: predictions, margins, SHAP
contributions) and export parity (to_xgboost_json reloaded by XGBoost).
exact-tier cases — including every objective and deterministic tree method,
plus categorical splits — must agree pointwise (1e-4 / 1e-5). quality cases
(row/column subsampling and DART) use an RMSE band because their RNG streams
differ. The train-only gblinear case is pointwise; gblinear XGBoost-JSON
import/export remains unsupported and is asserted explicitly. Histogram and
approximate quantile cuts are compared bit-for-bit.
See scripts/README.md for the case matrix and tolerances.
License and attribution
Licensed under the Apache License, Version 2.0; see LICENSE.
Copyright 2026 Brenden Matthews.
hessboost is a fork of sequoia-boost (Copyright 2026 Patrick Garrett, Apache-2.0).
It is an independent reimplementation of XGBoost (Copyright the XGBoost Contributors, Apache-2.0) built from XGBoost's public descriptions and papers; it contains no XGBoost source code. "XGBoost" is used descriptively, for algorithmic lineage and result compatibility. This project is not affiliated with or endorsed by the XGBoost project.