hessboost 0.1.0

A faithful, fast, pure-Rust reimplementation of XGBoost gradient boosting
Documentation
hessboost-0.1.0 has been yanked.

hessboost

crates.io docs.rs CI license

A faithful, fast, pure-Rust reimplementation of XGBoost gradient boosting with no C/C++ dependency and no FFI.

hessboost re-implements XGBoost's algorithms from scratch in idiomatic Rust. It includes the regularized second-order boosting objective, exact, histogram, and approximate tree construction, the full objective and metric catalog, monotone and interaction constraints, categorical splits, DART and gblinear boosters, TreeSHAP, and numeric-tree XGBoost-format model interop with multi-core (rayon) acceleration.

Objective, metric, and parameter names mirror XGBoost, so configurations transfer directly.

Built with AI. The implementation was generated with Claude (Anthropic's AI coding assistant) under human direction and review. It is AI-generated code: it is covered by unit, property, and doc tests plus CI-checked XGBoost 3.4.1 parity, but it may still contain bugs, subtle numerical errors, or wrong edge-case behavior. Review and validate it for your own use case. It is provided as-is, without warranty (see LICENSE). Issue reports and fixes are welcome.

Using AI coding agents? See AGENTS.md for a task-oriented guide.

Quick start

use hessboost::prelude::*;

fn main() -> Result<()> {
    // Dense features (row-major) + labels.
    let x: Vec<f32> = /* n_rows * n_cols values */ vec![0.0; 400];
    let y: Vec<f32> = vec![0.0; 100];

    let dtrain = DMatrix::from_dense(&x, 100, 4)?.with_labels(&y)?;

    let params = TrainingParams::builder()
        .objective("reg:squarederror")
        .tree_method(TreeMethod::Hist)   // fast histogram method
        .max_depth(6)
        .eta(0.1)
        .subsample(0.9)
        .lambda(1.0)
        .build()?;

    let model = train(&params, &dtrain, 200)?;
    let preds = model.predict(&dtrain)?;

    model.save_binary("model.bin")?;
    Ok(())
}

Examples

Runnable, self-contained examples live in examples/. Run any with cargo run --release --example <name>:

Example Shows
binary_classification binary:logistic, watched eval set, early stopping, AUC
multiclass multi:softprob, per-class probabilities, predict_class
ranking LambdaMART rank:ndcg over query groups
shap predict_contribs and predict_interactions (TreeSHAP)
model_io native binary / JSON and XGBoost-format model save & load
custom_objective custom loss and custom eval-metric hooks
constraints monotone + interaction constraints and categorical features
train_regression end-to-end regression with feature importance

Feature status

Implemented & tested

  • Boosters: gbtree, dart (tree dropout), and gblinear (linear model via coordinate descent).
  • Trees: tree_method = exact | hist | approx (approx uses hessian-weighted per-round binning), grow_policy = depthwise | lossguide, histogram binning with the parent−child subtraction trick, sparsity-aware missing-value handling, row/column subsampling (bytree/bylevel/bynode).
  • Regularization: lambda, alpha, gamma, min_child_weight, max_delta_step, max_depth, max_leaves, max_bin.
  • Objectives: reg:squarederror (alias reg:linear), reg:logistic, reg:pseudohubererror (huber_slope), binary:logistic, multi:softmax, multi:softprob, count:poisson, reg:gamma, reg:tweedie (tweedie_variance_power), learning-to-rank (rank:pairwise, rank:ndcg, rank:map, LambdaMART with lambdarank_num_pair_per_sample), and a user custom-objective hook. Intercepts are estimated per output exactly as XGBoost 3.4.1 does (label mean, class log-frequencies, or a Newton step).
  • Metrics: rmse, mae, logloss, error, auc, aucpr, mlogloss, merror, poisson/gamma/tweedie-nloglik, ndcg, map (with @k), and a custom-metric hook. Default metrics follow XGBoost (ndcg@k/map@k for ranking, tweedie-nloglik@rho for Tweedie), except reg:pseudohubererror, which reports mae because XGBoost's mphe is not implemented.
  • Constraints: monotone constraints and interaction constraints, supported in both the hist and exact builders.
  • Modeling: native categorical splits (hist and exact), per-instance base_margin (warm-start), TreeSHAP contributions (predict_contribs) and interaction values (predict_interactions), early stopping, feature importance (weight / gain / cover / totals), leaf-index and margin prediction.
  • Ecosystem: libsvm & CSV loaders, native binary + JSON model I/O, XGBoost-format JSON model import/export for numeric gbtree ensembles, k-fold cross-validation, multi-core histogram construction, and runtime-detected SIMD kernels: AArch64 NEON for objective and metric kernels, prediction transforms, multiclass operations, and histogram split evaluation. x86-64 AVX2/FMA for objective gradients, prediction transforms, and histogram split evaluation.

Not implemented: the UBJSON (binary) XGBoost model format, a GPU backend, distributed or external-memory training, and Python/CLI/C-ABI wrappers.

Performance

AArch64 builds use runtime-detected NEON kernels for objective gradients, probability transforms, and metric reductions. x86-64 builds use AVX2+FMA for exponential/sigmoid transforms, logistic and short-softmax gradients, and SSE2 for quantile bin search. Scalar fallbacks cover other CPUs, short inputs, and values outside the approximation ranges. Split search is scalar and follows XGBoost's f32 gain arithmetic exactly. Histogram training parallelizes data preparation and independent nodes, scales histogram tasks to node size, and reuses training-row partitions when that reduces prediction work. Leaves at max_depth skip histograms and split searches.

Compared with XGBoost

Measured on Apple M3 Max against XGBoost 3.4.1, the latest stable PyPI release checked on 2026-09-14 UTC. Both engines use the same dense f32 data and CPU hist parameters: 100 boosting rounds, depth 6, 256 bins, eta=0.1, and lambda=1. Times include fresh training-matrix preparation and training, and report the median of six fits after warmup.

Workload Threads hessboost XGBoost 3.4.1
Regression, 100k × 30 1 0.458 s 1.054 s
Regression, 100k × 30 4 0.201 s 0.366 s
Regression, 100k × 30 16 0.264 s 0.362 s
Regression, 50k × 128 1 1.153 s 3.200 s
Regression, 50k × 128 4 0.452 s 0.994 s
Regression, 50k × 128 16 0.428 s 0.669 s
Binary, 100k × 30 1 0.459 s 1.044 s
Binary, 100k × 30 4 0.197 s 0.366 s
Binary, 100k × 30 16 0.256 s 0.361 s
4-class, 50k × 30 1 1.094 s 2.506 s
4-class, 50k × 30 4 0.523 s 0.981 s
4-class, 50k × 30 16 0.746 s 1.219 s

hessboost has lower median fit time in all 12 configurations in this run. Single-thread speedups are 2.28× to 2.77×, four-thread speedups are 1.82× to 2.20×, and sixteen-thread speedups are 1.37× to 1.63×.

See Performance for held-out quality, workload definitions, kernel benchmarks, and reproduction commands.

Testing & parity

cargo test                          # unit + integration tests
cargo clippy --all-targets -- -D warnings

Numerical parity with XGBoost 3.4.1 is checked in CI by a fixture harness (scripts/gen_fixtures.py, tests/parity.rs, scripts/check_exports.py). Each case is checked three ways: train parity (same data and parameters, compare predictions), import parity (from_xgboost_json on the XGBoost model: predictions, margins, SHAP contributions) and export parity (to_xgboost_json reloaded by XGBoost). exact-tier cases — including every objective and deterministic tree method, plus categorical splits — must agree pointwise (1e-4 / 1e-5). quality cases (row/column subsampling and DART) use an RMSE band because their RNG streams differ. The train-only gblinear case is pointwise; gblinear XGBoost-JSON import/export remains unsupported and is asserted explicitly. Histogram and approximate quantile cuts are compared bit-for-bit.

uv run --with xgboost==3.4.1 --with numpy python scripts/gen_fixtures.py
cargo test --test parity --release -- --ignored --nocapture
uv run --with xgboost==3.4.1 --with numpy python scripts/check_exports.py

See scripts/README.md for the case matrix and tolerances.

License and attribution

Licensed under the Apache License, Version 2.0; see LICENSE. Copyright 2026 Brenden Matthews.

hessboost is a fork of sequoia-boost (Copyright 2026 Patrick Garrett, Apache-2.0).

It is an independent reimplementation of XGBoost (Copyright the XGBoost Contributors, Apache-2.0) built from XGBoost's public descriptions and papers; it contains no XGBoost source code. "XGBoost" is used descriptively, for algorithmic lineage and result compatibility. This project is not affiliated with or endorsed by the XGBoost project.