TenfloweRS
A pure Rust implementation of TensorFlow, providing a comprehensive deep learning framework with Rust's safety and performance guarantees.
Overview
TenfloweRS is the main convenience crate that re-exports all TenfloweRS subcrates, providing a unified API for deep learning in Rust. Built on the robust SciRS2 ecosystem, it offers:
- Production-Ready: Full-featured neural networks, training, and deployment
- High Performance: GPU acceleration, SIMD optimization, mixed precision
- Type Safety: Rust's type system prevents common ML bugs at compile time
- Cross-Platform: CPU, GPU (CUDA, Metal, Vulkan), and WebGPU support
- Ecosystem Integration: Seamless integration with SciRS2, NumRS2, and OptiRS
Quick Start
Add TenfloweRS to your Cargo.toml:
[]
= "0.1.2"
Basic Example
use *;
Build a Neural Network
use *;
Train a Model
use *;
Features
TenfloweRS provides several optional features:
Default Features
std: Standard library supportparallel: Parallel execution via Rayon
GPU Acceleration
gpu: GPU acceleration via WGPU (Metal, Vulkan, DirectX, WebGPU)cuda: CUDA support (Linux/Windows only)cudnn: cuDNN support (requires CUDA)opencl: OpenCL supportmetal: Metal support (macOS only)rocm: ROCm support (AMD GPUs)nccl: NCCL for distributed GPU training
BLAS Acceleration
blas: Generic BLAS supportblas-oxiblas: OxiBLAS acceleration (pure Rust)blas-accelerate: Apple Accelerate framework (macOS only)
Performance and Optimization
simd: SIMD vectorization optimizations
Serialization and I/O
serialize: Serialization support (JSON, MessagePack)compression: Compression support for checkpointsonnx: ONNX model import/export
Platform Support
wasm: WebAssembly support
Development
autograd: Automatic differentiation supportbenchmark: Benchmarking utilities
Language Bindings
python: Python bindings via PyO3 (requires Python environment)
Presets
minimal: Onlystd(smallest possible build)standard:std+parallel(same as the default features)full: Enable most features (gpu, blas-oxiblas, simd, serialize, compression, onnx, autograd)
Experimental
experimental: Opt-in to preview APIs not covered by stability guarantees
Enable GPU Support
[]
= { = "0.1.2", = ["gpu"] }
Enable All Features
[]
= { = "0.1.2", = ["full"] }
Architecture
TenfloweRS is organized into focused subcrates:
- tenflowers-core: Core tensor operations and device management (1,171 tests)
- tenflowers-autograd: Automatic differentiation engine (521 tests)
- tenflowers-neural: Neural network layers, models, and 150+ ML domains (11,596 tests)
- tenflowers-dataset: Data loading and preprocessing (660 tests)
- tenflowers-ffi: Python and C bindings (185 tests)
This meta crate re-exports all public APIs for convenience, including the tensor! macro and prelude module.
SciRS2 Integration
TenfloweRS is built on the SciRS2 scientific computing ecosystem:
TenfloweRS (Deep Learning Framework - TensorFlow-compatible API)
builds upon
OptiRS (ML Optimization Specialization)
builds upon
SciRS2 (Scientific Computing Foundation)
builds upon
ndarray, num-traits, etc. (Core Rust Scientific Stack)
This architecture provides:
- Advanced numerical operations via
scirs2-core - Automatic differentiation via
scirs2-autograd - Neural network abstractions via
scirs2-neural - Optimized algorithms via
optirs
Performance Benchmarks
Representative throughput figures on an AMD Ryzen 9 7950X (AVX2, 16 cores) and NVIDIA RTX 4090 (GPU).
CPU Tensor Operations (f32, release mode)
| Operation | Shape | TenfloweRS | Notes |
|---|---|---|---|
add |
[4096, 4096] | ~2.8 GB/s | SIMD-vectorized |
matmul |
[512, 512]² | ~35 GFLOPS | OpenBLAS backend |
relu |
[1M] | ~4.5 GB/s | Auto-vectorized |
softmax |
[batch=128, 1024] | ~890 MB/s | Numerically stable log-sum-exp |
GPU Operations (WGPU compute shaders, f32)
| Operation | Shape | Throughput |
|---|---|---|
matmul |
[2048, 2048]² | ~12 TFLOPS |
| Gaussian blur 5×5 | [H=512, W=512, C=3] | ~1800 MP/s |
| Random crop | [H=224, W=224, C=3] | ~3200 MP/s |
| Gaussian noise | [H=512, W=512, C=3] | ~2900 MP/s |
Data Pipeline
| Workload | Config | Throughput |
|---|---|---|
| CIFAR-10 prefetch | 4 workers, pinned | ~12,000 samples/s |
| ImageNet crop+resize+normalize | GPU transforms | ~2,400 samples/s |
| CSV streaming (1M rows) | SIMD stats | ~180 MB/s |
Numbers are indicative. Run
cargo bench -p tenflowers-datasetfor detailed measurements on your hardware.
Documentation
License
Licensed under the Apache License, Version 2.0 (LICENSE or http://www.apache.org/licenses/LICENSE-2.0).
Status
TenfloweRS v0.1.2 (2026-07-07). All 14,289 tests passing across the workspace (39 skipped), 0 clippy warnings, 0 TODO markers. The project comprises ~677K SLoC of Rust across 1,495 files in 6 published crates.