burn 0.22.0-pre.4

Flexible and Comprehensive Deep Learning Framework in Rust
Documentation
#![cfg_attr(not(feature = "std"), no_std)]
#![warn(missing_docs)]

//! # Burn
//!
//! Burn is a new comprehensive dynamic Deep Learning Framework built using Rust
//! with extreme flexibility, compute efficiency and portability as its primary goals.
//!
//! ## Performance
//!
//! Because we believe the goal of a deep learning framework is to convert computation
//! into useful intelligence, we have made performance a core pillar of Burn.
//! We strive to achieve top efficiency by leveraging multiple optimization techniques:
//!
//! - Automatic kernel fusion
//! - Asynchronous execution
//! - Thread-safe building blocks
//! - Intelligent memory management
//! - Automatic kernel selection
//! - Hardware specific features
//! - Custom Backend Extension
//!
//! ## Training & Inference
//!
//! The whole deep learning workflow is made easy with Burn, as you can monitor your training progress
//! with an ergonomic dashboard, and run inference everywhere from embedded devices to large GPU clusters.
//!
//! Burn was built from the ground up with training and inference in mind. It's also worth noting how Burn,
//! in comparison to frameworks like PyTorch, simplifies the transition from training to deployment,
//! eliminating the need for code changes.
//!
//! ## Backends
//!
//! Burn strives to be as fast as possible on as many hardwares as possible, with robust implementations.
//! We believe this flexibility is crucial for modern needs where you may train your models in the cloud,
//! then deploy on customer hardwares, which vary from user to user.
//!
//! Burn's backend architecture lets you swap backends while keeping the same model code. You can
//! enable multiple backends in the same application and choose the device for your tensors and
//! modules at runtime through [`tensor::Device`]. This gives you the freedom to use different
//! backends side by side and select the hardware best suited to each workload.
//!
//! Autodifferentiation and automatic kernel fusion integrate with the same tensor and module APIs,
//! so models benefit from these capabilities on supported backends without changing their implementation.
//!
//! - WGPU (WebGPU): Cross-Platform GPU Backend
//! - LibTorch: Backend using the LibTorch bindings (deprecated)
//! - Flex: Pure-Rust CPU backend (std, no_std, WebAssembly)
//! - Autodiff: Backend decorator that brings backpropagation to any backend
//! - Fusion: Backend decorator that brings kernel fusion to backends that support it
//!
//! # Quantization
//!
//! Quantization techniques perform computations and store tensors in lower precision data types like
//! 8-bit integer instead of floating point precision. There are multiple approaches to quantize a deep
//! learning model categorized as post-training quantization (PTQ) and quantization aware training (QAT).
//!
//! In post-training quantization, the model is trained in floating point precision and later converted
//! to the lower precision data type. There are two types of post-training quantization:
//!
//! 1. Static quantization: quantizes the weights and activations of the model. Quantizing the
//!    activations statically requires data to be calibrated (i.e., recording the activation values to
//!    compute the optimal quantization parameters with representative data).
//! 2. Dynamic quantization: quantized the weights ahead of time (like static quantization) but the
//!    activations are dynamically at runtime.
//!
//! Sometimes post-training quantization is not able to achieve acceptable task accuracy. In general,
//! this is where quantization-aware training (QAT) can be used: during training, fake-quantization
//! modules are inserted in the forward and backward passes to simulate quantization effects, allowing
//! the model to learn representations that are more robust to reduced precision.
//!
//! Burn does not currently support QAT. Only post-training quantization (PTQ) is implemented at this
//! time.
//!
//! Quantization support in Burn is currently in active development. It supports the following PTQ modes on some backends:
//! - Per-tensor and per-block quantization to 8-bit, 4-bit and 2-bit representations
//!
//! ## Feature Flags
//!
//! The following feature flags are available.
//! Default features include `std`, `optim` (and therefore `autodiff`), and `rl`, but no execution
//! backend.
//! Select a backend explicitly, for example `features = ["wgpu"]` or `["flex"]`.
//! Specialized operations are also opt-in, for example `features = ["flex", "signal"]`.
//! Backend-free builds can define tensor/model APIs without installing an execution backend.
//! `Device::default()` panics if no execution backend is available; graph capture remains
//! available through `Device::capture()` with the `capture` feature.
//!
//! - Training
//!   - `train`: Enables features `dataset` and `optim` and provides a training environment
//!   - `optim`: Enables optimizers and learning rate schedulers (implies `autodiff`)
//!   - `rl`: Enables reinforcement learning utilities
//!   - `tui`: Includes Text UI with progress bar and plots (requires `train`)
//!   - `metrics`: Includes system info metrics (CPU/GPU usage, etc.) (requires `train`)
//! - Dataset
//!   - `dataset`: Includes a datasets library
//!   - `audio`: Enables audio datasets (SpeechCommandsDataset)
//!   - `sqlite`: Stores datasets in an SQLite database, backed by [Turso](https://turso.tech/)
//!   - `sqlite-bundled`: Deprecated alias for `sqlite`
//!   - `vision`: Enables vision datasets (MnistDataset) and the `burn-vision` ops module
//! - Backends
//!   - `wgpu`: Makes available the WGPU backend
//!   - `webgpu`: Makes available the `wgpu` backend with the WebGPU Shading Language (WGSL) compiler
//!   - `vulkan`: Makes available the `wgpu` backend with the alternative SPIR-V compiler
//!   - `cuda`: Makes available the CUDA backend
//!   - `metal`: Makes available the Metal backend
//!   - `rocm`: Makes available the ROCm backend
//!   - `cpu`: Makes available the CubeCL CPU backend
//!   - `tch`: Makes available the LibTorch backend (deprecated - use a CubeCL backend instead)
//!   - `flex`: Makes available the Flex backend (pure-Rust CPU, std/no_std/WASM)
//!   - `ndarray`: Makes available the NdArray backend (deprecated - use `flex` instead)
//! - Backend specifications
//!   - `simd`: Enable SIMD codegen in the CPU backends
//!   - `rayon`: Enable multi-threaded execution in the CPU backends
//!   - `accelerate`: If supported, Accelerate will be used
//!   - `blas-netlib`: If supported, Blas Netlib will be use
//!   - `openblas`: If supported, Openblas will be use
//!   - `openblas-system`: If supported, Openblas installed on the system will be use
//!   - `autotune`: Enable running benchmarks to select the best kernel in backends that support it.
//!   - `autotune-checks`: Check that every autotune candidate produces the same output (debugging).
//!   - `x86-v4`: Enable AVX-512 matmul kernels in the Flex backend.
//!   - `apple-amx`: Enable the experimental Apple AMX matmul kernels in the Flex backend.
//!   - `template`: Enable template-based custom kernels in the WGPU backend.
//!   - `fusion`: Enable operation fusion in backends that support it.
//!   - `tracing`: Enable diagnostic tracing in the selected backends (disabled by default).
//! - Backend decorators
//!   - `autodiff`: Makes available the Autodiff backend
//! - Model Storage
//!   - `store`: Enables the `burn-store` snapshot tooling and burnpack stores; with `std`, this
//!     also includes SafeTensors
//!   - `safetensors`: Enables SafeTensors import and export in `no_std` builds (implies `store`)
//!   - `pytorch`: Enables PyTorch checkpoint import (implies `store`)
//! - Others:
//!   - `std`: Activates the standard library (deactivate for no_std)
//!   - `linalg`: Enables linear algebra operations
//!   - `capture`: Makes the non-executing graph capture backend available.
//!   - `ir`: Makes Burn's operation intermediate representation available.
//!   - `cubecl`: Re-exports CubeCL as `burn::cubecl` for writing custom kernels.
//!   - `signal`: Enables signal processing operations from `burn-signal`.
//!   - `extension`: Enables the backend extension API, including `Tensor::from_primitive`.
//!   - `remote`: Enables remote devices over Iroh; `remote-websocket` adds the WebSocket transport.
//!   - `remote-server`: Enables the remote server (implies `remote`).
//!   - `network`: Enables network utilities (currently, only a file downloader with progress bar)
//!
//! You can also check the details in sub-crates [`burn-core`](https://docs.rs/burn-core) and [`burn-train`](https://docs.rs/burn-train).
//!
//! ### Backend tracing
//!
//! Add `"tracing"` to the features of your `burn` dependency to compile backend instrumentation,
//! including autodiff and fusion spans. When depending directly on `burn-autodiff` or
//! `burn-fusion`, enable their `tracing` feature instead. These spans are opt-in: configuring a
//! tracing subscriber alone does not enable them. Configure your subscriber to include the
//! `trace` level to observe tensor operation spans.
//!
//! The feature propagates to enabled backends without selecting an additional backend. Normal
//! training logs remain available without this feature.

pub use burn_core::*;

/// Linear algebra operations.
#[cfg(feature = "linalg")]
pub mod linalg {
    pub use burn_linalg::*;
}

/// Core module infrastructure and neural-network initializers.
pub mod module {
    pub use burn_core::module::*;
    pub use burn_nn::Initializer;
}

/// Tensor types and compatibility re-exports.
pub mod tensor {
    pub use burn_core::tensor::*;

    /// Compatibility path for signal processing operations.
    #[cfg(feature = "signal")]
    pub mod signal {
        pub use burn_signal::*;
    }

    /// Compatibility path for linear algebra operations.
    #[cfg(feature = "linalg")]
    pub mod linalg {
        pub use burn_linalg::*;
    }
}

/// Train module
#[cfg(feature = "train")]
pub mod train {
    pub use burn_train::*;
}

/// Module for reinforcement learning.
#[cfg(feature = "rl")]
pub mod rl {
    pub use burn_rl::*;
}

#[cfg(feature = "remote-server")]
pub use burn_core::tensor::server;

/// Model storage and serialization: the non-generic record system (always available), plus,
/// with the `store` feature, the snapshot tooling and burnpack stores. The `safetensors` and
/// `pytorch` features add those importers.
pub mod store {
    pub use burn_core::store::*;
    #[cfg(feature = "store")]
    pub use burn_store::*;
}

/// Neural network module.
pub mod nn {
    pub use burn_nn::*;
}

pub use burn_std::config::{BurnConfig, config as runtime_config};

#[cfg(all(test, feature = "capture"))]
mod capture_tests {
    use crate::{module::Module, nn::BatchNormConfig, tensor::Device};

    #[test]
    fn capture_feature_exposes_the_user_facing_device_api() {
        let device = Device::capture();
        let captured = device
            .capture_scope(|scope| scope.complete([], []))
            .unwrap();

        assert!(captured.graph.operations.is_empty());
    }

    #[test]
    fn shared_running_state_moves_across_capture_scopes() {
        let module = BatchNormConfig::new(3).init(&Device::default());
        let first_device = Device::capture();
        let second_device = Device::capture();

        let first = first_device
            .capture_scope(|scope| {
                let _module = module.clone().to_device(&first_device);
                scope.complete([], [])
            })
            .unwrap();
        let second = second_device
            .capture_scope(|scope| {
                let _module = module.clone().to_device(&second_device);
                scope.complete([], [])
            })
            .unwrap();

        assert_eq!(first.values.len(), 4);
        assert_eq!(second.values.len(), 4);
    }
}

/// Optimizers module.
#[cfg(feature = "optim")]
pub mod optim {
    pub use burn_optim::*;
}

// For backward compat, `burn::lr_scheduler::*`
/// Learning rate scheduler module.
#[cfg(all(feature = "optim", feature = "std"))]
pub mod lr_scheduler {
    pub use burn_optim::lr_scheduler::*;
}
// For backward compat, `burn::grad_clipping::*`
/// Gradient clipping module.
#[cfg(feature = "optim")]
pub mod grad_clipping {
    pub use burn_optim::grad_clipping::*;
}

/// CubeCL module re-export.
#[cfg(feature = "cubecl")]
pub mod cubecl {
    pub use cubecl::*;
}

#[cfg(feature = "vision")]
/// Vision module.
pub mod vision {
    pub use burn_vision::*;
}

#[cfg(feature = "signal")]
/// Signal processing module.
pub mod signal {
    pub use burn_signal::*;
}

pub mod prelude {
    //! Structs and macros used by most projects. Add `use
    //! burn::prelude::*` to your code to quickly get started with
    //! Burn.
    pub use burn_core::prelude::*;

    pub use crate::nn;
}