1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
//! # Burn
//!
//! Burn is a new comprehensive dynamic Deep Learning Framework built using Rust
//! with extreme flexibility, compute efficiency and portability as its primary goals.
//!
//! ## Performance
//!
//! Because we believe the goal of a deep learning framework is to convert computation
//! into useful intelligence, we have made performance a core pillar of Burn.
//! We strive to achieve top efficiency by leveraging multiple optimization techniques:
//!
//! - Automatic kernel fusion
//! - Asynchronous execution
//! - Thread-safe building blocks
//! - Intelligent memory management
//! - Automatic kernel selection
//! - Hardware specific features
//! - Custom Backend Extension
//!
//! ## Training & Inference
//!
//! The whole deep learning workflow is made easy with Burn, as you can monitor your training progress
//! with an ergonomic dashboard, and run inference everywhere from embedded devices to large GPU clusters.
//!
//! Burn was built from the ground up with training and inference in mind. It's also worth noting how Burn,
//! in comparison to frameworks like PyTorch, simplifies the transition from training to deployment,
//! eliminating the need for code changes.
//!
//! ## Backends
//!
//! Burn strives to be as fast as possible on as many hardwares as possible, with robust implementations.
//! We believe this flexibility is crucial for modern needs where you may train your models in the cloud,
//! then deploy on customer hardwares, which vary from user to user.
//!
//! Burn's backend architecture lets you swap backends while keeping the same model code. You can
//! enable multiple backends in the same application and choose the device for your tensors and
//! modules at runtime through [`tensor::Device`]. This gives you the freedom to use different
//! backends side by side and select the hardware best suited to each workload.
//!
//! Autodifferentiation and automatic kernel fusion integrate with the same tensor and module APIs,
//! so models benefit from these capabilities on supported backends without changing their implementation.
//!
//! - WGPU (WebGPU): Cross-Platform GPU Backend
//! - LibTorch: Backend using the LibTorch bindings (deprecated)
//! - Flex: Pure-Rust CPU backend (std, no_std, WebAssembly)
//! - Autodiff: Backend decorator that brings backpropagation to any backend
//! - Fusion: Backend decorator that brings kernel fusion to backends that support it
//!
//! # Quantization
//!
//! Quantization techniques perform computations and store tensors in lower precision data types like
//! 8-bit integer instead of floating point precision. There are multiple approaches to quantize a deep
//! learning model categorized as post-training quantization (PTQ) and quantization aware training (QAT).
//!
//! In post-training quantization, the model is trained in floating point precision and later converted
//! to the lower precision data type. There are two types of post-training quantization:
//!
//! 1. Static quantization: quantizes the weights and activations of the model. Quantizing the
//! activations statically requires data to be calibrated (i.e., recording the activation values to
//! compute the optimal quantization parameters with representative data).
//! 2. Dynamic quantization: quantized the weights ahead of time (like static quantization) but the
//! activations are dynamically at runtime.
//!
//! Sometimes post-training quantization is not able to achieve acceptable task accuracy. In general,
//! this is where quantization-aware training (QAT) can be used: during training, fake-quantization
//! modules are inserted in the forward and backward passes to simulate quantization effects, allowing
//! the model to learn representations that are more robust to reduced precision.
//!
//! Burn does not currently support QAT. Only post-training quantization (PTQ) is implemented at this
//! time.
//!
//! Quantization support in Burn is currently in active development. It supports the following PTQ modes on some backends:
//! - Per-tensor and per-block quantization to 8-bit, 4-bit and 2-bit representations
//!
//! ## Feature Flags
//!
//! The following feature flags are available.
//! Default features include `std`, `optim` (and therefore `autodiff`), and `rl`, but no execution
//! backend.
//! Select a backend explicitly, for example `features = ["wgpu"]` or `["flex"]`.
//! Specialized operations are also opt-in, for example `features = ["flex", "signal"]`.
//! Backend-free builds can define tensor/model APIs without installing an execution backend.
//! `Device::default()` panics if no execution backend is available; graph capture remains
//! available through `Device::capture()` with the `capture` feature.
//!
//! - Training
//! - `train`: Enables features `dataset` and `optim` and provides a training environment
//! - `optim`: Enables optimizers and learning rate schedulers (implies `autodiff`)
//! - `rl`: Enables reinforcement learning utilities
//! - `tui`: Includes Text UI with progress bar and plots (requires `train`)
//! - `metrics`: Includes system info metrics (CPU/GPU usage, etc.) (requires `train`)
//! - Dataset
//! - `dataset`: Includes a datasets library
//! - `audio`: Enables audio datasets (SpeechCommandsDataset)
//! - `sqlite`: Stores datasets in an SQLite database, backed by [Turso](https://turso.tech/)
//! - `sqlite-bundled`: Deprecated alias for `sqlite`
//! - `vision`: Enables vision datasets (MnistDataset) and the `burn-vision` ops module
//! - Backends
//! - `wgpu`: Makes available the WGPU backend
//! - `webgpu`: Makes available the `wgpu` backend with the WebGPU Shading Language (WGSL) compiler
//! - `vulkan`: Makes available the `wgpu` backend with the alternative SPIR-V compiler
//! - `cuda`: Makes available the CUDA backend
//! - `metal`: Makes available the Metal backend
//! - `rocm`: Makes available the ROCm backend
//! - `cpu`: Makes available the CubeCL CPU backend
//! - `tch`: Makes available the LibTorch backend (deprecated - use a CubeCL backend instead)
//! - `flex`: Makes available the Flex backend (pure-Rust CPU, std/no_std/WASM)
//! - `ndarray`: Makes available the NdArray backend (deprecated - use `flex` instead)
//! - Backend specifications
//! - `simd`: Enable SIMD codegen in the CPU backends
//! - `rayon`: Enable multi-threaded execution in the CPU backends
//! - `accelerate`: If supported, Accelerate will be used
//! - `blas-netlib`: If supported, Blas Netlib will be use
//! - `openblas`: If supported, Openblas will be use
//! - `openblas-system`: If supported, Openblas installed on the system will be use
//! - `autotune`: Enable running benchmarks to select the best kernel in backends that support it.
//! - `autotune-checks`: Check that every autotune candidate produces the same output (debugging).
//! - `x86-v4`: Enable AVX-512 matmul kernels in the Flex backend.
//! - `apple-amx`: Enable the experimental Apple AMX matmul kernels in the Flex backend.
//! - `template`: Enable template-based custom kernels in the WGPU backend.
//! - `fusion`: Enable operation fusion in backends that support it.
//! - `tracing`: Enable diagnostic tracing in the selected backends (disabled by default).
//! - Backend decorators
//! - `autodiff`: Makes available the Autodiff backend
//! - Model Storage
//! - `store`: Enables the `burn-store` snapshot tooling and burnpack stores; with `std`, this
//! also includes SafeTensors
//! - `safetensors`: Enables SafeTensors import and export in `no_std` builds (implies `store`)
//! - `pytorch`: Enables PyTorch checkpoint import (implies `store`)
//! - Others:
//! - `std`: Activates the standard library (deactivate for no_std)
//! - `linalg`: Enables linear algebra operations
//! - `capture`: Makes the non-executing graph capture backend available.
//! - `ir`: Makes Burn's operation intermediate representation available.
//! - `cubecl`: Re-exports CubeCL as `burn::cubecl` for writing custom kernels.
//! - `signal`: Enables signal processing operations from `burn-signal`.
//! - `extension`: Enables the backend extension API, including `Tensor::from_primitive`.
//! - `remote`: Enables remote devices over Iroh; `remote-websocket` adds the WebSocket transport.
//! - `remote-server`: Enables the remote server (implies `remote`).
//! - `network`: Enables network utilities (currently, only a file downloader with progress bar)
//!
//! You can also check the details in sub-crates [`burn-core`](https://docs.rs/burn-core) and [`burn-train`](https://docs.rs/burn-train).
//!
//! ### Backend tracing
//!
//! Add `"tracing"` to the features of your `burn` dependency to compile backend instrumentation,
//! including autodiff and fusion spans. When depending directly on `burn-autodiff` or
//! `burn-fusion`, enable their `tracing` feature instead. These spans are opt-in: configuring a
//! tracing subscriber alone does not enable them. Configure your subscriber to include the
//! `trace` level to observe tensor operation spans.
//!
//! The feature propagates to enabled backends without selecting an additional backend. Normal
//! training logs remain available without this feature.
pub use *;
/// Linear algebra operations.
/// Core module infrastructure and neural-network initializers.
/// Tensor types and compatibility re-exports.
/// Train module
/// Module for reinforcement learning.
pub use server;
/// Model storage and serialization: the non-generic record system (always available), plus,
/// with the `store` feature, the snapshot tooling and burnpack stores. The `safetensors` and
/// `pytorch` features add those importers.
/// Neural network module.
pub use ;
/// Optimizers module.
// For backward compat, `burn::lr_scheduler::*`
/// Learning rate scheduler module.
// For backward compat, `burn::grad_clipping::*`
/// Gradient clipping module.
/// CubeCL module re-export.
/// Vision module.
/// Signal processing module.