burn-dataset 0.22.0

Library with simple dataset APIs for creating ML data pipelines
docs.rs failed to build burn-dataset-0.22.0
Please check the build logs for more information.
See Builds for ideas on how to fix a failed build, or Metadata for how to configure docs.rs builds.
If you believe this is docs.rs' fault, open an issue.
Visit the last successful build: burn-dataset-0.22.0-pre.4

Burn Dataset

Datasets, transformations and data sources for Burn

Current Crates.io Version Documentation license

Applications enable the dataset feature of burn (also enabled by train) and use this crate as burn::data::dataset.

  • Dataset is random access by index. InMemDataset holds items in memory, and the sqlite feature adds SqliteDataset for datasets larger than memory.
  • transform: mapping, sampling, shuffling, windowing, partial views, selection and composition of datasets, applied lazily.
  • source: dataset downloads, including Hugging Face datasets.
  • vision, nlp, audio: ready-made datasets such as MNIST, CIFAR, image folders, AG News and Speech Commands, behind the features of the same names.

See the dataset chapter of the Burn Book.

Feature Flags

  • sqlite: SQLite-backed datasets, using the Turso engine. sqlite-bundled is a deprecated alias.
  • vision, nlp, audio: domain datasets and loaders. builtin-sources enables the downloadable vision and NLP datasets.
  • dataframe: datasets backed by a Polars DataFrame.
  • network: file downloads.
  • fake: generated datasets for tests.
  • tracing: instrument operations with the tracing crate.

Try the audio dataset with:

cargo run -p burn-dataset --example speech_commands --features audio

Part of the Burn deep learning framework. See the Burn Book and the API documentation.