webdataset
High-performance sequential dataset loading from tar archives, in the WebDataset format.
A WebDataset is a set of tar files ("shards"). Inside a shard, the files that share a basename make up one training sample:
imagenet-000000.tar
n03991062_24866.jpg n03991062_24866.cls
n03995372_9042.jpg n03995372_9042.cls
That is the entire format. There is no index, no metadata file and no conversion step. Because reading is sequential, a shard streams at the full bandwidth of the device or the network link, and a dataset is just a list of URLs.
use ;
use ;
What this crate gives you
- A re-runnable pipeline of stages, so
repeatandwith_epochwork. - Decoders for
.txt,.cls,.json,.npy,.npz,.ten,.cbor,.mpand images, plus gzip chaining. - Shard lists with brace expansion, resampling, and node and worker splitting.
- A threaded loader, and an async pipeline over
futures_io::AsyncRead. - Writers, from a single archive to a numbered series of shards.
Features
| feature | adds |
|---|---|
threads (default) |
per-shard read-ahead and the multi-worker loader |
subprocess (default) |
pipe: and the curl/gsutil/ais schemes |
yaml (default) |
multi-source dataset specifications |
async |
reading from any AsyncRead, yielding a Stream |
image |
.jpg, .png and friends |
libjpeg |
decode JPEG with libjpeg-turbo, matching Pillow bit for bit |
msgpack, cbor, npz |
those formats |
zstd, bzip2, xz |
shards in those containers |
wasm-js |
host randomness on wasm32-unknown-unknown |
full |
every format, plus threads, subprocesses and async |
Parity with the reference implementation
This is a port, not a binding, and its output is checked against the Python
library field by field — 5,646 fields across every bundled shard and every
imagespec, all identical. See the
workspace README for how that is
measured and what the deliberate differences are.
License: BSD-3-Clause