Skip to main content

DatasetBuilder

Struct DatasetBuilder 

Source
pub struct DatasetBuilder<T: H5Type> { /* private fields */ }
Expand description

A fluent builder for creating datasets.

Obtained from H5File::new_dataset::<T>().

let file = H5File::create("builder.h5").unwrap();
let ds = file.new_dataset::<f32>()
    .shape(&[10, 20])
    .create("temperatures")
    .unwrap();

Implementations§

Source§

impl<T: H5Type> DatasetBuilder<T>

Source

pub fn shape<S: AsRef<[usize]>>(self, dims: S) -> Self

Set the dataset dimensions.

This is required before calling create, unless null was called instead. Use an empty slice &[] for a scalar (0-dimensional) dataset.

Source

pub fn scalar(self) -> Self

Create a scalar (0-dimensional) dataset holding a single value.

Source

pub fn null(self) -> Self

Create a dataset with the NULL dataspace: no elements at all.

Distinct from scalar, which holds exactly one element. A NULL dataset holds zero bytes of data and cannot be written to — write_raw and write_raw_bytes return an error, and it cannot be chunked or filtered, matching h5py’s h5py.Empty.

Source

pub fn chunk(self, chunk_dims: &[usize]) -> Self

Set chunk dimensions for chunked storage.

When set, the dataset uses chunked storage with the extensible array index. You should also call max_shape or resizable to allow extending.

Source

pub fn resizable(self) -> Self

Make all dimensions unlimited (resizable).

This sets max_dims to u64::MAX for all dimensions.

Source

pub fn max_shape(self, max: &[Option<usize>]) -> Self

Set maximum dimensions. None means unlimited for that dimension.

Source

pub fn compact(self) -> Self

Store the raw data inside the dataset’s object header — H5Pset_layout(dcpl, H5D_COMPACT).

A compact dataset costs no data block and no second seek to read, which suits the small per-run constants an analysis file is full of. It is bounded by what one object header message can hold (MAX_COMPACT_DATA bytes) and it is fixed in size: chunk, a filter, and an unlimited max_shape are all rejected at create, as libhdf5 rejects them.

let file = H5File::create("compact.h5").unwrap();
let ds = file.new_dataset::<i32>()
    .shape([16])
    .compact()
    .create("data")
    .unwrap();
ds.write_raw(&(0..16i32).collect::<Vec<_>>()).unwrap();
Source

pub fn early_allocation(self) -> Self

Allocate the whole of a chunked dataset’s storage at create — H5Pset_alloc_time(dcpl, H5D_ALLOC_TIME_EARLY), h5py’s alloc_time=h5d.ALLOC_TIME_EARLY.

Every chunk exists, holding the fill value, before anything is written, so an unwritten chunk costs a read of fill bytes rather than a miss. On a fixed-shape unfiltered dataset that is also what lets libhdf5 pick its cheapest chunk index — the implicit index, which is no index at all: the chunks are one contiguous run in grid order and a chunk’s address is arithmetic. This builder makes the same choice under the same conditions, so such a dataset is written with no index structure in the file.

Ignored by storage that has no chunk grid to allocate: contiguous, compact and NULL-dataspace datasets.

let file = H5File::create("implicit.h5").unwrap();
let ds = file.new_dataset::<i32>()
    .shape([16])
    .chunk(&[4])
    .early_allocation()
    .create("data")
    .unwrap();
ds.write_raw(&(0..16i32).collect::<Vec<_>>()).unwrap();
Source

pub fn deflate(self, level: u32) -> Self

Enable deflate (gzip) compression with the given level (0-9).

Requires chunked storage (call .chunk() before .create()). Level 0 = no compression, 9 = maximum compression. Default is 6.

Source

pub fn shuffle(self) -> Self

Enable the shuffle filter — H5Pset_shuffle(dcpl), h5py’s shuffle=True.

Shuffle reorders a chunk’s bytes by their position within an element, which typically improves how well a compressor behind it does on numeric data. It is a permutation, not a compressor: on its own it leaves the chunk exactly as large as it was, which is what H5Pset_shuffle without a compressor writes. Combine it with deflate to compress the shuffled stream. Requires chunked storage.

The element width the filter records is the dataset’s, so a datatype override is what it follows when the stored element is not T itself.

Source

pub fn shuffle_deflate(self, level: u32) -> Self

Enable shuffle + deflate compression — the same pipeline as .shuffle().deflate(level).

Shuffle reorders bytes by position within elements before compression, which typically improves compression ratios for numeric data. Requires chunked storage.

Source

pub fn zstd(self, level: u32) -> Self

Enable Zstandard compression with the given level (1-22, default 3).

Requires chunked storage (call .chunk() before .create()).

Source

pub fn filter_pipeline(self, pipeline: FilterPipeline) -> Self

Set a custom filter pipeline for compression.

This takes precedence over deflate and shuffle_deflate. Requires chunked storage.

Source

pub fn datatype(self, dt: DatatypeMessage) -> Self

Override the stored element datatype.

By default the dataset is created with the datatype derived from the Rust type parameter T (H5Type::hdf5_type). Use this to store a different on-disk datatype than the in-memory element type — for example a reduced-precision fixed-point type that matches an N-bit filter (see FilterPipeline::nbit). The element byte size of the override must equal T::element_size(); the N-bit filter packs the significant bits within that fixed footprint.

Source

pub fn committed_type(self, path: &str) -> Self

Build the dataset on the committed (named) datatype at path — h5py’s dtype=f["name"], H5Dcreate2 with a committed type id.

The dataset does not describe its type: its header stores a pointer to that object, so the type is defined once and every dataset sharing it is guaranteed to agree. The type comes from the committed object, so this supersedes both T and datatype.

The path is resolved at create, which fails when no committed datatype is there — commit it with H5File::commit_datatype first.

let file = H5File::create("committed.h5").unwrap();
file.commit_datatype("temperature", DatatypeMessage::f64_type()).unwrap();
file.new_dataset::<f64>()
    .committed_type("temperature")
    .shape([4])
    .create("readings")
    .unwrap();
Source

pub fn object_references(self) -> Self

Store object references — h5py’s h5py.ref_dtype.

The elements are written with write_object_references and name objects by path. The element width is the file’s address size, so the datatype is resolved at create rather than here; it overrides both T and any datatype call.

let file = H5File::create("refs.h5").unwrap();
file.new_dataset::<i32>().shape([4]).create("target").unwrap();
let refs = file.new_dataset::<u64>()
    .object_references()
    .shape([1])
    .create("refs")
    .unwrap();
refs.write_object_references(&["/target"]).unwrap();
file.close().unwrap();
Source

pub fn std_object_references(self) -> Self

Store revised object references — the 1.12 H5T_STD_REF.

Same paths and same write_object_references call as object_references; only the stored element differs, carrying the reference’s kind alongside the address so one datatype can hold every reference kind. h5py cannot read it, so prefer the pre-1.12 form for files h5py will open.

let file = H5File::create("stdrefs.h5").unwrap();
file.new_dataset::<i32>().shape([4]).create("target").unwrap();
let refs = file.new_dataset::<u64>()
    .std_object_references()
    .shape([1])
    .create("refs")
    .unwrap();
refs.write_object_references(&["/target"]).unwrap();
file.close().unwrap();
Source

pub fn std_region_references(self) -> Self

Store revised region references — H5R_DATASET_REGION2, written into the same H5T_STD_REF datatype std_object_references makes.

The elements are written with write_std_region_references. What distinguishes them from the pre-1.12 region_references is the element, not the datatype: a 1.12 element names its own kind, so one dataset of this type may hold object, region and attribute references together. h5py 3.15 cannot read any of them, so prefer the pre-1.12 form for files h5py will open.

let file = H5File::options().libver(LibverBound::V112).create("stdregions.h5").unwrap();
file.new_dataset::<i32>().shape([8]).create("target").unwrap();
let refs = file.new_dataset::<u64>()
    .std_region_references()
    .shape([1])
    .create("refs")
    .unwrap();
let rows = Selection::Hyperslab {
    rank: 1,
    form: Hyperslab::Blocks(vec![HyperslabBlock { start: vec![0], end: vec![2] }]),
};
refs.write_std_region_references(&[("/target", rows)]).unwrap();
file.close().unwrap();
Source

pub fn attribute_references(self) -> Self

Store attribute references — H5R_ATTR, the one reference kind with no pre-1.12 form, in the same H5T_STD_REF datatype std_object_references makes.

The elements are written with write_attribute_references and name an object and one of its attributes. h5py 3.15 cannot read them.

let file = H5File::options().libver(LibverBound::V112).create("attrrefs.h5").unwrap();
let target = file.new_dataset::<i32>().shape([4]).create("target").unwrap();
target.new_attr::<i32>().shape([3]).create("note").unwrap()
    .write_array(&[7i32, 8, 9]).unwrap();
let refs = file.new_dataset::<u64>()
    .attribute_references()
    .shape([1])
    .create("refs")
    .unwrap();
refs.write_attribute_references(&[("/target", "note")]).unwrap();
file.close().unwrap();
Source

pub fn region_references(self) -> Self

Store dataset region references — h5py’s h5py.regionref_dtype.

The elements are written with write_region_references and name a dataset plus a selection over it. The element is a global-heap id, so its width follows the file’s address size and the datatype is resolved at create rather than here; it overrides both T and any datatype call.

let file = H5File::create("regions.h5").unwrap();
file.new_dataset::<i32>().shape([8]).create("target").unwrap();
let refs = file.new_dataset::<u64>()
    .region_references()
    .shape([1])
    .create("refs")
    .unwrap();
let rows = Selection::Hyperslab {
    rank: 1,
    form: Hyperslab::Blocks(vec![HyperslabBlock { start: vec![0], end: vec![2] }]),
};
refs.write_region_references(&[("/target", rows)]).unwrap();
file.close().unwrap();
Source

pub fn external(self, files: &[(&str, u64, u64)]) -> Self

Keep the raw data in files outside this one — H5Pset_external, h5py’s external=[(name, offset, size)].

Each entry is a file name, the byte offset in it where that entry’s region starts, and how many bytes of the dataset it holds; the entries concatenate, in order, into the dataset’s bytes and together must cover them. A relative name is resolved against HDF5_EXTFILE_PREFIX the way libhdf5 resolves it, so the same name reads back through this crate and through h5py. The storage is contiguous by definition, which rules out chunk, a filter, compact, null and either reference kind.

The named files are created on first write and never truncated, so several datasets may own disjoint ranges of one file.

The last entry may take the unlimited size external_file_list::UNLIMITED (H5O_EFL_UNLIMITED), which makes it absorb the whole rest of the dataset however far it grows. A dataset whose max_shape is unlimited must have one, since no finite reservation could cover it, and only the first dimension may be extendible — both H5D__efl_construct’s rules.

let file = H5File::create("ext.h5").unwrap();
let ds = file.new_dataset::<i32>()
    .shape([16])
    .external(&[("ext.raw", 0, 64)])
    .create("data")
    .unwrap();
ds.write_raw(&(0..16i32).collect::<Vec<_>>()).unwrap();
Source

pub fn efile_prefix(self, prefix: impl Into<String>) -> Self

H5Pset_efile_prefix on the dapl H5Dcreate2 takes — the directory the raw data files named by external are created under, and looked for under on every later write through this handle.

H5D__create builds dset->shared->extfile_prefix from the dapl (H5Dint.c:1318) and H5D__efl_write joins each slot name against it with the same single-path H5_combine_path the read side uses (H5Defl.c:429-431) — so this decides where the bytes land, and the prefix a later reader names must agree for it to find them.

Measured under libhdf5 1.14.6 and 2.0.0: writing through a dapl that names a directory creates the raw data file there and nowhere else, and a directory that does not exist fails the write outright rather than being created.

It shares DatasetAccess::efile_prefix’s rules, both being H5D__build_file_prefix: HDF5_EXTFILE_PREFIX shadows this outright (H5Dint.c:1084-1090), a leading ${ORIGIN} stands for the directory holding the HDF5 file (:1105-1113), and "." or "" means no prefix (:1098-1102), which leaves a stored name to resolve against the process’s current directory.

Ignored by a dataset that names no external files, which has no slot name to join.

Source

pub fn virtual_mapping( self, virtual_selection: Selection, source_file: &str, source_dataset: &str, source_selection: Selection, ) -> Self

Map part of this dataset onto part of a dataset in another file, making it virtual — H5Pset_virtual, one VirtualLayout[...] = VirtualSource(...) assignment in h5py.

The arguments are H5Pset_virtual’s, in its order: which elements of this dataset the mapping fills, the file and dataset the data comes from, and which elements of that source dataset it comes from. Call it once per mapping; they apply in the order given, which is the order libhdf5 resolves overlapping ones in.

The source file is named exactly as stored — resolved against HDF5_VDS_PREFIX, or the virtual dataset’s own directory, when the file is read — and "." means this file. Nothing is opened or checked here: a source that does not exist yet is legal, and reads of the unmapped or unresolvable parts return the fill_value.

A virtual dataset stores nothing of its own, which rules out chunk, a filter, compact, null, external and either reference kind — and makes writing to it an error, since its elements belong to the source datasets.

An unlimited (H5S_UNLIMITED) selection is written as one: such a mapping grows with its source, and the dataset’s extent in that dimension is whatever the sources reachable when it is opened supply (H5D__virtual_set_extent_unlim). Give it a max_shape unlimited in the same dimension, as libhdf5 requires of the dataspace behind one.

A source name may carry libhdf5’s printf-style substitutions: %b is the block index and %% an escaped literal %. One such mapping stands for the family of source datasets that fill the successive blocks of an unlimited virtual selection, so it is legal only with an unlimited virtual selection over a limited source selection, and the dataset’s extent stops at the first block whose source is missing.

let file = H5File::create("vds.h5").unwrap();
let ds = file.new_dataset::<i32>()
    .shape([16])
    .virtual_mapping(Selection::All, "src.h5", "src", Selection::All)
    .create("vds")
    .unwrap();
Source

pub fn fill_value(self, value: T) -> Self

Set a user-defined fill value for unwritten elements.

Without this, datasets use the HDF5 default zero-fill. When set, the value is written into the dataset’s fill-value message (fill_defined = 2), so HDF5 readers treat unallocated chunks and unwritten regions as this value rather than zero.

let file = H5File::create("fv.h5").unwrap();
let ds = file.new_dataset::<f32>()
    .shape(&[100])
    .fill_value(f32::NAN)
    .create("data")
    .unwrap();
Source

pub fn fill_time(self, time: FillTime) -> Self

Set when the fill value is written into allocated storage — H5Pset_fill_time.

Without this, a dataset gets FillTime::IfSet (H5D_CRT_FILL_TIME_DEF), the default every dataset creation property list carries. FillTime::Never applies to a dataset with no fill value too: it only stops this writer’s own eager tiling of the value into newly allocated storage, not the default zero-fill that storage already has, so its only observable effect is on a dataset that also calls fill_value.

let file = H5File::create("fv.h5").unwrap();
let ds = file.new_dataset::<f32>()
    .shape(&[100])
    .fill_value(f32::NAN)
    .fill_time(FillTime::Never)
    .create("data")
    .unwrap();
Source

pub fn create(self, name: &str) -> Result<H5Dataset>

Finalize and create the dataset with the given name.

The name is the link name within the root group (e.g. "data" or "group1/data" once nested groups are supported).

Auto Trait Implementations§

§

impl<T> !RefUnwindSafe for DatasetBuilder<T>

§

impl<T> !Send for DatasetBuilder<T>

§

impl<T> !Sync for DatasetBuilder<T>

§

impl<T> !UnwindSafe for DatasetBuilder<T>

§

impl<T> Freeze for DatasetBuilder<T>
where PhantomData<T>: Freeze,

§

impl<T> Unpin for DatasetBuilder<T>
where PhantomData<T>: Unpin,

§

impl<T> UnsafeUnpin for DatasetBuilder<T>

Blanket Implementations§

Source§

impl<T> Any for T
where T: 'static + ?Sized,

Source§

fn type_id(&self) -> TypeId

Gets the TypeId of self. Read more
Source§

impl<T> Borrow<T> for T
where T: ?Sized,

Source§

fn borrow(&self) -> &T

Immutably borrows from an owned value. Read more
Source§

impl<T> BorrowMut<T> for T
where T: ?Sized,

Source§

fn borrow_mut(&mut self) -> &mut T

Mutably borrows from an owned value. Read more
Source§

impl<T> From<T> for T

Source§

fn from(t: T) -> T

Returns the argument unchanged.

Source§

impl<T, U> Into<U> for T
where U: From<T>,

Source§

fn into(self) -> U

Calls U::from(self).

That is, this conversion is whatever the implementation of From<T> for U chooses to do.

Source§

impl<T, S> SimdFrom<T, S> for T
where S: Simd,

Source§

fn simd_from(_simd: S, value: T) -> T

Source§

impl<F, T, S> SimdInto<T, S> for F
where T: SimdFrom<F, S>, S: Simd,

Source§

fn simd_into(self, simd: S) -> T

Source§

impl<T, U> TryFrom<U> for T
where U: Into<T>,

Source§

type Error = !

The type returned in the event of a conversion error.
Source§

fn try_from(value: U) -> Result<T, !>

Performs the conversion.
Source§

impl<T, U> TryInto<U> for T
where U: TryFrom<T>,

Source§

type Error = <U as TryFrom<T>>::Error

The type returned in the event of a conversion error.
Source§

fn try_into(self) -> Result<U, <U as TryFrom<T>>::Error>

Performs the conversion.