pub struct DatasetBuilder<T: H5Type> { /* private fields */ }Expand description
A fluent builder for creating datasets.
Obtained from H5File::new_dataset::<T>().
let file = H5File::create("builder.h5").unwrap();
let ds = file.new_dataset::<f32>()
.shape(&[10, 20])
.create("temperatures")
.unwrap();Implementations§
Source§impl<T: H5Type> DatasetBuilder<T>
impl<T: H5Type> DatasetBuilder<T>
Sourcepub fn null(self) -> Self
pub fn null(self) -> Self
Create a dataset with the NULL dataspace: no elements at all.
Distinct from scalar, which holds exactly one
element. A NULL dataset holds zero bytes of data and cannot be
written to — write_raw and
write_raw_bytes return an error, and
it cannot be chunked or filtered, matching h5py’s h5py.Empty.
Sourcepub fn resizable(self) -> Self
pub fn resizable(self) -> Self
Make all dimensions unlimited (resizable).
This sets max_dims to u64::MAX for all dimensions.
Sourcepub fn max_shape(self, max: &[Option<usize>]) -> Self
pub fn max_shape(self, max: &[Option<usize>]) -> Self
Set maximum dimensions. None means unlimited for that dimension.
Sourcepub fn compact(self) -> Self
pub fn compact(self) -> Self
Store the raw data inside the dataset’s object header —
H5Pset_layout(dcpl, H5D_COMPACT).
A compact dataset costs no data block and no second seek to read, which
suits the small per-run constants an analysis file is full of. It is
bounded by what one object header message can hold
(MAX_COMPACT_DATA bytes) and it
is fixed in size: chunk, a filter, and an unlimited
max_shape are all rejected at
create, as libhdf5 rejects them.
let file = H5File::create("compact.h5").unwrap();
let ds = file.new_dataset::<i32>()
.shape([16])
.compact()
.create("data")
.unwrap();
ds.write_raw(&(0..16i32).collect::<Vec<_>>()).unwrap();Sourcepub fn early_allocation(self) -> Self
pub fn early_allocation(self) -> Self
Allocate the whole of a chunked dataset’s storage at create —
H5Pset_alloc_time(dcpl, H5D_ALLOC_TIME_EARLY), h5py’s
alloc_time=h5d.ALLOC_TIME_EARLY.
Every chunk exists, holding the fill value, before anything is written, so an unwritten chunk costs a read of fill bytes rather than a miss. On a fixed-shape unfiltered dataset that is also what lets libhdf5 pick its cheapest chunk index — the implicit index, which is no index at all: the chunks are one contiguous run in grid order and a chunk’s address is arithmetic. This builder makes the same choice under the same conditions, so such a dataset is written with no index structure in the file.
Ignored by storage that has no chunk grid to allocate: contiguous, compact and NULL-dataspace datasets.
let file = H5File::create("implicit.h5").unwrap();
let ds = file.new_dataset::<i32>()
.shape([16])
.chunk(&[4])
.early_allocation()
.create("data")
.unwrap();
ds.write_raw(&(0..16i32).collect::<Vec<_>>()).unwrap();Sourcepub fn deflate(self, level: u32) -> Self
pub fn deflate(self, level: u32) -> Self
Enable deflate (gzip) compression with the given level (0-9).
Requires chunked storage (call .chunk() before .create()).
Level 0 = no compression, 9 = maximum compression. Default is 6.
Sourcepub fn shuffle(self) -> Self
pub fn shuffle(self) -> Self
Enable the shuffle filter — H5Pset_shuffle(dcpl), h5py’s
shuffle=True.
Shuffle reorders a chunk’s bytes by their position within an element,
which typically improves how well a compressor behind it does on
numeric data. It is a permutation, not a compressor: on its own it
leaves the chunk exactly as large as it was, which is what
H5Pset_shuffle without a compressor writes. Combine it with
deflate to compress the shuffled stream. Requires
chunked storage.
The element width the filter records is the dataset’s, so a
datatype override is what it follows when the
stored element is not T itself.
Sourcepub fn shuffle_deflate(self, level: u32) -> Self
pub fn shuffle_deflate(self, level: u32) -> Self
Enable shuffle + deflate compression — the same pipeline as
.shuffle().deflate(level).
Shuffle reorders bytes by position within elements before compression, which typically improves compression ratios for numeric data. Requires chunked storage.
Sourcepub fn zstd(self, level: u32) -> Self
pub fn zstd(self, level: u32) -> Self
Enable Zstandard compression with the given level (1-22, default 3).
Requires chunked storage (call .chunk() before .create()).
Sourcepub fn filter_pipeline(self, pipeline: FilterPipeline) -> Self
pub fn filter_pipeline(self, pipeline: FilterPipeline) -> Self
Set a custom filter pipeline for compression.
This takes precedence over deflate and
shuffle_deflate. Requires chunked storage.
Sourcepub fn datatype(self, dt: DatatypeMessage) -> Self
pub fn datatype(self, dt: DatatypeMessage) -> Self
Override the stored element datatype.
By default the dataset is created with the datatype derived from the
Rust type parameter T (H5Type::hdf5_type). Use this to store a
different on-disk datatype than the in-memory element type — for
example a reduced-precision fixed-point type that matches an N-bit
filter (see FilterPipeline::nbit). The element byte size of the
override must equal T::element_size(); the N-bit filter packs the
significant bits within that fixed footprint.
Sourcepub fn committed_type(self, path: &str) -> Self
pub fn committed_type(self, path: &str) -> Self
Build the dataset on the committed (named) datatype at path —
h5py’s dtype=f["name"], H5Dcreate2 with a committed type id.
The dataset does not describe its type: its header stores a pointer to
that object, so the type is defined once and every dataset sharing it
is guaranteed to agree. The type comes from the committed object, so
this supersedes both T and datatype.
The path is resolved at create, which fails when no
committed datatype is there — commit it with
H5File::commit_datatype
first.
let file = H5File::create("committed.h5").unwrap();
file.commit_datatype("temperature", DatatypeMessage::f64_type()).unwrap();
file.new_dataset::<f64>()
.committed_type("temperature")
.shape([4])
.create("readings")
.unwrap();Sourcepub fn object_references(self) -> Self
pub fn object_references(self) -> Self
Store object references — h5py’s h5py.ref_dtype.
The elements are written with
write_object_references and
name objects by path. The element width is the file’s address size, so
the datatype is resolved at create rather than here;
it overrides both T and any datatype call.
let file = H5File::create("refs.h5").unwrap();
file.new_dataset::<i32>().shape([4]).create("target").unwrap();
let refs = file.new_dataset::<u64>()
.object_references()
.shape([1])
.create("refs")
.unwrap();
refs.write_object_references(&["/target"]).unwrap();
file.close().unwrap();Sourcepub fn std_object_references(self) -> Self
pub fn std_object_references(self) -> Self
Store revised object references — the 1.12 H5T_STD_REF.
Same paths and same write_object_references
call as object_references; only the stored
element differs, carrying the reference’s kind alongside the address so
one datatype can hold every reference kind. h5py cannot read it, so
prefer the pre-1.12 form for files h5py will open.
let file = H5File::create("stdrefs.h5").unwrap();
file.new_dataset::<i32>().shape([4]).create("target").unwrap();
let refs = file.new_dataset::<u64>()
.std_object_references()
.shape([1])
.create("refs")
.unwrap();
refs.write_object_references(&["/target"]).unwrap();
file.close().unwrap();Sourcepub fn std_region_references(self) -> Self
pub fn std_region_references(self) -> Self
Store revised region references — H5R_DATASET_REGION2, written into
the same H5T_STD_REF datatype
std_object_references makes.
The elements are written with
write_std_region_references.
What distinguishes them from the pre-1.12
region_references is the element, not the
datatype: a 1.12 element names its own kind, so one dataset of this type
may hold object, region and attribute references together. h5py 3.15
cannot read any of them, so prefer the pre-1.12 form for files h5py will
open.
let file = H5File::options().libver(LibverBound::V112).create("stdregions.h5").unwrap();
file.new_dataset::<i32>().shape([8]).create("target").unwrap();
let refs = file.new_dataset::<u64>()
.std_region_references()
.shape([1])
.create("refs")
.unwrap();
let rows = Selection::Hyperslab {
rank: 1,
form: Hyperslab::Blocks(vec![HyperslabBlock { start: vec![0], end: vec![2] }]),
};
refs.write_std_region_references(&[("/target", rows)]).unwrap();
file.close().unwrap();Sourcepub fn attribute_references(self) -> Self
pub fn attribute_references(self) -> Self
Store attribute references — H5R_ATTR, the one reference kind with no
pre-1.12 form, in the same H5T_STD_REF datatype
std_object_references makes.
The elements are written with
write_attribute_references and
name an object and one of its attributes. h5py 3.15 cannot read them.
let file = H5File::options().libver(LibverBound::V112).create("attrrefs.h5").unwrap();
let target = file.new_dataset::<i32>().shape([4]).create("target").unwrap();
target.new_attr::<i32>().shape([3]).create("note").unwrap()
.write_array(&[7i32, 8, 9]).unwrap();
let refs = file.new_dataset::<u64>()
.attribute_references()
.shape([1])
.create("refs")
.unwrap();
refs.write_attribute_references(&[("/target", "note")]).unwrap();
file.close().unwrap();Sourcepub fn region_references(self) -> Self
pub fn region_references(self) -> Self
Store dataset region references — h5py’s h5py.regionref_dtype.
The elements are written with
write_region_references and name
a dataset plus a selection over it. The element is a global-heap id, so
its width follows the file’s address size and the datatype is resolved
at create rather than here; it overrides both T and
any datatype call.
let file = H5File::create("regions.h5").unwrap();
file.new_dataset::<i32>().shape([8]).create("target").unwrap();
let refs = file.new_dataset::<u64>()
.region_references()
.shape([1])
.create("refs")
.unwrap();
let rows = Selection::Hyperslab {
rank: 1,
form: Hyperslab::Blocks(vec![HyperslabBlock { start: vec![0], end: vec![2] }]),
};
refs.write_region_references(&[("/target", rows)]).unwrap();
file.close().unwrap();Sourcepub fn external(self, files: &[(&str, u64, u64)]) -> Self
pub fn external(self, files: &[(&str, u64, u64)]) -> Self
Keep the raw data in files outside this one — H5Pset_external,
h5py’s external=[(name, offset, size)].
Each entry is a file name, the byte offset in it where that entry’s
region starts, and how many bytes of the dataset it holds; the entries
concatenate, in order, into the dataset’s bytes and together must cover
them. A relative name is resolved against HDF5_EXTFILE_PREFIX the way
libhdf5 resolves it, so the same name reads back through this crate and
through h5py. The storage is contiguous by definition, which rules out
chunk, a filter, compact,
null and either reference kind.
The named files are created on first write and never truncated, so several datasets may own disjoint ranges of one file.
The last entry may take the unlimited size
external_file_list::UNLIMITED
(H5O_EFL_UNLIMITED), which makes it absorb the whole rest of the
dataset however far it grows. A dataset whose
max_shape is unlimited must have one, since no
finite reservation could cover it, and only the first dimension may be
extendible — both H5D__efl_construct’s rules.
let file = H5File::create("ext.h5").unwrap();
let ds = file.new_dataset::<i32>()
.shape([16])
.external(&[("ext.raw", 0, 64)])
.create("data")
.unwrap();
ds.write_raw(&(0..16i32).collect::<Vec<_>>()).unwrap();Sourcepub fn efile_prefix(self, prefix: impl Into<String>) -> Self
pub fn efile_prefix(self, prefix: impl Into<String>) -> Self
H5Pset_efile_prefix on the dapl H5Dcreate2 takes — the directory
the raw data files named by external are created
under, and looked for under on every later write through this handle.
H5D__create builds dset->shared->extfile_prefix from the dapl
(H5Dint.c:1318) and H5D__efl_write joins each slot name against it
with the same single-path H5_combine_path the read side uses
(H5Defl.c:429-431) — so this decides where the bytes land, and the
prefix a later reader names must agree for it to find them.
Measured under libhdf5 1.14.6 and 2.0.0: writing through a dapl that names a directory creates the raw data file there and nowhere else, and a directory that does not exist fails the write outright rather than being created.
It shares DatasetAccess::efile_prefix’s rules, both being
H5D__build_file_prefix: HDF5_EXTFILE_PREFIX shadows this outright
(H5Dint.c:1084-1090), a leading ${ORIGIN} stands for the directory
holding the HDF5 file (:1105-1113), and "." or "" means no prefix
(:1098-1102), which leaves a stored name to resolve against the
process’s current directory.
Ignored by a dataset that names no external files, which has no slot name to join.
Sourcepub fn virtual_mapping(
self,
virtual_selection: Selection,
source_file: &str,
source_dataset: &str,
source_selection: Selection,
) -> Self
pub fn virtual_mapping( self, virtual_selection: Selection, source_file: &str, source_dataset: &str, source_selection: Selection, ) -> Self
Map part of this dataset onto part of a dataset in another file, making
it virtual — H5Pset_virtual, one VirtualLayout[...] = VirtualSource(...) assignment in h5py.
The arguments are H5Pset_virtual’s, in its order: which elements of
this dataset the mapping fills, the file and dataset the data comes
from, and which elements of that source dataset it comes from. Call it
once per mapping; they apply in the order given, which is the order
libhdf5 resolves overlapping ones in.
The source file is named exactly as stored — resolved against
HDF5_VDS_PREFIX, or the virtual dataset’s own directory, when the
file is read — and "." means this file. Nothing is opened or checked
here: a source that does not exist yet is legal, and reads of the
unmapped or unresolvable parts return the fill_value.
A virtual dataset stores nothing of its own, which rules out
chunk, a filter, compact,
null, external and either reference
kind — and makes writing to it an error, since its elements belong to
the source datasets.
An unlimited (H5S_UNLIMITED) selection is written as one: such a
mapping grows with its source, and the dataset’s extent in that
dimension is whatever the sources reachable when it is opened supply
(H5D__virtual_set_extent_unlim). Give it a
max_shape unlimited in the same dimension, as
libhdf5 requires of the dataspace behind one.
A source name may carry libhdf5’s printf-style substitutions: %b
is the block index and %% an escaped literal %. One such mapping
stands for the family of source datasets that fill the successive
blocks of an unlimited virtual selection, so it is legal only with an
unlimited virtual selection over a limited source selection, and the
dataset’s extent stops at the first block whose source is missing.
let file = H5File::create("vds.h5").unwrap();
let ds = file.new_dataset::<i32>()
.shape([16])
.virtual_mapping(Selection::All, "src.h5", "src", Selection::All)
.create("vds")
.unwrap();Sourcepub fn fill_value(self, value: T) -> Self
pub fn fill_value(self, value: T) -> Self
Set a user-defined fill value for unwritten elements.
Without this, datasets use the HDF5 default zero-fill. When set,
the value is written into the dataset’s fill-value message
(fill_defined = 2), so HDF5 readers treat unallocated chunks and
unwritten regions as this value rather than zero.
let file = H5File::create("fv.h5").unwrap();
let ds = file.new_dataset::<f32>()
.shape(&[100])
.fill_value(f32::NAN)
.create("data")
.unwrap();Sourcepub fn fill_time(self, time: FillTime) -> Self
pub fn fill_time(self, time: FillTime) -> Self
Set when the fill value is written into allocated storage —
H5Pset_fill_time.
Without this, a dataset gets FillTime::IfSet
(H5D_CRT_FILL_TIME_DEF), the default every dataset creation
property list carries. FillTime::Never applies to a dataset with
no fill value too: it only stops this writer’s own eager tiling of
the value into newly allocated storage, not the default zero-fill
that storage already has, so its only observable effect is on a
dataset that also calls fill_value.
let file = H5File::create("fv.h5").unwrap();
let ds = file.new_dataset::<f32>()
.shape(&[100])
.fill_value(f32::NAN)
.fill_time(FillTime::Never)
.create("data")
.unwrap();