pub struct H5Dataset { /* private fields */ }Expand description
A handle to an HDF5 dataset, supporting typed read and write operations.
The dataset holds a shared reference to the file’s I/O backend, so it
remains valid even if the originating H5File is
moved or dropped (they share ownership via Rc).
Implementations§
Source§impl H5Dataset
impl H5Dataset
Sourcepub fn total_elements(&self) -> usize
pub fn total_elements(&self) -> usize
Return the total number of elements in the dataset.
Sourcepub fn element_size(&self) -> usize
pub fn element_size(&self) -> usize
Return the size of one element in bytes.
Sourcepub fn datatype(&self) -> Result<DatatypeMessage>
pub fn datatype(&self) -> Result<DatatypeMessage>
Return the element datatype as parsed from the file (read mode only).
Unlike element_size, which reports only the
byte width, this exposes the full datatype: its class (integer vs
floating-point vs string vs compound …), signedness, byte order and
bit precision. Callers that must reconstruct the exact stored type —
for example to map it to a NumPy / Arrow dtype — should use this
instead of inferring a type from the byte width, which cannot
distinguish u8 from i8 (both 1 byte) or i32 from f32 (both 4
bytes).
§Errors
Returns an error if the file is in write mode, or if the dataset can no longer be found in the reader’s metadata.
let file = H5File::open("data.h5").unwrap();
let ds = file.dataset("image").unwrap();
match ds.datatype().unwrap() {
DatatypeMessage::FixedPoint { size, signed, .. } => {
println!("integer: {} bytes, signed={}", size, signed);
}
DatatypeMessage::FloatingPoint { size, .. } => {
println!("float: {} bytes", size);
}
other => println!("other type: {other}"),
}Sourcepub fn chunk_dims(&self) -> Option<Vec<usize>>
pub fn chunk_dims(&self) -> Option<Vec<usize>>
Return the chunk dimensions, if this is a chunked dataset.
Sourcepub fn is_chunked(&self) -> bool
pub fn is_chunked(&self) -> bool
Return whether this is a chunked dataset.
Sourcepub fn attr_names(&self) -> Result<Vec<String>>
pub fn attr_names(&self) -> Result<Vec<String>>
Return the names of all attributes on this dataset (read mode only).
Sourcepub fn attr(&self, attr_name: &str) -> Result<H5Attribute>
pub fn attr(&self, attr_name: &str) -> Result<H5Attribute>
Open an attribute by name (read mode only).
Sourcepub fn new_attr<T: 'static>(&self) -> AttrBuilder<'_, T>
pub fn new_attr<T: 'static>(&self) -> AttrBuilder<'_, T>
Start building a new attribute on this dataset.
Returns a fluent builder. Call .shape(()) for a scalar attribute
and .create("name") to finalize.
§Example
let file = H5File::create("attr.h5").unwrap();
let ds = file.new_dataset::<f32>().shape(&[10]).create("data").unwrap();
let attr = ds.new_attr::<VarLenUnicode>().shape(()).create("units").unwrap();
attr.write_scalar(&VarLenUnicode("meters".to_string())).unwrap();Sourcepub fn write_raw<T: H5Type>(&self, data: &[T]) -> Result<()>
pub fn write_raw<T: H5Type>(&self, data: &[T]) -> Result<()>
Write a typed slice holding the dataset’s whole image.
The slice length must match the total number of elements declared by
the dataset shape. The data is reinterpreted as raw bytes and written
to the file: to the contiguous data block, or — for a chunked dataset —
scattered across its chunk grid, through the filter pipeline if one is
set. To write only part of a dataset, use
write_slice.
§Errors
Returns an error if:
- The file is in read mode.
- The data length does not match the declared shape.
Sourcepub fn write_raw_bytes(&self, bytes: &[u8]) -> Result<()>
pub fn write_raw_bytes(&self, bytes: &[u8]) -> Result<()>
Write the raw byte image of the whole dataset directly.
Takes the same layouts as write_raw: a contiguous
data block, or a chunk grid the image is scattered across.
Unlike write_raw, this is not generic over an
H5Type carrier, so it works for element types that have no matching
Rust primitive — in particular a runtime
CompoundType of arbitrary size set via
DatasetBuilder::datatype. bytes.len() must equal
product(shape) * element_size, where element_size is taken from the
dataset’s on-disk datatype.
let file = H5File::create("c.h5").unwrap();
let ct = CompoundType {
members: vec![
("id".to_string(), i32::hdf5_type(), 0),
("val".to_string(), f64::hdf5_type(), 4),
],
total_size: 12,
};
let ds = file
.new_dataset::<u8>()
.datatype(ct.to_datatype())
.shape(&[2])
.create("records")
.unwrap();
let mut bytes = Vec::new();
bytes.extend_from_slice(&1i32.to_le_bytes());
bytes.extend_from_slice(&2.5f64.to_le_bytes());
bytes.extend_from_slice(&2i32.to_le_bytes());
bytes.extend_from_slice(&3.5f64.to_le_bytes());
ds.write_raw_bytes(&bytes).unwrap();Sourcepub fn write_chunk(&self, chunk_idx: usize, data: &[u8]) -> Result<()>
pub fn write_chunk(&self, chunk_idx: usize, data: &[u8]) -> Result<()>
Write a single chunk to a chunked dataset.
chunk_idx is the linear chunk index (typically the frame number for
streaming datasets). data is the raw byte data for one chunk.
For datasets with two or more unlimited dimensions (v2 B-tree index),
use write_chunk_at instead.
Sourcepub fn write_chunk_raw(
&self,
chunk_idx: usize,
data: &[u8],
filter_mask: u32,
) -> Result<()>
pub fn write_chunk_raw( &self, chunk_idx: usize, data: &[u8], filter_mask: u32, ) -> Result<()>
Write an already-filtered (pre-compressed) chunk verbatim, recording
the caller-supplied filter_mask. The bytes are stored as-is without
running the dataset’s filter pipeline — the HDF5 “direct chunk write”
(H5Dwrite_chunk, formerly H5DOwrite_chunk) operation.
chunk_idx is the linear chunk index (the frame number for streaming
datasets), exactly as for write_chunk. data is
the already-filtered bytes of one chunk — its length is the stored
(compressed) size, not the uncompressed chunk size.
filter_mask is a bitfield: bit i set means filter i of the
dataset’s pipeline was not applied to this chunk and must be skipped
on read. Pass 0 when the full pipeline was already applied upstream (the
common case: a codec plugin handed you compressed frames).
The dataset must be chunked and filtered; an unfiltered chunk index
has no slot to record a stored size or mask. A v2-B-tree-indexed dataset
(two or more unlimited dimensions) has no fixed chunk grid to linearize
against, so address its chunks with
write_chunk_raw_at instead.
§Reading back
Both this crate’s reader and libhdf5/h5py honor the per-chunk
filter_mask: a chunk written with any mask round-trips correctly, with
the reader skipping exactly the filters the mask marks as not applied.
Sourcepub fn write_chunk_at(&self, chunk_coords: &[usize], data: &[u8]) -> Result<()>
pub fn write_chunk_at(&self, chunk_coords: &[usize], data: &[u8]) -> Result<()>
Write a single chunk to a v2-B-tree-indexed dataset, addressed by its chunk-grid coordinates (one per dimension).
This is the entry point for datasets with two or more unlimited
dimensions. The dataset’s logical dimensions are extended to cover
the written chunk. data is the raw bytes of one full chunk.
let file = H5File::create("bt2.h5").unwrap();
let ds = file.new_dataset::<i32>()
.shape(&[0, 0])
.chunk(&[2, 2])
.max_shape(&[None, None])
.create("grid")
.unwrap();
let chunk = [0i32, 1, 2, 3];
let bytes: Vec<u8> = chunk.iter().flat_map(|v| v.to_le_bytes()).collect();
ds.write_chunk_at(&[0, 0], &bytes).unwrap();Sourcepub fn write_chunk_raw_at(
&self,
chunk_coords: &[usize],
data: &[u8],
filter_mask: u32,
) -> Result<()>
pub fn write_chunk_raw_at( &self, chunk_coords: &[usize], data: &[u8], filter_mask: u32, ) -> Result<()>
Write an already-filtered chunk verbatim to a chunked dataset, addressed by its chunk-grid coordinates.
The coordinate-addressed twin of
write_chunk_raw, and the form a
v2-B-tree-indexed dataset needs: with two or more unlimited dimensions
there is no fixed chunk grid for a linear index to mean anything against.
As with write_chunk_at, the dataset’s logical dimensions are extended
to cover the written chunk.
data is the already-filtered bytes of one chunk — its length is the
stored size — and filter_mask bit i set means filter i of the
pipeline was not applied and must be skipped on read. Pass 0 when the
full pipeline already ran upstream.
The dataset must be chunked and filtered; an unfiltered chunk index has no slot to record a stored size or mask.
Sourcepub fn write_chunks_batch(&self, chunks: &[(usize, &[u8])]) -> Result<()>
pub fn write_chunks_batch(&self, chunks: &[(usize, &[u8])]) -> Result<()>
Write multiple chunks in a batch, optionally compressing in parallel.
chunks is a slice of (chunk_index, raw_data) pairs. When a filter
pipeline is configured and the parallel feature is enabled, all
chunks are compressed concurrently via rayon.
Sourcepub fn append<T: H5Type>(&self, data: &[T]) -> Result<()>
pub fn append<T: H5Type>(&self, data: &[T]) -> Result<()>
Append data along the first dimension of a chunked dataset.
data must contain a whole number of “frames” — slices along
dimension 0. For example, if the dataset has shape [N, H, W]
and chunk_dims = [1, H, W], then data.len() must be a
multiple of H * W.
This method writes the necessary chunks and extends the dataset shape automatically.
let file = H5File::create("append.h5").unwrap();
let ds = file.new_dataset::<f64>()
.shape(&[0, 3])
.chunk(&[1, 3])
.max_shape(&[None, Some(3)])
.create("data")
.unwrap();
ds.append(&[1.0, 2.0, 3.0]).unwrap(); // shape becomes [1, 3]
ds.append(&[4.0, 5.0, 6.0, 7.0, 8.0, 9.0]).unwrap(); // shape becomes [3, 3]Sourcepub fn extend(&self, new_dims: &[usize]) -> Result<()>
pub fn extend(&self, new_dims: &[usize]) -> Result<()>
Extend the dimensions of a chunked dataset.
Sourcepub fn set_extent(&self, new_dims: &[usize]) -> Result<()>
pub fn set_extent(&self, new_dims: &[usize]) -> Result<()>
Set the logical extent of a chunked dataset, growing or shrinking any dimension.
Unlike extend, which only grows, this can reduce a
dimension — for example to correct an over-extended frame count
after writing a partial multi-frame chunk. Shrinking changes the
logical dataspace only: data in chunks beyond the new extent stays
in the file but is no longer visible on read, exactly as libhdf5’s
H5Dset_extent behaves. The new extent must not exceed the
dataset’s maximum dimensions.
Sourcepub fn read_slice<T: H5Type>(
&self,
starts: &[usize],
counts: &[usize],
) -> Result<Vec<T>>
pub fn read_slice<T: H5Type>( &self, starts: &[usize], counts: &[usize], ) -> Result<Vec<T>>
Read a slice (hyperslab) of the dataset as a typed vector.
starts and counts define the N-dimensional selection:
starts[d] = first index along dim d, counts[d] = how many elements.
Sourcepub fn write_slice<T: H5Type>(
&self,
starts: &[usize],
counts: &[usize],
data: &[T],
) -> Result<()>
pub fn write_slice<T: H5Type>( &self, starts: &[usize], counts: &[usize], data: &[T], ) -> Result<()>
Write a typed slice to a sub-region of the dataset.
starts and counts define the N-dimensional selection, which must lie
inside the dataset’s current extent.
Works for both contiguous and chunked datasets. For a chunked dataset only the chunks the selection touches are rewritten — a partially covered chunk is read back, patched, and written again, so updating one row of an appendable dataset costs the chunks that row crosses rather than the whole dataset. Elements of a touched chunk that the selection does not cover keep their stored value, or the dataset’s fill value if the chunk did not exist yet.
Sourcepub fn write_vlen_strings_slice(
&self,
start: usize,
strings: &[&str],
) -> Result<()>
pub fn write_vlen_strings_slice( &self, start: usize, strings: &[&str], ) -> Result<()>
Replace elements start .. start + strings.len() of a 1-D
variable-length string dataset.
The extent and every element outside the range are left alone, and the cost is the new strings plus the chunks holding their references — not the column. The dataset’s character set is enforced: a non-ASCII replacement in a dataset that declares ASCII is rejected rather than stored under a datatype that misdescribes it.
The global heap objects the replaced references pointed at are freed — the same reclaim libhdf5 performs on an overwrite — so updating one element repeatedly reuses space rather than growing the file. A collection emptied by the update returns its block to the allocator. Under SWMR nothing is freed, because a reader may still be following those references.
let file = H5File::open_rw("meta.h5").unwrap();
let ds = file.dataset_writer("notes").unwrap();
ds.write_vlen_strings_slice(42, &["replacement"]).unwrap();
file.close().unwrap();Sourcepub fn read_vlen_strings(&self) -> Result<Vec<String>>
pub fn read_vlen_strings(&self) -> Result<Vec<String>>
Read variable-length strings from a dataset.
This handles h5py-style vlen string datasets that store strings as global heap references. Returns one String per element.
Sourcepub fn read_vlen_bytes(&self) -> Result<Vec<Vec<u8>>>
pub fn read_vlen_bytes(&self) -> Result<Vec<Vec<u8>>>
Read variable-length byte arrays from a dataset.
This handles vlen byte-array datasets (a vlen sequence of u8, e.g.
those written by write_vlen_bytes)
that store each element as a global heap reference. Returns one
Vec<u8> per element.
Sourcepub fn read_strings(&self) -> Result<Vec<String>>
pub fn read_strings(&self) -> Result<Vec<String>>
Read a string dataset, fixed-width or variable-length, as one String
per element.
The width of a FixedString dataset is whatever the file says, so a
24-byte label column and a 100-byte one are read by the same call. The
padding rule the datatype declares decides where each element ends —
null-terminated (0), null-padded (1) or space-padded (2) — and its
character set decides how the remaining bytes are decoded: ASCII (0)
requires 7-bit bytes, UTF-8 (1) requires valid UTF-8. An element that
violates either is an error naming the element, not a silent
substitution; read_strings_lossy is the
call that accepts such a file, replacing what it cannot decode.
let file = H5File::open("labels.h5").unwrap();
let labels = file.dataset("names").unwrap().read_strings().unwrap();Sourcepub fn read_strings_lossy(&self) -> Result<Vec<String>>
pub fn read_strings_lossy(&self) -> Result<Vec<String>>
read_strings, but bytes that do not decode
under the dataset’s character set become U+FFFD instead of an error.
Producers do mislabel the character set — a file that declares ASCII while storing Latin-1 or UTF-8 bytes reads here and not there.
Sourcepub fn read_raw<T: H5Type>(&self) -> Result<Vec<T>>
pub fn read_raw<T: H5Type>(&self) -> Result<Vec<T>>
Read the entire dataset as a typed vector.
The raw bytes are read from the file and reinterpreted as T. The
caller must ensure that T matches the datatype used when the dataset
was written.
§Errors
Returns an error if:
- The file is in write mode.
- The raw data size is not a multiple of
T::element_size().
Sourcepub fn read_raw_bytes(&self) -> Result<Vec<u8>>
pub fn read_raw_bytes(&self) -> Result<Vec<u8>>
Read the raw byte image of a dataset without an H5Type carrier.
The counterpart to write_raw_bytes: returns
the element bytes verbatim regardless of the on-disk element type, so a
runtime CompoundType whose records have
no matching Rust primitive can be read back and decoded by the caller.
Sourcepub fn read_raw_into<T: H5Type>(&self, out: &mut [T]) -> Result<()>
pub fn read_raw_into<T: H5Type>(&self, out: &mut [T]) -> Result<()>
Read the whole dataset into a caller-provided buffer, with no allocation.
out must have exactly product(dims) elements (the dataset’s element
count) and T::element_size() must match the dataset’s on-disk element
size, otherwise an error is returned and out is left unspecified. The
zero-copy counterpart of read_raw: the bytes are read
straight into out rather than into a fresh Vec, so a pinned /
page-locked host buffer can be filled in one pass and DMA’d to a GPU
without the extra staging copy a read_raw + copy-into-pinned would
incur. Works for every layout (contiguous, compact, and chunked under
any index); for chunked data each decoded chunk is scattered directly
into out.
let file = H5File::open("data.h5").unwrap();
let ds = file.dataset("frames").unwrap();
let n: usize = ds.shape().iter().product();
let mut buf = vec![0u16; n]; // or a pinned host allocation
ds.read_raw_into(&mut buf).unwrap();Sourcepub fn read_slice_into<T: H5Type>(
&self,
out: &mut [T],
starts: &[usize],
counts: &[usize],
) -> Result<()>
pub fn read_slice_into<T: H5Type>( &self, out: &mut [T], starts: &[usize], counts: &[usize], ) -> Result<()>
Read a hyperslab into a caller-provided buffer, with no allocation.
out must have exactly product(counts) elements and
T::element_size() must match the dataset’s element size. The zero-copy
counterpart of read_slice and the slice analogue of
read_raw_into: only chunks overlapping the
selection are read, and the selected bytes land directly in out — the
entry point for reading one frame / block straight into a pinned host
buffer for an H2D transfer.
let file = H5File::open("vol.h5").unwrap();
let ds = file.dataset("vol").unwrap(); // shape [nz, ny, nx]
let (ny, nx) = (ds.shape()[1], ds.shape()[2]);
let mut frame = vec![0f32; ny * nx]; // or a pinned host allocation
ds.read_slice_into(&mut frame, &[5, 0, 0], &[1, ny, nx]).unwrap();