pub struct H5Dataset { /* private fields */ }Expand description
A handle to an HDF5 dataset, supporting typed read and write operations.
The dataset holds a shared reference to the file’s I/O backend, so it
remains valid even if the originating H5File is
moved or dropped (they share ownership via Rc).
Implementations§
Source§impl H5Dataset
impl H5Dataset
Sourcepub fn total_elements(&self) -> usize
pub fn total_elements(&self) -> usize
Return the total number of elements in the dataset.
0 for a NULL dataspace (is_null) — unlike a scalar,
whose shape() is the same empty Vec but which holds exactly one
element, so shape().iter().product() cannot be used here.
Sourcepub fn element_size(&self) -> usize
pub fn element_size(&self) -> usize
Return the size of one element in bytes.
Sourcepub fn is_null(&self) -> bool
pub fn is_null(&self) -> bool
Return whether this dataset has the NULL dataspace: no elements at
all, distinct from a scalar dataset (rank 0, exactly one element) —
both report the same empty shape. See
DatasetBuilder::null.
Sourcepub fn datatype(&self) -> Result<DatatypeMessage>
pub fn datatype(&self) -> Result<DatatypeMessage>
Return the element datatype as parsed from the file (read mode only).
Unlike element_size, which reports only the
byte width, this exposes the full datatype: its class (integer vs
floating-point vs string vs compound …), signedness, byte order and
bit precision. Callers that must reconstruct the exact stored type —
for example to map it to a NumPy / Arrow dtype — should use this
instead of inferring a type from the byte width, which cannot
distinguish u8 from i8 (both 1 byte) or i32 from f32 (both 4
bytes).
§Errors
Returns an error if the file is in write mode, or if the dataset can no longer be found in the reader’s metadata.
let file = H5File::open("data.h5").unwrap();
let ds = file.dataset("image").unwrap();
match ds.datatype().unwrap() {
DatatypeMessage::FixedPoint { size, signed, .. } => {
println!("integer: {} bytes, signed={}", size, signed);
}
DatatypeMessage::FloatingPoint { size, .. } => {
println!("float: {} bytes", size);
}
other => println!("other type: {other}"),
}Sourcepub fn chunk_dims(&self) -> Option<Vec<usize>>
pub fn chunk_dims(&self) -> Option<Vec<usize>>
Return the chunk dimensions, if this is a chunked dataset.
Sourcepub fn is_chunked(&self) -> bool
pub fn is_chunked(&self) -> bool
Return whether this is a chunked dataset.
Sourcepub fn storage_layout(&self) -> Result<StorageLayout>
pub fn storage_layout(&self) -> Result<StorageLayout>
Return the dataset’s storage layout class (read mode only).
§Errors
Returns an error if the file is in write mode, or if the dataset can no longer be found in the reader’s metadata.
Sourcepub fn chunk_index(&self) -> Result<Option<ChunkIndex>>
pub fn chunk_index(&self) -> Result<Option<ChunkIndex>>
Return the chunk index structure this dataset’s layout uses, or
None for a dataset that is not chunked (read mode only).
§Errors
Returns an error if the file is in write mode, or if the dataset can no longer be found in the reader’s metadata.
Sourcepub fn filters(&self) -> Result<Vec<Filter>>
pub fn filters(&self) -> Result<Vec<Filter>>
Return this dataset’s filter pipeline (read mode only), in application order. Empty when the dataset has no filter pipeline message at all — an unfiltered dataset, not an error.
§Errors
Returns an error if the file is in write mode, or if the dataset can no longer be found in the reader’s metadata.
Sourcepub fn fill_value(&self) -> Result<FillValue>
pub fn fill_value(&self) -> Result<FillValue>
Return this dataset’s fill-value state (read mode only).
§Errors
Returns an error if the file is in write mode, or if the dataset can no longer be found in the reader’s metadata.
Sourcepub fn fill_time(&self) -> Result<FillTime>
pub fn fill_time(&self) -> Result<FillTime>
Return when this dataset’s fill value is written into allocated
storage (read mode only) — H5Pget_fill_time.
§Errors
Returns an error if the file is in write mode, or if the dataset can no longer be found in the reader’s metadata.
Sourcepub fn alloc_time(&self) -> Result<AllocTime>
pub fn alloc_time(&self) -> Result<AllocTime>
Return when this dataset’s raw-data storage is allocated (read mode
only) — H5Pget_alloc_time.
§Errors
Returns an error if the file is in write mode, or if the dataset can no longer be found in the reader’s metadata.
Sourcepub fn external_files(&self) -> Result<Vec<ExternalFileSegment>>
pub fn external_files(&self) -> Result<Vec<ExternalFileSegment>>
Return this dataset’s external raw-data file segments (read mode only), in the order the dataset’s logical byte range concatenates them. Empty for a dataset whose data lives in this file.
§Errors
Returns an error if the file is in write mode, or if the dataset can no longer be found in the reader’s metadata.
Sourcepub fn max_shape(&self) -> Result<Vec<Option<usize>>>
pub fn max_shape(&self) -> Result<Vec<Option<usize>>>
Return this dataset’s maximum dimension sizes (read mode only):
None in a dimension marks that axis unlimited. A dataset with no
maximum-dimensions message reports its current shape (max == current
— the upstream convention for a fixed-extent dataset).
§Errors
Returns an error if the file is in write mode, or if the dataset can no longer be found in the reader’s metadata.
Sourcepub fn virtual_mappings(&self) -> Result<Vec<VirtualMapping>>
pub fn virtual_mappings(&self) -> Result<Vec<VirtualMapping>>
Return this dataset’s virtual-dataset source/virtual mappings (read mode only), in on-disk order. Empty for any dataset whose layout is not virtual, and for a virtual dataset that has no mappings yet.
§Errors
Returns an error if the file is in write mode, or if the dataset can no longer be found in the reader’s metadata.
Sourcepub fn attr_names(&self) -> Result<Vec<String>>
pub fn attr_names(&self) -> Result<Vec<String>>
Return the names of all attributes on this dataset (read mode only).
Sourcepub fn attr_unreadable_reason(&self, attr_name: &str) -> Result<Option<String>>
pub fn attr_unreadable_reason(&self, attr_name: &str) -> Result<Option<String>>
Why the attribute attr_name on this dataset cannot be read, or None
when it can be.
An attribute whose message this crate cannot decode is still listed by
attr_names — the object header carries it — and
this says what stands in the way. Opening it through
attr fails with the same text.
Sourcepub fn attrs_unreadable_reason(&self) -> Result<Option<String>>
pub fn attrs_unreadable_reason(&self) -> Result<Option<String>>
Why this dataset’s attribute set cannot be listed, or None when it
can be.
The object-scope counterpart of
attr_unreadable_reason. A dense
attribute set is indexed by name hash, so a heap or index that will not
read yields no names to hang a per-attribute reason on;
attr_names then returns the failure rather than a
short list, and this reports it without an attribute name.
Sourcepub fn attr_storage(&self) -> Result<AttributeStorage>
pub fn attr_storage(&self) -> Result<AttributeStorage>
This dataset’s own compact-vs-dense attribute storage — the
equivalent of h5py.h5o.get_info(did.id).meta_size.attr.index_size
being nonzero (read mode only).
Sourcepub fn header_attr_count(&self) -> Result<u64>
pub fn header_attr_count(&self) -> Result<u64>
This dataset’s own object-header attribute count — the equivalent of
h5py.h5o.get_info(did.id).num_attrs (read mode only).
Sourcepub fn attr(&self, attr_name: &str) -> Result<H5Attribute>
pub fn attr(&self, attr_name: &str) -> Result<H5Attribute>
Open an attribute by name (read mode only).
Sourcepub fn new_attr<T: 'static>(&self) -> AttrBuilder<'_, T>
pub fn new_attr<T: 'static>(&self) -> AttrBuilder<'_, T>
Start building a new attribute on this dataset.
Returns a fluent builder. Call .shape(()) for a scalar attribute
and .create("name") to finalize.
§Example
let file = H5File::create("attr.h5").unwrap();
let ds = file.new_dataset::<f32>().shape(&[10]).create("data").unwrap();
let attr = ds.new_attr::<VarLenUnicode>().shape(()).create("units").unwrap();
attr.write_scalar(&VarLenUnicode("meters".to_string())).unwrap();Sourcepub fn write_raw<T: H5Type>(&self, data: &[T]) -> Result<()>
pub fn write_raw<T: H5Type>(&self, data: &[T]) -> Result<()>
Write a typed slice holding the dataset’s whole image.
The slice length must match the total number of elements declared by
the dataset shape. The data is reinterpreted as raw bytes and written
to the file: to the contiguous data block, or — for a chunked dataset —
scattered across its chunk grid, through the filter pipeline if one is
set. To write only part of a dataset, use
write_slice.
§Errors
Returns an error if:
- The file is in read mode.
- The data length does not match the declared shape.
Sourcepub fn write_raw_bytes(&self, bytes: &[u8]) -> Result<()>
pub fn write_raw_bytes(&self, bytes: &[u8]) -> Result<()>
Write the raw byte image of the whole dataset directly.
Takes the same layouts as write_raw: a contiguous
data block, or a chunk grid the image is scattered across.
Unlike write_raw, this is not generic over an
H5Type carrier, so it works for element types that have no matching
Rust primitive — in particular a runtime
CompoundType of arbitrary size set via
DatasetBuilder::datatype. bytes.len() must equal
product(shape) * element_size, where element_size is taken from the
dataset’s on-disk datatype.
let file = H5File::create("c.h5").unwrap();
let ct = CompoundType {
members: vec![
("id".to_string(), i32::hdf5_type(), 0),
("val".to_string(), f64::hdf5_type(), 4),
],
total_size: 12,
};
let ds = file
.new_dataset::<u8>()
.datatype(ct.to_datatype())
.shape(&[2])
.create("records")
.unwrap();
let mut bytes = Vec::new();
bytes.extend_from_slice(&1i32.to_le_bytes());
bytes.extend_from_slice(&2.5f64.to_le_bytes());
bytes.extend_from_slice(&2i32.to_le_bytes());
bytes.extend_from_slice(&3.5f64.to_le_bytes());
ds.write_raw_bytes(&bytes).unwrap();Sourcepub fn write_chunk(&self, chunk_idx: usize, data: &[u8]) -> Result<()>
pub fn write_chunk(&self, chunk_idx: usize, data: &[u8]) -> Result<()>
Write a single chunk to a chunked dataset.
chunk_idx is the linear chunk index (typically the frame number for
streaming datasets). data is the raw byte data for one chunk.
For datasets with two or more unlimited dimensions (v2 B-tree index),
use write_chunk_at instead.
Sourcepub fn write_chunk_raw(
&self,
chunk_idx: usize,
data: &[u8],
filter_mask: u32,
) -> Result<()>
pub fn write_chunk_raw( &self, chunk_idx: usize, data: &[u8], filter_mask: u32, ) -> Result<()>
Write an already-filtered (pre-compressed) chunk verbatim, recording
the caller-supplied filter_mask. The bytes are stored as-is without
running the dataset’s filter pipeline — the HDF5 “direct chunk write”
(H5Dwrite_chunk, formerly H5DOwrite_chunk) operation.
chunk_idx is the linear chunk index (the frame number for streaming
datasets), exactly as for write_chunk. data is
the already-filtered bytes of one chunk — its length is the stored
(compressed) size, not the uncompressed chunk size.
filter_mask is a bitfield: bit i set means filter i of the
dataset’s pipeline was not applied to this chunk and must be skipped
on read. Pass 0 when the full pipeline was already applied upstream (the
common case: a codec plugin handed you compressed frames).
The dataset must be chunked and filtered; an unfiltered chunk index
has no slot to record a stored size or mask. A v2-B-tree-indexed dataset
(two or more unlimited dimensions) has no fixed chunk grid to linearize
against, so address its chunks with
write_chunk_raw_at instead.
§Reading back
Both this crate’s reader and libhdf5/h5py honor the per-chunk
filter_mask: a chunk written with any mask round-trips correctly, with
the reader skipping exactly the filters the mask marks as not applied.
Sourcepub fn write_chunk_at(&self, chunk_coords: &[usize], data: &[u8]) -> Result<()>
pub fn write_chunk_at(&self, chunk_coords: &[usize], data: &[u8]) -> Result<()>
Write a single chunk to a v2-B-tree-indexed dataset, addressed by its chunk-grid coordinates (one per dimension).
This is the entry point for datasets with two or more unlimited
dimensions. The dataset’s logical dimensions are extended to cover
the written chunk. data is the raw bytes of one full chunk.
let file = H5File::create("bt2.h5").unwrap();
let ds = file.new_dataset::<i32>()
.shape(&[0, 0])
.chunk(&[2, 2])
.max_shape(&[None, None])
.create("grid")
.unwrap();
let chunk = [0i32, 1, 2, 3];
let bytes: Vec<u8> = chunk.iter().flat_map(|v| v.to_le_bytes()).collect();
ds.write_chunk_at(&[0, 0], &bytes).unwrap();Sourcepub fn write_chunk_raw_at(
&self,
chunk_coords: &[usize],
data: &[u8],
filter_mask: u32,
) -> Result<()>
pub fn write_chunk_raw_at( &self, chunk_coords: &[usize], data: &[u8], filter_mask: u32, ) -> Result<()>
Write an already-filtered chunk verbatim to a chunked dataset, addressed by its chunk-grid coordinates.
The coordinate-addressed twin of
write_chunk_raw, and the form a
v2-B-tree-indexed dataset needs: with two or more unlimited dimensions
there is no fixed chunk grid for a linear index to mean anything against.
As with write_chunk_at, the dataset’s logical dimensions are extended
to cover the written chunk.
data is the already-filtered bytes of one chunk — its length is the
stored size — and filter_mask bit i set means filter i of the
pipeline was not applied and must be skipped on read. Pass 0 when the
full pipeline already ran upstream.
The dataset must be chunked and filtered; an unfiltered chunk index has no slot to record a stored size or mask.
Sourcepub fn write_chunks_batch(&self, chunks: &[(usize, &[u8])]) -> Result<()>
pub fn write_chunks_batch(&self, chunks: &[(usize, &[u8])]) -> Result<()>
Write multiple chunks in a batch, optionally compressing in parallel.
chunks is a slice of (chunk_index, raw_data) pairs. When a filter
pipeline is configured and the parallel feature is enabled, all
chunks are compressed concurrently via rayon.
Sourcepub fn append<T: H5Type>(&self, data: &[T]) -> Result<()>
pub fn append<T: H5Type>(&self, data: &[T]) -> Result<()>
Append data along the first dimension of a chunked dataset.
data must contain a whole number of “frames” — slices along
dimension 0. For example, if the dataset has shape [N, H, W]
and chunk_dims = [1, H, W], then data.len() must be a
multiple of H * W.
This method writes the necessary chunks and extends the dataset shape automatically.
let file = H5File::create("append.h5").unwrap();
let ds = file.new_dataset::<f64>()
.shape(&[0, 3])
.chunk(&[1, 3])
.max_shape(&[None, Some(3)])
.create("data")
.unwrap();
ds.append(&[1.0, 2.0, 3.0]).unwrap(); // shape becomes [1, 3]
ds.append(&[4.0, 5.0, 6.0, 7.0, 8.0, 9.0]).unwrap(); // shape becomes [3, 3]Sourcepub fn extend(&self, new_dims: &[usize]) -> Result<()>
pub fn extend(&self, new_dims: &[usize]) -> Result<()>
Extend the dimensions of a chunked dataset.
Sourcepub fn set_extent(&self, new_dims: &[usize]) -> Result<()>
pub fn set_extent(&self, new_dims: &[usize]) -> Result<()>
Set the logical extent of a chunked dataset, growing or shrinking any dimension.
Unlike extend, which only grows, this can reduce a
dimension — for example to correct an over-extended frame count
after writing a partial multi-frame chunk. Shrinking prunes the
stored chunks the way libhdf5’s H5Dset_extent does: a chunk
entirely beyond the new extent is removed from the chunk index and
its storage freed for reuse, and a chunk the new extent cuts
through has its out-of-extent region overwritten with the fill
value — so growing the extent back exposes fill values, not the
old data. The new extent must not exceed the dataset’s maximum
dimensions.
Sourcepub fn read_slice<T: H5Type>(
&self,
starts: &[usize],
counts: &[usize],
) -> Result<Vec<T>>
pub fn read_slice<T: H5Type>( &self, starts: &[usize], counts: &[usize], ) -> Result<Vec<T>>
Read a slice (hyperslab) of the dataset as a typed vector.
starts and counts define the N-dimensional selection:
starts[d] = first index along dim d, counts[d] = how many elements.
Sourcepub fn read_hyperslab<T: H5Type>(
&self,
start: &[usize],
stride: &[usize],
count: &[usize],
block: &[usize],
) -> Result<Vec<T>>
pub fn read_hyperslab<T: H5Type>( &self, start: &[usize], stride: &[usize], count: &[usize], block: &[usize], ) -> Result<Vec<T>>
Read a strided hyperslab as a typed vector — h5py’s stepped slicing
(ds[a:b:s]) or the general start/stride/count/block form of
H5Sselect_hyperslab.
One entry per dimension: start[d] is the first index, stride[d]
the spacing between selected blocks (all-1 is the same selection
read_slice reads), count[d] how many blocks,
and block[d] how many contiguous elements each block covers. The
returned vector is row-major over count[d] * block[d] per
dimension — exactly the shape h5py’s stepped slicing produces.
let file = H5File::open("data.h5").unwrap();
let ds = file.dataset("series").unwrap(); // shape [100]
// Python: ds[0:100:2] — every other element.
let evens: Vec<f64> = ds.read_hyperslab(&[0], &[2], &[50], &[1]).unwrap();Sourcepub fn read_points<T: H5Type>(&self, points: &[Vec<usize>]) -> Result<Vec<T>>
pub fn read_points<T: H5Type>(&self, points: &[Vec<usize>]) -> Result<Vec<T>>
Read a list of coordinates in one call, as a typed vector — h5py fancy indexing with a coordinate list.
points[i] is a coordinate with one entry per dimension. The
returned vector holds one element per point, in the same order as
points, regardless of the dataset’s rank.
let file = H5File::open("data.h5").unwrap();
let ds = file.dataset("grid").unwrap(); // shape [10, 10]
// Python: ds[np.array([[0, 0], [3, 4], [9, 9]])]
let picked: Vec<f64> = ds.read_points(&[vec![0, 0], vec![3, 4], vec![9, 9]]).unwrap();Sourcepub fn read_chunk_raw_at(
&self,
chunk_coords: &[usize],
) -> Result<(Vec<u8>, u32)>
pub fn read_chunk_raw_at( &self, chunk_coords: &[usize], ) -> Result<(Vec<u8>, u32)>
Read one chunk’s raw (still-filtered) bytes and its filter mask,
addressed by chunk-grid coordinates — the read half of
write_chunk_raw_at and the HDF5 “direct
chunk read” (H5Dread_chunk, formerly H5DOread_chunk; h5py’s
Dataset.id.read_direct_chunk).
The bytes are exactly what is stored on disk: filtered/compressed if
the dataset has a filter pipeline, with no decompression applied. The
returned u32 is the chunk’s filter mask: bit i set means filter
i of the pipeline was not applied to this particular chunk and
must be skipped when reversing it.
Err if the dataset is not chunked, chunk_coords has the wrong
rank, or the chunk at those coordinates has never been written.
let file = H5File::open("data.h5").unwrap();
let ds = file.dataset("frames").unwrap();
let (raw, filter_mask) = ds.read_chunk_raw_at(&[0, 0]).unwrap();Sourcepub fn write_slice<T: H5Type>(
&self,
starts: &[usize],
counts: &[usize],
data: &[T],
) -> Result<()>
pub fn write_slice<T: H5Type>( &self, starts: &[usize], counts: &[usize], data: &[T], ) -> Result<()>
Write a typed slice to a sub-region of the dataset.
starts and counts define the N-dimensional selection, which must lie
inside the dataset’s current extent.
Works for both contiguous and chunked datasets. For a chunked dataset only the chunks the selection touches are rewritten — a partially covered chunk is read back, patched, and written again, so updating one row of an appendable dataset costs the chunks that row crosses rather than the whole dataset. Elements of a touched chunk that the selection does not cover keep their stored value, or the dataset’s fill value if the chunk did not exist yet.
Sourcepub fn write_vlen_strings_slice(
&self,
start: usize,
strings: &[&str],
) -> Result<()>
pub fn write_vlen_strings_slice( &self, start: usize, strings: &[&str], ) -> Result<()>
Replace elements start .. start + strings.len() of a 1-D
variable-length string dataset.
The extent and every element outside the range are left alone, and the cost is the new strings plus the chunks holding their references — not the column. The dataset’s character set is enforced: a non-ASCII replacement in a dataset that declares ASCII is rejected rather than stored under a datatype that misdescribes it.
The global heap objects the replaced references pointed at are freed — the same reclaim libhdf5 performs on an overwrite — so updating one element repeatedly reuses space rather than growing the file. A collection emptied by the update returns its block to the allocator. Under SWMR nothing is freed, because a reader may still be following those references.
let file = H5File::open_rw("meta.h5").unwrap();
let ds = file.dataset_writer("notes").unwrap();
ds.write_vlen_strings_slice(42, &["replacement"]).unwrap();
file.close().unwrap();Sourcepub fn read_vlen_strings(&self) -> Result<Vec<String>>
pub fn read_vlen_strings(&self) -> Result<Vec<String>>
Read variable-length strings from a dataset.
This handles h5py-style vlen string datasets that store strings as global heap references. Returns one String per element.
Sourcepub fn read_vlen_bytes(&self) -> Result<Vec<Vec<u8>>>
pub fn read_vlen_bytes(&self) -> Result<Vec<Vec<u8>>>
Read variable-length byte arrays from a dataset.
This handles vlen byte-array datasets (a vlen sequence of u8, e.g.
those written by write_vlen_bytes)
that store each element as a global heap reference. Returns one
Vec<u8> per element.
Sourcepub fn write_object_references(&self, paths: &[&str]) -> Result<()>
pub fn write_object_references(&self, paths: &[&str]) -> Result<()>
Write object references naming paths into elements 0..paths.len()
— h5py’s refs[i] = f['/target'].ref.
The dataset must have been created with
object_references. A path names
a dataset or a group (/ is the root group) and must already exist;
what reaches the file is the target’s object header address, which is
assigned when the file is finalized. Elements left unwritten read back
as null references.
Sourcepub fn set_scale(&self, name: Option<&str>) -> Result<()>
pub fn set_scale(&self, name: Option<&str>) -> Result<()>
Mark this dataset as a dimension scale — H5DSset_scale, h5py’s
ds.make_scale(name).
Writes the CLASS attribute as the fixed-length null-terminated
string DIMENSION_SCALE (16 bytes, the width H5DSis_scale checks
for) and, when name is given, NAME the same way. A dataset that
has scales attached to it cannot become one. Write mode only.
let file = H5File::create("scales.h5").unwrap();
let x = file.new_dataset::<f64>().shape([4]).create("x").unwrap();
x.write_raw(&[0.0, 0.5, 1.0, 1.5]).unwrap();
x.set_scale(Some("x")).unwrap();Sourcepub fn attach_scale(&self, axis: usize, scale: &H5Dataset) -> Result<()>
pub fn attach_scale(&self, axis: usize, scale: &H5Dataset) -> Result<()>
Attach scale as a dimension scale of this dataset’s axis —
H5DSattach_scale, h5py’s ds.dims[axis].attach_scale(scale).
Records the attachment in this dataset’s DIMENSION_LIST and the
scale’s REFERENCE_LIST, and marks scale as a dimension scale if it
is not one yet. An axis may carry several scales: each attach appends.
Attaching a scale already on that axis changes nothing. Both datasets
must belong to this file, in write mode; a scale cannot have scales of
its own, a dataset that is a scale cannot have scales attached, and
axis must be below the rank (a scalar dataset counts as rank 1).
let file = H5File::create("scales.h5").unwrap();
let data = file.new_dataset::<i32>().shape([2, 3]).create("data").unwrap();
let x = file.new_dataset::<f64>().shape([3]).create("x").unwrap();
x.set_scale(Some("x")).unwrap();
data.attach_scale(1, &x).unwrap();Sourcepub fn write_region_references(
&self,
targets: &[(&str, Selection)],
) -> Result<()>
pub fn write_region_references( &self, targets: &[(&str, Selection)], ) -> Result<()>
Write region references over targets into elements
0..targets.len() — h5py’s refs[i] = f['/target'].regionref[0:3].
The dataset must have been created with
region_references. Each target is
the path of an existing dataset and a Selection over it, which
must fit that dataset’s extent — the rule H5Rcreate applies. What
reaches the file is a global-heap object holding the target’s object
header address (assigned when the file is finalized) and the serialized
selection. Elements left unwritten read back as null references.
let file = H5File::create("regions.h5").unwrap();
file.new_dataset::<i32>().shape([4, 6]).create("m").unwrap();
let refs = file.new_dataset::<u64>()
.region_references()
.shape([1])
.create("refs")
.unwrap();
let points = Selection::Points(PointSelection {
rank: 2,
points: vec![vec![0, 1], vec![3, 5]],
});
refs.write_region_references(&[("/m", points)]).unwrap();
file.close().unwrap();Sourcepub fn write_std_region_references(
&self,
targets: &[(&str, Selection)],
) -> Result<()>
pub fn write_std_region_references( &self, targets: &[(&str, Selection)], ) -> Result<()>
Write revised region references over targets into elements
0..targets.len() — H5Rcreate_region plus H5Dwrite of an
H5T_STD_REF dataset.
The dataset must have been created with
std_region_references or one
of its two siblings, which make the same datatype. Each target is the
path of an existing dataset and a Selection over it, which must
fit that dataset’s extent. What reaches the file is a global-heap blob
holding the target’s object header address (assigned when the file is
finalized) and the serialized selection, and an element carrying the
blob’s id and its byte count. Elements left unwritten read back as null
references.
Sourcepub fn write_attribute_references(&self, targets: &[(&str, &str)]) -> Result<()>
pub fn write_attribute_references(&self, targets: &[(&str, &str)]) -> Result<()>
Write attribute references naming targets into elements
0..targets.len() — H5Rcreate_attr plus H5Dwrite of an
H5T_STD_REF dataset.
Each target is the path of an existing object — a dataset, a group, or
/ for the root group — and the name of an attribute it already
carries. There is no pre-1.12 form of this reference kind, so the
dataset must have been created with
attribute_references or one of
its two siblings. Elements left unwritten read back as null references.
Sourcepub fn read_references(&self) -> Result<Vec<Reference>>
pub fn read_references(&self) -> Result<Vec<Reference>>
Read a reference dataset’s elements, each resolved to the object it names.
Every reference kind is read: the pre-1.12 pair h5py writes —
Reference (an object header address) and RegionReference (a heap id
whose heap object holds the target plus a serialized selection) — and
the 1.12 H5T_STD_REF trio, H5R_OBJECT2, H5R_DATASET_REGION2 and
H5R_ATTR. An object reference comes back as Reference::Object
carrying the target’s path, a region reference as
Reference::Region, whose bounds is the
selection’s bounding box — libhdf5’s H5Sget_select_bounds — and an
attribute reference as Reference::Attr, which adds the attribute’s
name.
A 1.12 reference written into a file other than its target’s carries
that file’s name, and Reference::file reports it; the path is then
a path inside that file, resolved by opening it under the name the
reference carries, and None when nothing is there.
let file = H5File::open("refs.h5").unwrap();
for r in file.dataset("refs").unwrap().read_references().unwrap() {
println!("{:?} {:?}", r.path(), r.bounds());
}Sourcepub fn read_strings(&self) -> Result<Vec<String>>
pub fn read_strings(&self) -> Result<Vec<String>>
Read a string dataset, fixed-width or variable-length, as one String
per element.
The width of a FixedString dataset is whatever the file says, so a
24-byte label column and a 100-byte one are read by the same call. The
padding rule the datatype declares decides where each element ends —
null-terminated (0), null-padded (1) or space-padded (2) — and its
character set decides how the remaining bytes are decoded: ASCII (0)
requires 7-bit bytes, UTF-8 (1) requires valid UTF-8. An element that
violates either is an error naming the element, not a silent
substitution; read_strings_lossy is the
call that accepts such a file, replacing what it cannot decode.
let file = H5File::open("labels.h5").unwrap();
let labels = file.dataset("names").unwrap().read_strings().unwrap();Sourcepub fn read_strings_lossy(&self) -> Result<Vec<String>>
pub fn read_strings_lossy(&self) -> Result<Vec<String>>
read_strings, but bytes that do not decode
under the dataset’s character set become U+FFFD instead of an error.
Producers do mislabel the character set — a file that declares ASCII while storing Latin-1 or UTF-8 bytes reads here and not there.
Sourcepub fn read_raw<T: H5Type>(&self) -> Result<Vec<T>>
pub fn read_raw<T: H5Type>(&self) -> Result<Vec<T>>
Read the entire dataset as a typed vector.
The raw bytes are read from the file and reinterpreted as T. The
caller must ensure that T matches the datatype used when the dataset
was written.
§Errors
Returns an error if:
- The file is in write mode.
- The raw data size is not a multiple of
T::element_size().
Sourcepub fn read_raw_bytes(&self) -> Result<Vec<u8>>
pub fn read_raw_bytes(&self) -> Result<Vec<u8>>
Read the raw byte image of a dataset without an H5Type carrier.
The counterpart to write_raw_bytes: returns
the element bytes verbatim regardless of the on-disk element type, so a
runtime CompoundType whose records have
no matching Rust primitive can be read back and decoded by the caller.
Sourcepub fn read_numeric_as<T: ReadNumeric>(&self) -> Result<Vec<T>>
pub fn read_numeric_as<T: ReadNumeric>(&self) -> Result<Vec<T>>
Read a numeric dataset as T, converting each element from the
on-disk datatype.
Unlike read_raw, which requires T’s size to match
the stored element size exactly, this inspects the dataset’s datatype
message — class, signedness, byte order, width — and converts per
element:
- integer → integer: checked; a stored value that does not fit in
Tis an error naming the element index and value, never a silent wrap. f32source →f64: exact widening.f64source →f32, float → integer, and integer → float are rejected asTypeMismatch.
Big-endian sources are decoded according to the datatype’s byte order,
which read_raw’s size-only check would misread.
let file = H5File::open("data.h5").unwrap();
let ds = file.dataset("counts").unwrap(); // stored as e.g. i16
let counts = ds.read_numeric_as::<i64>().unwrap();Sourcepub fn read_numeric_slice_as<T: ReadNumeric>(
&self,
starts: &[usize],
counts: &[usize],
) -> Result<Vec<T>>
pub fn read_numeric_slice_as<T: ReadNumeric>( &self, starts: &[usize], counts: &[usize], ) -> Result<Vec<T>>
Read a slice (hyperslab) of a numeric dataset as T, with the same
per-element datatype conversion as
read_numeric_as.
starts and counts define the N-dimensional selection exactly as in
read_slice.
Sourcepub fn read_raw_into<T: H5Type>(&self, out: &mut [T]) -> Result<()>
pub fn read_raw_into<T: H5Type>(&self, out: &mut [T]) -> Result<()>
Read the whole dataset into a caller-provided buffer, with no allocation.
out must have exactly product(dims) elements (the dataset’s element
count) and T::element_size() must match the dataset’s on-disk element
size, otherwise an error is returned and out is left unspecified. The
zero-copy counterpart of read_raw: the bytes are read
straight into out rather than into a fresh Vec, so a pinned /
page-locked host buffer can be filled in one pass and DMA’d to a GPU
without the extra staging copy a read_raw + copy-into-pinned would
incur. Works for every layout (contiguous, compact, and chunked under
any index); for chunked data each decoded chunk is scattered directly
into out.
let file = H5File::open("data.h5").unwrap();
let ds = file.dataset("frames").unwrap();
let n: usize = ds.shape().iter().product();
let mut buf = vec![0u16; n]; // or a pinned host allocation
ds.read_raw_into(&mut buf).unwrap();Sourcepub fn read_slice_into<T: H5Type>(
&self,
out: &mut [T],
starts: &[usize],
counts: &[usize],
) -> Result<()>
pub fn read_slice_into<T: H5Type>( &self, out: &mut [T], starts: &[usize], counts: &[usize], ) -> Result<()>
Read a hyperslab into a caller-provided buffer, with no allocation.
out must have exactly product(counts) elements and
T::element_size() must match the dataset’s element size. The zero-copy
counterpart of read_slice and the slice analogue of
read_raw_into: only chunks overlapping the
selection are read, and the selected bytes land directly in out — the
entry point for reading one frame / block straight into a pinned host
buffer for an H2D transfer.
let file = H5File::open("vol.h5").unwrap();
let ds = file.dataset("vol").unwrap(); // shape [nz, ny, nx]
let (ny, nx) = (ds.shape()[1], ds.shape()[2]);
let mut frame = vec![0f32; ny * nx]; // or a pinned host allocation
ds.read_slice_into(&mut frame, &[5, 0, 0], &[1, ny, nx]).unwrap();Trait Implementations§
Source§impl Drop for H5Dataset
impl Drop for H5Dataset
Source§fn drop(&mut self)
fn drop(&mut self)
Closing a virtual dataset’s last handle closes the source files that
open was holding, which is what H5D__virtual_reset_layout does at
the last H5Dclose (H5Dvirtual.c:709-710, closing each
source_dset->dset at :955 and with it the file that dataset kept
open). Nothing else in this crate can end a virtual open, so this is
where the reader is told.
The token is dropped before the reader is asked, so the reader’s
Weak already reads dead for the handle going away here. A write-mode
handle’s token belongs to the writer’s own external file prefix, which
has no source files to close, so it takes no lock either.