Expand description
Apache Arrow IPC feature buffer reader for dag-ml-data.
This crate is a workspace member so the workspace cargo build /
cargo test commands always compile it, but the C ABI wiring in
dag-ml-data-capi is gated behind the arrow-ipc feature, so the
Arrow dependency is opt-in at the cdylib boundary for hosts that
ship libdag_ml_data_capi.so. Workspace consumers running the gate
still pay the compile-time cost.
The reader maps each Arrow RecordBatch whose schema metadata carries
dag_ml_data.feature_set_id to a NumericFeatureMatrixF64Columnar,
then resolves the whole stream to a NumericFeatureBufferStore. The
mapping is total:
Float64Arraycolumn →NumericFeatureMatrixF64Columnar.columns[i]- Arrow validity bitmap → per-column
validity_masks[i] observation_idUtf8Arraycolumn →observation_ids- feature column names →
feature_names
Mandatory schema metadata keys:
dag_ml_data.feature_set_id— the feature-set id of the bufferdag_ml_data.representation_id— the representation id of the buffer
Constants§
- OBSERVATION_
ID_ COLUMN - Column name holding per-row observation ids in each record batch.
- SCHEMA_
METADATA_ FEATURE_ SET_ ID - Arrow schema-metadata key carrying the source feature-set id.
- SCHEMA_
METADATA_ REPRESENTATION_ ID - Arrow schema-metadata key carrying the representation id.
Functions§
- read_
buffers_ from_ ipc_ file - Parse an Arrow IPC file (with footer + magic) from any
Read + Seeksource into aNumericFeatureBufferStore. Same per-batch contract asread_buffers_from_ipc_stream. - read_
buffers_ from_ ipc_ path - Convenience: read an Arrow IPC file from disk.
- read_
buffers_ from_ ipc_ stream - Parse an Arrow IPC stream from any
Readsource into aNumericFeatureBufferStore. Each top-level record batch becomes one buffer; the batches’ schemas must declare bothdag_ml_data.feature_set_idanddag_ml_data.representation_idin theirmetadatamap and expose anobservation_idUTF-8 column.