Skip to main content

Crate dag_ml_data_arrow

Crate dag_ml_data_arrow 

Source
Expand description

Apache Arrow IPC feature buffer reader for dag-ml-data.

This crate is a workspace member so the workspace cargo build / cargo test commands always compile it, but the C ABI wiring in dag-ml-data-capi is gated behind the arrow-ipc feature, so the Arrow dependency is opt-in at the cdylib boundary for hosts that ship libdag_ml_data_capi.so. Workspace consumers running the gate still pay the compile-time cost.

The reader maps each Arrow RecordBatch whose schema metadata carries dag_ml_data.feature_set_id to a NumericFeatureMatrixF64Columnar, then resolves the whole stream to a NumericFeatureBufferStore. The mapping is total:

  • Float64Array column → NumericFeatureMatrixF64Columnar.columns[i]
  • Arrow validity bitmap → per-column validity_masks[i]
  • observation_id Utf8Array column → observation_ids
  • feature column names → feature_names

Mandatory schema metadata keys:

  • dag_ml_data.feature_set_id — the feature-set id of the buffer
  • dag_ml_data.representation_id — the representation id of the buffer

Constants§

OBSERVATION_ID_COLUMN
Column name holding per-row observation ids in each record batch.
SCHEMA_METADATA_FEATURE_SET_ID
Arrow schema-metadata key carrying the source feature-set id.
SCHEMA_METADATA_REPRESENTATION_ID
Arrow schema-metadata key carrying the representation id.

Functions§

read_buffers_from_ipc_file
Parse an Arrow IPC file (with footer + magic) from any Read + Seek source into a NumericFeatureBufferStore. Same per-batch contract as read_buffers_from_ipc_stream.
read_buffers_from_ipc_path
Convenience: read an Arrow IPC file from disk.
read_buffers_from_ipc_stream
Parse an Arrow IPC stream from any Read source into a NumericFeatureBufferStore. Each top-level record batch becomes one buffer; the batches’ schemas must declare both dag_ml_data.feature_set_id and dag_ml_data.representation_id in their metadata map and expose an observation_id UTF-8 column.