pub enum ChunkSize {
Fixed(usize),
Adaptive,
Custom(fn(u64) -> usize),
}Expand description
How much data a streaming engine reads per chunk.
The unit is bytes, everywhere. Chunk size bounds how much of the source
is resident at once, so it is expressed in the unit that actually bounds
memory rather than in rows, whose width varies per dataset. Every engine and
binding agrees: ChunkSize::Fixed(65_536) and Python’s chunk_size=65536
both mean 64 KiB per chunk, whether the source is a file, a byte stream, or
a URL.
Chunk size never changes what a profile contains — only the granularity at which the source is read, progress is emitted, and chunk-level stop conditions are evaluated.
Variants§
Fixed(usize)
Fixed chunk size in bytes.
Adaptive
Let the engine choose the chunk size (default).
Each engine resolves this against what it knows: the incremental engine
derives a size from its memory limit and the file size, while the async
reader — whose source has no length to adapt to — uses a fixed working-set
target. Unlike Fixed, the resulting size is not a
guarantee, so it is not something to assert against.
Custom(fn(u64) -> usize)
Custom sizing function, given the source size in bytes and returning a chunk size in bytes. Cannot derive Debug/Clone with a function pointer.
Implementations§
Source§impl ChunkSize
impl ChunkSize
Sourcepub fn calculate(&self, file_size_bytes: u64) -> usize
pub fn calculate(&self, file_size_bytes: u64) -> usize
Resolve to a concrete chunk size in bytes for a source of the given size.
This is a standalone helper, not the path any engine takes: Fixed and
Custom resolve exactly as an engine would, but Adaptive here is
derived from system available memory, whereas an engine derives it
from its own configured memory limit. Do not use this to predict what an
engine will do with Adaptive.