pub struct Documents<R, O: Options = Standard> { /* private fields */ }Expand description
A reader turned into a series of BEVE values.
One value is held at a time. The buffer keeps a working window over the
stream, compacted as values are consumed, so a file of a million records
costs roughly one record plus one read, not a million records. This is the
answer to a BEVE file too large to hold: from_slice
wants the whole document, and this wants one value of it.
Pick the constructor that matches how the producer laid the values out:
Documents::array for the elements of one big array, Documents::values
for whole documents back to back.
let file = structio::to_beve(&vec![Record { id: 1 }, Record { id: 2 }]);
let mut docs = structio::beve::Documents::array(&file[..]);
let ids: Vec<u64> = docs.iter::<Record>().map(|r| r.unwrap().id).collect();
assert_eq!(ids, [1, 2]);O is the read policy every value is read under. The
constructors give you Standard; Documents::with_options changes it,
as one more link in the same builder chain that sets the size limits.
Implementations§
Source§impl<R: Read> Documents<R>
impl<R: Read> Documents<R>
Sourcepub fn array(reader: R) -> Self
pub fn array(reader: R) -> Self
The elements of a single top-level array.
Both array forms work. A generic array holds whole values; a typed one
holds a block with a single header for the lot, and its elements are
handed out one at a time all the same, so a file that is one enormous
Vec<f64> streams as f64s.
Sourcepub fn values(reader: R) -> Self
pub fn values(reader: R) -> Self
Whole BEVE documents one after another.
A single document is the one-value case, but note that it buys nothing
over from_beve_reader there: one value is
buffered whole either way.
Source§impl<R: Read, O: Options> Documents<R, O>
impl<R: Read, O: Options> Documents<R, O>
Sourcepub fn with_options<P: Options>(self) -> Documents<R, P>
pub fn with_options<P: Options>(self) -> Documents<R, P>
Read every value under the policy P instead.
Sourcepub fn max_value(self, bytes: usize) -> Self
pub fn max_value(self, bytes: usize) -> Self
Fail rather than buffer more than bytes for a single value.
Unlimited by default. Set this when the producer is not trusted: BEVE states its own extents, so a hostile document can claim a length it never delivers, and this is what stops the window growing to meet it. Reads are clipped so the window never runs more than a byte past the limit before the failure is noticed.
Sourcepub fn read_size(self, bytes: usize) -> Self
pub fn read_size(self, bytes: usize) -> Self
How many bytes to request per read. Defaults to 64 KiB.
It sizes the window as well as the read: the buffer is allocated on the first fill and holds one chunk, so this is the knob for a caller who is decoding a small document that is already in memory and does not want 64 KiB of buffer behind it. Set it larger than a value and the value is still buffered whole; the window grows to whatever one value needs.
Sourcepub fn into_inner(self) -> R
pub fn into_inner(self) -> R
Recover the underlying reader, discarding the window.
Reading is done a chunk at a time, so bytes past the last value
returned have usually already been taken from the reader, and those are
lost. This mirrors io::BufReader::into_inner, which is lossy for the
same reason. Use it to finish with a reader, not to hand a live stream
on to something else; Documents::into_parts is the lossless form.
Sourcepub fn into_parts(self) -> (R, Vec<u8>)
pub fn into_parts(self) -> (R, Vec<u8>)
Recover the underlying reader together with the bytes already taken from it that did not become a value.
Concatenating the returned bytes with everything still in the reader reconstructs the remainder of the stream exactly, which is what makes it safe to hand a partly consumed stream to something else.
The bytes are empty when framing has failed, since the position in the stream is no longer known and there is nothing honest to resume from.
Sourcepub fn next_value<'a, T: Read<'a> + Default>(
&'a mut self,
) -> Option<StreamResult<T>>
pub fn next_value<'a, T: Read<'a> + Default>( &'a mut self, ) -> Option<StreamResult<T>>
The next value, which may borrow from the stream buffer.
The borrow is of self, so the reader cannot advance while the value is
alive. That is what makes zero-copy &str and &[u8] fields work here:
they point into the window, and the window is pinned until you drop
them. For values that own their data, Documents::iter is an ordinary
iterator and reads better in a loop.
None means the stream ended, either cleanly or because framing has
already failed and the failure was reported.
A value that fails to read is reported and skipped; the framing is still intact, so the next value is read normally. That is what makes per-record error recovery work for a file of records.
A failure to frame is different: the position in the input is no
longer known, so there is nothing honest to resume from. It is reported
once and ends the stream, rather than being returned forever and turning
the natural while let loop into a spin.
Sourcepub fn next_value_into<T: for<'de> Read<'de>>(
&mut self,
value: &mut T,
) -> Option<StreamResult<()>>
pub fn next_value_into<T: for<'de> Read<'de>>( &mut self, value: &mut T, ) -> Option<StreamResult<()>>
The next value, read into one you already have.
Mirrors read_beve_into: value keeps its
allocations between calls, so a loop over a million records of the same
shape settles into doing no allocation at all.