pub struct Documents<R, O: Options = Standard> { /* private fields */ }Expand description
A reader turned into a series of JSON values.
One value is held at a time. The buffer keeps a working window over the stream, compacted as values are consumed, so a file of a million records costs roughly one record plus one read, not a million records.
Pick the constructor that matches how the producer laid the values out:
Documents::lines for newline-delimited records, Documents::array
for the elements of one big array, Documents::values for bare values
back to back.
O is the read policy every value is read under. The
constructors give you Standard; Documents::with_options changes it,
as one more link in the same builder chain that sets the size limits.
Implementations§
Source§impl<R: Read> Documents<R>
impl<R: Read> Documents<R>
Sourcepub fn lines(reader: R) -> Self
pub fn lines(reader: R) -> Self
Newline-delimited JSON: one value per line, blank lines ignored.
Sourcepub fn values(reader: R) -> Self
pub fn values(reader: R) -> Self
Whole JSON values one after another, separated by optional whitespace.
A single document is the one-value case, but note that it buys nothing
over from_reader there: one value is buffered
whole either way.
Source§impl<R: Read, O: Options> Documents<R, O>
impl<R: Read, O: Options> Documents<R, O>
Sourcepub fn with_options<P: Options>(self) -> Documents<R, P>
pub fn with_options<P: Options>(self) -> Documents<R, P>
Read every value under the policy P instead.
use structio::{Documents, SkipUnknown};
let input = &b"{\"id\":1,\"note\":\"ignored\"}"[..];
let mut docs = Documents::lines(input).with_options::<SkipUnknown>();
assert_eq!(docs.iter::<Rec>().next().unwrap().unwrap(), Rec { id: 1 });Sourcepub fn max_value(self, bytes: usize) -> Self
pub fn max_value(self, bytes: usize) -> Self
Fail rather than buffer more than bytes for a single value.
Unlimited by default. Set this when the producer is not trusted: it is what stops a stream that never closes a bracket from consuming all available memory. Reads are clipped so the window never runs more than a byte past the limit before the failure is noticed.
Sourcepub fn read_size(self, bytes: usize) -> Self
pub fn read_size(self, bytes: usize) -> Self
How many bytes to request per read. Defaults to 64 KiB.
It sizes the window as well as the read: the buffer is allocated on the first fill and holds one chunk, so this is the knob for a caller who is decoding a small document that is already in memory and does not want 64 KiB of buffer behind it. Set it larger than a value and the value is still buffered whole; the window grows to whatever one value needs.
Sourcepub fn into_inner(self) -> R
pub fn into_inner(self) -> R
Recover the underlying reader, discarding the window.
Reading is done a chunk at a time, so bytes past the last value
returned have usually already been taken from the reader, and those are
lost. This mirrors io::BufReader::into_inner, which is lossy for the
same reason. Use it to finish with a reader, not to hand a live stream
on to something else; Documents::into_parts is the lossless form.
Sourcepub fn into_parts(self) -> (R, Vec<u8>)
pub fn into_parts(self) -> (R, Vec<u8>)
Recover the underlying reader together with the bytes already taken from it that did not become a value.
Concatenating the returned bytes with everything still in the reader reconstructs the remainder of the stream exactly, which is what makes it safe to hand a partly consumed stream to something else.
The bytes are empty when framing has failed, since the position in the stream is no longer known and there is nothing honest to resume from.
Sourcepub fn next_value<'a, T: Read<'a> + Default>(
&'a mut self,
) -> Option<StreamResult<T>>
pub fn next_value<'a, T: Read<'a> + Default>( &'a mut self, ) -> Option<StreamResult<T>>
The next value, which may borrow from the stream buffer.
The borrow is of self, so the reader cannot advance while the value
is alive. That is what makes zero-copy &str fields work here: they
point into the window, and the window is pinned until you drop them.
For values that own their data, Documents::iter is an ordinary
iterator and reads better in a loop.
None means the stream ended, either cleanly or because framing has
already failed and the failure was reported.
A value that fails to parse is reported and skipped; the framing is still intact, so the next value is read normally. That is what makes per-record error recovery work for a file of records.
A failure to frame is different: the position in the input is no
longer known, so there is nothing honest to resume from. It is reported
once and ends the stream, rather than being returned forever and
turning the natural while let loop into a spin.
This is the borrowing half of a pair, and the only streaming form open
to a type that borrows from the document it was read out of. The 'a is
what makes that work: the value’s input lifetime is the borrow of the
reader, so the window cannot be filled, compacted or dropped while a
field still points into it. Documents::next_value_into is the owning
half, and says what the difference costs.
Sourcepub fn next_value_into<T: for<'de> Read<'de>>(
&mut self,
value: &mut T,
) -> Option<StreamResult<()>>
pub fn next_value_into<T: for<'de> Read<'de>>( &mut self, value: &mut T, ) -> Option<StreamResult<()>>
The next value, read into one you already have.
Mirrors read_into: value keeps its allocations
between calls, so a loop over a million records of the same shape
settles into doing no allocation at all.
The for<'de> bound is a soundness requirement rather than a
convenience, and it is what a type with a borrowing field, such as one
holding a &'de str or a Raw, fails to satisfy.
value is a &mut T that outlives this call in both directions: it was
alive before the window was filled and is still alive after the window
has compacted the bytes this value was read out of. A T that had
borrowed from the window would then be holding a pointer into bytes that
have moved. Only a T that can be read under any input lifetime
borrows from no input at all, and for<'de> is how that is spelled.
So the two forms divide the work rather than overlapping, and which one
a type has is decided by whether it borrows.
Documents::next_value is the borrowing form, and a borrowing type
has only that one: Documents::iter takes ReadOwned, which is
this bound plus the Default it needs to construct each value. What such a type gives up is the allocation reuse, which
costs it nothing for the fields that borrow, those being subslices of
the window either way, and costs it a rebuild per value for any owned
field standing beside them. That is a real limitation of the streaming
path and not a bound anybody can loosen: the alternative is not slower,
it is unsound.