1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
//! Reading and writing JSON that does not fit, or has not arrived.
//!
//! The batch entry points ([`from_str`](crate::from_str),
//! [`to_string`](crate::to_string)) want the whole document in memory at once.
//! This module covers the cases where that is not what you have: output going
//! to a socket, input arriving in pieces, or a file of records too large to
//! hold.
//!
//! # Writing
//!
//! [`to_writer`] serializes straight into an [`io::Write`], draining as it
//! goes. Peak memory is the writer's buffer plus the largest single scalar,
//! whatever the size of the document.
//!
//! ```no_run
//! # #[derive(Default)] struct Record { id: u64 }
//! # structio::object!(Record { id });
//! # fn main() -> std::io::Result<()> {
//! let file = std::fs::File::create("out.json")?;
//! structio::to_writer(&Record { id: 7 }, file)?;
//! # Ok(())
//! # }
//! ```
//!
//! # Reading a sequence
//!
//! [`Documents`] turns a reader into a series of values, holding only one at a
//! time. It handles the three shapes a JSON stream comes in: newline-delimited
//! records, the elements of one large array, and bare values back to back.
//!
//! ```
//! # #[derive(Default, Debug, PartialEq)] struct Record { id: u64 }
//! # structio::object!(Record { id });
//! let input = b"{\"id\":1}\n{\"id\":2}\n" as &[u8];
//! let mut docs = structio::Documents::lines(input);
//! let ids: Vec<u64> = docs
//! .iter::<Record>()
//! .map(|r| r.unwrap().id)
//! .collect();
//! assert_eq!(ids, [1, 2]);
//! ```
//!
//! Values that borrow from the input are available too, through
//! [`Documents::next_value`], which holds the reader still for as long as the value
//! lives.
//!
//! # Reading from chunks you are handed
//!
//! [`Feed`] is the same machine driven from the other side: push bytes in as
//! they arrive, and take values out as they complete. Chunks may split a value
//! anywhere, including inside a string or a number.
//!
//! ```
//! # #[derive(Default, Debug, PartialEq)] struct Record { id: u64 }
//! # structio::object!(Record { id });
//! let mut feed = structio::Feed::values();
//! feed.push(b"{\"id\":1}{\"i");
//! assert_eq!(feed.next_value::<Record>().unwrap().unwrap(), Record { id: 1 });
//! assert!(feed.next_value::<Record>().is_none());
//! feed.push(b"d\":2}");
//! assert_eq!(feed.next_value::<Record>().unwrap().unwrap(), Record { id: 2 });
//! ```
//!
//! # What suspends, and what does not
//!
//! Both readers suspend and resume at any byte: the structural scan that finds
//! a value's extent carries its whole state across chunk boundaries. What they
//! do not do is hand back a half-filled struct. A value is parsed once its
//! bytes are all present, by the same [`Parser`](crate::json::Parser) the batch API
//! uses, which is what keeps streamed and slurped documents from ever
//! disagreeing about what a document means, and keeps borrowed `&str` fields
//! working.
//!
//! The practical consequence is that memory is bounded by the largest single
//! *value*, not by the largest document. For [`Mode::Array`] and
//! [`Mode::Lines`] that is one record, which is the case the format exists to
//! serve. A single enormous object read as one value is still buffered whole.
//!
//! # Untrusted input
//!
//! There is no size limit by default, because failing on a legitimately large
//! record is worse than the alternative for the common case. When the producer
//! is not trusted, set one with [`Documents::max_value`] or
//! [`Feed::max_value`]; a value that exceeds it fails with
//! [`ErrorCode::DocumentTooLarge`](crate::ErrorCode::DocumentTooLarge) rather
//! than growing the buffer. On the pull side reads are clipped so the window
//! never runs more than a byte past the limit; on the push side the limit
//! bounds what is retained, not the size of a chunk you hand to
//! [`Feed::push`].
//!
//! A value that fails to *parse* is reported and skipped, so one bad record in
//! a file does not end the run. A failure to *frame* is terminal: the position
//! in the input is no longer known, so it is reported once and the stream ends
//! there.
//!
//! One divergence from the batch API is deliberate: an input holding nothing
//! but whitespace is an empty stream, with zero values and no error, in every
//! mode. [`from_str`](crate::from_str) rejects the same input, because there a
//! document was asked for and none was found. An empty file is a normal thing
//! for a stream of records to be.
use io;
use crate;
use crateWriter;
use crate;
pub use Feed;
pub use ;
pub use Mode;
// Not re-exported: `StreamError` is format independent and lives in
// [`error`](crate::error), reachable as `structio::StreamError`. Naming it here
// too would imply the JSON side owns it, which it has not since BEVE arrived.
use crate;
/// Serialize a value straight into an [`io::Write`].
///
/// The document is drained to `out` as it is produced rather than assembled
/// first, so this holds [`DEFAULT_SINK_BUFFER`](crate::json::writer::DEFAULT_SINK_BUFFER)
/// bytes plus the largest single value written, however large the document is.
/// A long string is that largest value and is buffered whole, so the bound is
/// only as good as the longest string in the document.
///
/// Wrapping `out` in a [`BufWriter`](std::io::BufWriter) is redundant: this
/// already buffers and hands the sink whole blocks. `out` is not flushed.
///
/// The error is `out`'s own, handed back unchanged. A [`Write`] implementation
/// returns nothing to fail with, so the sink is the only thing on this path
/// that can go wrong and there is no crate error type to fold it into. A
/// caller wanting one error type across reading and writing has
/// [`StreamError`], which converts from [`io::Error`] with `?`.
/// [`to_writer`] under an explicit [write policy](crate::Options).
/// [`to_writer`] with an explicit buffer size.
/// [`to_writer_buffered`] under an explicit [write policy](crate::Options).
/// Parse one JSON document from an [`io::Read`].
///
/// The reader is drained into memory and then parsed, so this is a
/// convenience rather than a way to bound memory: the whole document is held
/// at once. For a stream of many values, or for a document larger than you
/// want resident, use [`Documents`].
///
/// `T` may not borrow from the input, since the buffer does not outlive the
/// call. [`Documents::next_value`] is the borrowing form.
/// [`from_reader`] under an explicit [read policy](crate::Options).