1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
//! Brotli compression, in safe Rust.
//!
//! `mbrotli` implements every Brotli quality as a port of Google's reference
//! encoder. Its compression APIs emit identical bytes for equivalent stream
//! settings, including empty inputs; C's one-shot rewrites are deliberately
//! omitted. Differential verification uses equivalent C streaming settings.
//! There is no `unsafe` in
//! this crate, and the SIMD instruction set is resolved once per compressor
//! rather than inside any loop.
//!
//! There is no decoder: round-trip verification uses Google's C decoder, and
//! this crate compresses only.
//!
//! # The shape of the API
//!
//! ```text
//! EncoderConfig what every stream is encoded with
//! Compressor a reusable encoder, and the workspace it owns
//! StreamConfig what one stream knows about itself
//! EncoderSession one stream, driven a chunk at a time
//! PreparedDictionary immutable knowledge many compressors can share
//! EncoderReader adapters over a session
//! EncoderWriter
//! ```
//!
//! A [`Compressor`] is stateful, and every encoding method takes `&mut self`.
//! That is the whole design: the encoder's hash tables, sliding window and
//! histograms are expensive to build and cheap to reuse, so a compressor keeps
//! them and hands them to the next call. Reuse is what ordinary code gets,
//! rather than something to opt into.
//!
//! One compressor belongs to one worker. To compress separate streams in
//! parallel, build one per worker with [`Compressor::fork_empty`]; a lock around
//! a single compressor would serialise the compression itself, not merely the access.
//! To compress one input across workers, use
//! [`ParallelCompressor`](compressor::parallel::ParallelCompressor) as shown below.
//!
//! # Choosing a quality
//!
//! | Quality | What it does | Typical use |
//! | --- | --- | --- |
//! | 0 | One pass, static entropy codes | Fastest, largest output |
//! | 1 | Two passes, per-block entropy codes | Fast |
//! | 2 | Greedy matching with the format's fixed codes | Fast |
//! | 3 | Greedy matching, one prefix code per stream | Balanced |
//! | 4 | Adds block splitting and histogram optimisation | Balanced, denser |
//! | 5 | Adds an extensive search and literal context modelling | Densest of these |
//! | 6 to 9 | Wider match search, more cached distances, richer context models | Denser, slower |
//! | 10, 11 | Binary-tree matching and a Zopfli dynamic program | Densest, slowest |
//!
//! [`EncoderConfig::default`] is quality 11, which mirrors the reference
//! encoder's default and is far slower than most callers want. For online
//! compression, say so:
//!
//! ```
//! use mbrotli::{Compressor, EncoderConfig, Quality};
//!
//! let mut encoder = Compressor::new(EncoderConfig::default().with_quality(Quality::Q5))?;
//! let payload = "the quick brown fox ".repeat(500);
//!
//! let compressed = encoder.compress(payload.as_bytes())?;
//!
//! assert!(compressed.len() < payload.len() / 100);
//! # Ok::<(), Box<dyn std::error::Error>>(())
//! ```
//!
//! # Reusing a compressor and its destination
//!
//! [`Compressor::compress_into`] is the entry point to reach for when there is
//! more than one thing to compress. It appends to a destination the caller
//! owns, so both the encoder's workspace and the output buffer are reused; a
//! warm compressor writing into a destination that is already big enough
//! allocates nothing beyond the tables it grows into over its first streams.
//!
//! ```
//! use mbrotli::{Compressor, EncoderConfig, Quality};
//!
//! let mut encoder = Compressor::new(EncoderConfig::default().with_quality(Quality::Q5))?;
//! let mut output = Vec::new();
//!
//! for payload in [&b"first"[..], b"second", b"third"] {
//! output.clear();
//! encoder.compress_into(payload, &mut output)?;
//! }
//! # Ok::<(), Box<dyn std::error::Error>>(())
//! ```
//!
//! # Streaming
//!
//! [`Compressor::writer`] compresses everything written to it. `Write` has no
//! closing hook and a meta-block boundary need not land on a byte boundary, so
//! the stream is terminated explicitly:
//!
//! ```
//! use mbrotli::io::FinishError;
//! use mbrotli::{Compressor, EncoderConfig, InputSize, Quality};
//! use std::io::Write;
//!
//! let mut encoder = Compressor::new(EncoderConfig::default().with_quality(Quality::Q5))?;
//! let payload = b"streamed in chunks".repeat(50);
//!
//! let streamed = {
//! let stream = InputSize::Exact(payload.len() as u64).into();
//! let mut sink = encoder.writer(Vec::new(), stream)?;
//! for chunk in payload.chunks(64) {
//! sink.write_all(chunk)?;
//! }
//! sink.finish().map_err(FinishError::into_error)?
//! };
//!
//! assert_eq!(streamed, encoder.compress(&payload)?);
//! # Ok::<(), Box<dyn std::error::Error>>(())
//! ```
//!
//! Declaring [`InputSize::Exact`] is what makes that last assertion hold:
//! qualities four and five choose their match finder from how much input is
//! coming, so a stream that does not say produces different — equally valid —
//! bytes. [`Compressor::reader`] is the pull-shaped counterpart, and
//! [`Compressor::start`] is the state machine both are built on.
//!
//! # Parallel compression
//!
//! [`ParallelCompressor`](compressor::parallel::ParallelCompressor) splits an
//! input into independent segments and assembles them into one Brotli stream.
//! The caller schedules the tasks; this example uses scoped threads and
//! automatically sized memory staging:
//!
//! ```
//! use mbrotli::compressor::parallel::{
//! BatchConfig, ParallelCompressor, ParallelConfig, TaskCount,
//! };
//! use mbrotli::{EncoderConfig, Quality};
//!
//! let payload = vec![b'a'; 8 << 20];
//! let mut encoder = ParallelCompressor::new(
//! EncoderConfig::default().with_quality(Quality::Q5),
//! ParallelConfig::default(),
//! )?;
//! let mut batch = encoder.prepare_slice(
//! &payload,
//! BatchConfig::auto(TaskCount::try_from(2)?),
//! )?;
//! let tasks = batch.take_tasks()?;
//! std::thread::scope(|scope| {
//! for task in tasks {
//! scope.spawn(move || task.run());
//! }
//! });
//! let mut compressed = Vec::new();
//! let result = batch.finish_into(&mut compressed)?;
//!
//! assert_eq!(result.stats.effective_tasks, 2);
//! assert!(compressed.len() < payload.len());
//! # Ok::<(), Box<dyn std::error::Error>>(())
//! ```
//!
//! Tasks are capped at the segment count; the default segment size is 4 MiB.
//! Finishing collects task errors and assembles segments in input order.
//! Independent segments can produce different bytes and compressed sizes from
//! serial compression, while fixed segment settings give identical output
//! across task counts and execution orders.
//!
//! # Large Window Brotli
//!
//! [RFC 9841] widens the sliding window past what RFC 7932 can express. Which
//! header a stream carries is part of the window itself: build one with
//! [`Window::standard`] or [`Window::large`], never by widening a number.
//!
//! ```
//! use mbrotli::{Compressor, EncoderConfig, Quality, Window};
//!
//! let config = EncoderConfig::default()
//! .with_quality(Quality::Q5)
//! .with_window(Window::large(30)?);
//! let mut encoder = Compressor::new(config)?;
//!
//! let compressed = encoder.compress("large window ".repeat(1000).as_bytes())?;
//!
//! // The stream carries the RFC 9841 header, so it needs a decoder expecting one.
//! assert_eq!(compressed[0], 0b0001_0001);
//! # Ok::<(), Box<dyn std::error::Error>>(())
//! ```
//!
//! Qualities 0, 1 and 2 write distances through a model built for the RFC 7932
//! alphabet, so [`Compressor::new`] refuses a Large Window there rather than
//! quietly dropping the request.
//!
//! # Shared dictionaries
//!
//! RFC 9841 also lets a caller attach up to fifteen LZ77 prefix dictionaries in
//! front of a stream. A [`PreparedDictionary`](dictionary::PreparedDictionary)
//! is immutable and holds no per-stream state, so any number of compressors may
//! borrow one at once without a lock.
//!
//! ```
//! use mbrotli::dictionary::DictionaryBuilder;
//! use mbrotli::{Compressor, EncoderConfig, Quality};
//!
//! let dictionary = DictionaryBuilder::new()
//! .add_prefix(&b"HTTP/1.1 200 OK\r\nContent-Type: "[..])
//! .build()?;
//! let mut encoder = Compressor::new(EncoderConfig::default().with_quality(Quality::Q5))?;
//!
//! let payload = b"Content-Type: text/html; charset=utf-8";
//! assert!(
//! encoder.compress_with_dictionary(&dictionary, payload)?.len()
//! < encoder.compress(payload)?.len()
//! );
//! # Ok::<(), Box<dyn std::error::Error>>(())
//! ```
//!
//! Below quality five no match finder can consult a dictionary, and one handed
//! to such a compressor is refused with
//! [`EncodeError::DictionaryUnsupportedForQuality`] rather than ignored: a
//! stream compressed without the dictionary it was given decodes perfectly
//! well, which is what would make the mistake invisible.
//!
//! The `experimental` feature adds serialized shared dictionaries, custom word
//! and transform indexes, headerless stream continuations, and the separate
//! Shared Brotli framing writer. Equivalent-C-streaming byte comparisons do not
//! cover every extension. Rust API/backend identity and decoder compatibility
//! remain required for equivalent stream settings.
//!
//! [RFC 9841]: https://www.rfc-editor.org/rfc/rfc9841.html
// The port is safe Rust by construction: the bit writer, the match scans and
// the SIMD kernels all shed their bounds checks through `as_chunks`,
// `first_chunk` and const-generic widths rather than through raw pointers.
// `forbid` rather than `deny`, so no module can opt back in.
//
// The differential unit tests inside `core::hq` and `core::rfc9841` call
// Google's C encoder through `google-brotli-ffi` to compare a stage against
// its reference, which is unavoidably `unsafe`. Those live behind `cfg(test)`
// and reach nothing that ships, so the ban is on everything but the test
// build rather than weakened to a `deny` the shipped code could opt out of.
pub use framing;
pub use ;
pub use ;