fast-floe
An optimized, multi-provider, spec-complianct implementation of the Fast Lightweight Online Encryption (FLOE) scheme for authenticated encryption of large files and byte streams with bounded-memory and random access.
Introduction
fast-floe implements the Fast Lightweight Online Encryption (FLOE)
specification in Rust. FLOE divides a message into independently authenticated
and encrypted segments of user chosen length. A recipient can encrypt/decrypt a
large message one segment at a time and read/seek randomly into the ciphertext.
FLOE is built for really big data, streaming data where you might not know size ahead of time, data that might arrive out-of-order, and random-access into encrypted data.
Callers select the segment length at encryption time to trade off throughput for random-access overhead (smaller segment length == cheaper random access, and larger segment length == more throughput).
Multiple cryptography providers are supported (aws-lc-rs, boring, ring, and RustCrypto).
They all interoperate with each other (i.e. provider choice does not affect wire format).
Quick start
Add fast-floe from crates.io to
Cargo.toml. Rust refers to the crate as fast_floe.
[]
= "0.3"
aws-lc-rs is the default cryptographic provider, but can be overridden, see the "Cryptographic Providers" section below.
use ;
#
encrypt returns one complete FLOE message: an authenticated header followed
by encrypted segments. decrypt authenticates the message before returning its
plaintext.
The three inputs are:
Key: a 32-byte secret. useKey::from_bytesto import existing key material andKey::generateto generate a new random key.AAD: context that must be authenticated with the ciphertext, such as an session ID or protocol version. FLOE does not store the AAD. Decryption requires the same AAD bytes to succeed.Parameters: the encrypted segment length. FLOE supports all segments lengths from 64 to 4,294,967,294 (u32::MAX- 2) bytes.
Choose an API
Ordered below from simplest (and more high-level) to more advanced (and low-level). Use the highest layer that fits your needs:
| Need | API | Input |
|---|---|---|
| Encrypt or decrypt bytes already in memory | encrypt, decrypt |
One-shot complete message |
Process a file, socket, or std::io adapter |
fast_floe::io |
Read or Write |
| Read selected authenticated ranges | fast_floe::random_access |
Read + Seek ciphertext |
| Exchange segments, possibly out-of-order | fast_floe::online |
Streaming data |
| Process segments manually or in parallel | fast_floe::low_level |
Experts needing control |
Whole messages
Use the quick-start encrypt and decrypt functions when the complete input
and output fit in memory. Both return a new Vec<u8>.
decrypt uses the authenticated profile (segment size) read from the header.
If your application wants to choose the profile instead, use decrypt_with_parameters.
Streams with std::io
EncryptReader and DecryptReader keep memory bounded while composing with
ordinary Rust I/O:
use ;
use ;
use ;
#
Read EncryptReader to EOF to complete encryption. EncryptWriter is available
when a Write interface fits better.
DecryptReader::finish authenticates unread ciphertext and rejects trailing
bytes. Use finish_frame when non-FLOE data follows the FLOE message.
IMPORTANT: always call finish or try_finish to complete the FLOE message.
flush will not emit a partial non-final FLOE segment, and dropping an
unfinished writer leaves a truncated message.
Authenticated random access
random_access::Reader presents a seekable ciphertext as authenticated
plaintext. It implements Read + Seek and also reads explicit ranges:
use Cursor;
use Reader;
use ;
#
Reader construction authenticates the header and final segment, which establishes the
FLOE profile (segment size) and complete length. Each requested segment is authenticated
before its bytes are returned. The underlying seekable source must remain unchanged while
the reader is in use. Use Reader::new_with_length when non-FLOE data follows the last segment.
Segment-oriented processing
Use online::Encryptor and online::Decryptor when your transport already
works with packets, buffers, or other segment-like things.
use ;
use ;
use ;
#
All non-final plaintext segments must have exactly
parameters.plaintext_segment_length() bytes. The final segment may be shorter,
including empty. Final encryption consumes the encryptor. Decryption reads the
final marker from each segment, authenticates it, and finish detects a missing
final segment (e.g. detects truncation). If the input ends exactly on a segment
boundary, the next loop iteration emits an empty final segment.
Low-level API
fast_floe::low_level exposes the FLOE specification's details. A
MessageLayout supplies the correct position, lengths, offsets, and finality
for every segment and you should use it when you know the plaintext length in advance:
use start_encryption;
use ;
#
Parallel or concurrent workloads can call state.into_shared() and then shared.fork()
for each distinct worker/thread. SharedEncryptionContext and SharedDecryptionContext
are thread-safe (e.g. they are Send + Sync).
Low-level encryption requires the caller to handle all FLOE invariants: process every position once, produce exactly one final segment, leave no gaps, and process nothing after the final segment. Breaking these rules can break message security. Prefer the misuse-resistant higher-level APIs.
low_level::SegmentBuffer supports reusable in-place storage. If you need to control
allocation, *_raw methods are available for use w/ your own buffer management.
Notes
- AAD is authenticated but is not stored in the ciphertext. Store or derive it separately and reproduce it exactly.
- Each segment is released only after its own authentication succeeds. A stream consumer may receive valid early segments before a later corruption or truncation is discovered. Callers should stage plaintext until finalization when the application requires all-or-nothing release.
- Segment prefixes and layout calculations describe framing. Treat them as untrusted until the corresponding header or segment authenticates.
Segment sizes
FLOE supports all segment lengths between 64 and 4,294,967,294 (u32::MAX - 2) bytes, inclusive.
Any size in that range is valid, including odd lengths and non-powers-of-2. The constant
Parameters::VALID_SEGMENT_LENGTHS encodes this range.
Use Parameters::from_segment_length() to construct Parameters with your desired length,
or use one of the Parameters::SEGMENT_* convenience constants.
Only the encrypted segment size differs, all other FLOE parameters (AES-256-GCM, HKDF-SHA-384, IV length) are the same.
Parameters::plaintext_layout and Parameters::ciphertext_layout calculate a
complete random_access::MessageLayout. Use them for storage sizing and manual
segment processing.
Cryptographic providers
This crate uses different providers to implement AES-256-GCM, HKDF-SHA-384,
and random-number generation. The crate default is aws-lc-rs, but all
of these are supported:
| Feature | Provider crate(s) |
|---|---|
aws-lc-rs (default) |
aws-lc-rs |
boring |
boring |
ring |
ring |
rustcrypto |
RustCrypto aes-gcm and rand_chacha |
Provider choice does not change the FLOE wire format. They are all compatible with each other.
To use an alternate provider:
= { = "0.3", = false, = ["ring"] }
Provider features are additive. With exactly one compiled provider,
Key::generate and Key::from_bytes use it automatically. But with several,
you must bind one to the key with Key::generate_with_provider or Key::from_bytes_with_provider.
Examples and development
The repository includes two file examples:
serial_fileuses the bounded-memorystd::ioadapters.manual_fileuses message layouts and low-level segment operations.
Run them with:
cargo run --example serial_file -- encrypt INPUT OUTPUT 64_HEX_KEY [4k|1m]
cargo run --example manual_file -- encrypt INPUT OUTPUT 64_HEX_KEY [4k|1m]
For a non-default provider, add --no-default-features --features ring
before --.
Run the default-provider test and documentation checks with:
RUSTDOCFLAGS="-D warnings"
Run all provider and segment-size benchmarks with:
Benchmark results
Median throughput in GiB/s with each provider compiled solo with -C target-cpu=native and a
single codegen unit. Each benchmark sizes its buffer to be at least 4x the
detected last-level-cache size to ensure we're reaching actual memory, not spinning in cache.
The "into" columns encrypt or decrypt into a separate output buffer (using scatter/gather when the provider supports it) while "in place" overwrite their input.
AMD Zen 5 9950X (VAES, AVX-512), Rust 1.97.1
FLOE segment-oriented performance in GiB/sec, higher is better
| Provider | Segments | Encrypt into | Encrypt in place | Decrypt into | Decrypt in place |
|---|---|---|---|---|---|
aws-lc-rs |
1 MiB | 13.88 | 21.50 | 12.05 | 22.25 |
boring |
1 MiB | 11.98 | 17.72 | 8.56 | 19.29 |
ring |
1 MiB | 6.18 | 11.87 | 6.43 | 12.90 |
rustcrypto |
1 MiB | 4.84 | 4.87 | 4.93 | 4.94 |
aws-lc-rs |
4 KiB | 8.85 | 10.15 | 9.18 | 11.31 |
boring |
4 KiB | 10.95 | 14.76 | 7.98 | 16.12 |
ring |
4 KiB | 5.99 | 9.15 | 6.91 | 10.31 |
rustcrypto |
4 KiB | 4.67 | 4.84 | 4.81 | 4.59 |
Compared to "bare" AES-256-GCM from each provider, FLOE "in_place" on Zen5 is within ~5% on
1 MiB segments for aws-lc-rs, ring, and rustcrypto (boring is ~12% slower), and
roughly 5%-45% slower on 4 KiB segments.
The std::io adapters and the random-access reader have additional overhead on top of the segment
operations above. The Reader's operate on whole-segments which are copy-free. Sub-segment
sized reads are buffered (one copy).
| Provider | Segments | EncryptWriter | EncryptReader | DecryptReader | Reader (seq) | Reader (range) |
|---|---|---|---|---|---|---|
aws-lc-rs |
1 MiB | 18.07 | 13.79 | 12.35 | 10.85 | 9.48 |
boring |
1 MiB | 16.62 | 13.27 | 10.22 | 10.56 | 6.96 |
ring |
1 MiB | 8.22 | 7.48 | 7.34 | 7.55 | 5.63 |
rustcrypto |
1 MiB | 4.89 | 4.18 | 4.24 | 3.87 | 3.92 |
aws-lc-rs |
4 KiB | 11.91 | 10.89 | 9.03 | 6.81 | 6.93 |
boring |
4 KiB | 14.96 | 9.42 | 8.55 | 8.97 | 6.81 |
ring |
4 KiB | 8.12 | 7.29 | 7.82 | 7.10 | 5.96 |
rustcrypto |
4 KiB | 4.97 | 4.43 | 3.96 | 3.78 | 4.03 |
Apple M3, Rust 1.97.1
FLOE segment-oriented performance in GiB/sec, higher is better.
| Provider | Segments | Encrypt into | Encrypt in place | Decrypt into | Decrypt in place |
|---|---|---|---|---|---|
aws-lc-rs |
1 MiB | 8.63 | 8.65 | 8.75 | 8.76 |
boring |
1 MiB | 7.22 | 7.20 | 6.08 | 7.15 |
ring |
1 MiB | 6.08 | 7.20 | 6.03 | 7.17 |
rustcrypto |
1 MiB | 3.56 | 3.63 | 3.62 | 3.65 |
aws-lc-rs |
4 KiB | 6.97 | 7.39 | 7.40 | 7.68 |
boring |
4 KiB | 6.51 | 6.22 | 5.82 | 6.64 |
ring |
4 KiB | 5.46 | 6.09 | 5.77 | 6.14 |
rustcrypto |
4 KiB | 3.45 | 3.47 | 3.48 | 3.45 |
Compared to "bare" AES-256-GCM from each provider, FLOE on the M3 is within ~1% on 1 MiB segments and ~15%-25% slower on 4 KiB segments.
License
Licensed under the Apache License 2.0.