runsync-transfer 0.1.0

High-throughput P2P file transfer engine: adaptive compression, end-to-end AEAD, parallel chunked pipeline over QUIC or any async transport.
Documentation
# runsync-transfer

A peer-to-peer file transfer engine in Rust. The sender compresses and encrypts
chunks, the receiver decrypts and decompresses them, and both ends spread the
work across every core while the link stays busy.

Library only — no UI, no binary you have to run. It plugs into an application
that already owns a QUIC connection.

```toml
[dependencies]
runsync-transfer = "0.1"
```

## What it does

| | |
|---|---|
| **Parallel** | Chunks compress, encrypt, and land on disk concurrently across N streams and N CPU workers. No shared file cursor, no reassembly buffer. |
| **Adaptive compression** | Per-chunk entropy probe plus an extension table, and a per-file verdict so a file that does not compress stops being tried. `.flac` and `.mp4` ship raw; text and PCM get compressed. Incompressible input is never inflated. |
| **Audio coder** | A lossless LPC coder for every PCM layout a WAV can hold — 8-bit unsigned, 16/24/32-bit signed, 32-bit float. zstd manages ~1.03× on real `.wav`; this reaches **1.58×** across a real music library, beating reference FLAC's ~1.53× on the same files. Every chunk is decoded and compared before it is accepted, so a codec bug can cost speed but never a file. |
| **End-to-end encryption** | X25519 + HKDF-SHA256 + AES-256-GCM or ChaCha20-Poly1305, *inside* the transport's TLS, so a relay carries bytes it cannot read. |
| **Large files** | 64-bit offsets, positional I/O, constant memory. A 100 GB file costs the same resident bytes as a 100 MB one. |
| **Many files** | One manifest, one connection, chunks from different files interleaved across streams. |
| **Sparse files** | All-zero chunks cross as a flag. A 100 GB VM image with 1 GB of data transfers as ~1 GB and keeps its holes. |
| **Delta sync** | The receiver hashes what it already has; the sender recognises unchanged chunks and sends 28 bytes instead of a megabyte. Re-sending an unchanged 3.4 GiB tree moves **0.1 MiB** in **4.2 s** — chunk hashes are cached between runs, and the partial file is seeded by a copy-on-write clone, so an unchanged file costs a `stat`. |
| **Resume** | An interrupted transfer restarts from the chunks already on disk. |
| **Verification** | BLAKE3 per chunk, folded into a per-file root, checked before the file is renamed into place. |

## Using it

The engine talks to a [`Transport`], not to a socket. If your application
already has a `quinn::Connection`, hand it over — no second endpoint, no extra
handshake, no separate certificate story:

```rust
use runsync_transfer::{send, receive, Config, Source, QuicTransport};
use std::sync::Arc;

// Sending side
let transport = Arc::new(QuicTransport::from_connection(conn));
let stats = send(transport, &[Source::new("/data/album")], &Config::default(), None).await?;
println!("{stats}");   // 412/412 files, 8.11 GiB of 8.11 GiB (100.0%), 240 MiB/s …

// Receiving side
let transport = Arc::new(QuicTransport::from_connection(conn));
let stats = receive(transport, "/dest", &Config::default(), None).await?;
```

`send` and `receive` never close the connection. It is yours: multiplex other
traffic over it, or run another transfer on it when this one returns.

### Progress

```rust
let cb: runsync_transfer::ProgressFn = Arc::new(|p| {
    println!("{:.1}% — {}/s — eta {:?}", p.fraction() * 100.0,
             runsync_transfer::human_bytes(p.throughput() as u64), p.eta());
});
send(transport, &sources, &cfg, Some(cb)).await?;
```

### Encryption

`Secrecy::TransportOnly` (the default) relies on QUIC/TLS 1.3. That is the right
answer when both endpoints terminate their own TLS and you trust every hop. When
traffic crosses a relay you do not control, add the end-to-end layer:

```rust
use runsync_transfer::{Config, Secrecy, crypto};

// Symmetric: both sides hold the same 32 bytes, shared out of band.
let cfg = Config::default().with_secrecy(Secrecy::Psk(psk));

// Or asymmetric: each side pins the other's X25519 public key.
let (our_secret, our_public) = crypto::generate_identity();
let cfg = Config::default().with_secrecy(Secrecy::Static {
    our_secret,
    peer_public,
});
```

Both give forward secrecy — the session keys come from a fresh ephemeral
exchange, so recorded traffic stays unreadable even if the PSK or identity key
later leaks. A peer cannot unilaterally downgrade to `TransportOnly`: mismatched
modes fail the handshake.

### Bringing your own transport

Implement four methods and the engine will run over anything:

```rust
#[async_trait::async_trait]
impl Transport for MyTransport {
    async fn open_uni(&self)   -> Result<BoxSend> { .. }   // data streams
    async fn accept_uni(&self) -> Result<BoxRecv> { .. }
    async fn open_bi(&self)    -> Result<(BoxSend, BoxRecv)> { .. }  // control
    async fn accept_bi(&self)  -> Result<(BoxSend, BoxRecv)> { .. }
    fn close(&self, code: u32, reason: &[u8]) { .. }
}
```

`transport::mem` implements it over in-process pipes, which is how the test
suite runs the full protocol without a socket.

## Tuning

```rust
Config::default()          // 1 MiB chunks, adaptive zstd — a sane middle
Config::throughput()       // 4 MiB chunks, LZ4 — for a link faster than your CPU
Config::bandwidth_saving() // zstd level 9 — for a slow or metered link
```

The knob that matters most on a long fat path is `streams`. The one that bounds
memory is `streams × queue_depth × chunk_size` — see `Config::memory_budget()`.

The bundled QUIC helpers already raise the flow-control windows
(`bulk_transport_config()`); stock QUIC defaults are sized for request/response
traffic and will cap a bulk transfer well below link rate on a high-latency
path. If you supply your own `quinn::Connection`, apply the same treatment or
transfers over a fat pipe will be flow-control bound.

## Performance

Against reference zstd 1.5.7 at the same level and block size
(`zstd -b3 -B1048576`, the comparison that isolates our overhead rather than the
codec's), ratios match to the digit and speed matches or beats it nearly across
the board:

| corpus | ratio (zstd → ours) | compress MB/s | decompress MB/s |
|---|---|---|---|
| `logs.txt` | 7.06 → 7.06 | 819 → **835** | 2317 → **3067** |
| `source.txt` | 13.15 → 13.15 | 1422 → **1454** | 4664 → **6048** |
| `audio_pcm.wav` | 1.11 → 1.11 | 506 → **608** | 1180 → **1416** |
| `media.flac` | 1.00 → 1.00 | 9250 → 8118 | 31155 → **50961** |

In the engine's normal adaptive mode `media.flac` runs at **33.7 GB/s**, because
the fastest way to compress entropy-coded data is not to. A 100 GiB disk image
holding 1 GiB of data crosses as ~1 GiB and keeps its holes.

Full methodology, per-stage costs, the zero-byte cases, and the optimisation
history are in [BENCHMARKS.md](BENCHMARKS.md).

## Design notes

**Chunks are independent.** Each is compressed and sealed on its own. That costs
a little compression ratio versus one solid stream, and buys three things the
engine needs: any worker can process any chunk in any order, a resumed transfer
can skip individual chunks, and a corrupt chunk cannot poison the ones after it.

**Frames are self-describing.** File id, chunk index, and lengths ride in a
28-byte header, so frames may arrive on any stream in any order. The receiver
writes each one at its own offset with `pwrite` and never buffers a file.

**Nothing is committed early.** Files are written to a `.rst-part` sidecar and
renamed into place only after the last chunk lands and the hash matches. A crash
leaves an obvious partial file that resume picks up, never a truncated file that
looks finished.

**Resume state is durable in the right order.** A chunk is recorded as present
only after the data file has been flushed. The reverse order would produce a
state file claiming chunks a power loss discarded.

**Hashing is order-independent.** A flat BLAKE3 over a file has to be computed
sequentially, which would serialise a pipeline whose whole point is finishing
chunks out of order. Instead each chunk is hashed on its own and the ordered
concatenation of chunk hashes is hashed again. Equally strong end to end, and
both sides can compute it in parallel.

## The receiver treats the manifest as hostile

A transfer is a remote peer writing to your filesystem.

- Paths are validated by component. Absolute paths, `..`, root components,
  Windows prefixes, and all-dot names are rejected. Backslashes are normalised
  first, so `..\..` cannot slip past on a platform that honours them.
- Symlinks are created **last**, after every regular file is committed, and only
  when the target lexically resolves inside the destination. Creating them
  earlier would let a link to `/etc` turn a later write into a write outside the
  root.
- Every length read off the wire is bounded before it sizes an allocation.
- A frame naming an unknown file, an out-of-range chunk, or a plaintext length
  that disagrees with its position is a protocol error.
- On an encrypted session, an unsealed frame is rejected rather than accepted as
  a downgrade.

`tests/hostile.rs` drives the receiver with a peer that tries each of these.

## Testing

```
cargo test --features test-certs        # unit + integration + adversarial + QUIC
cargo build --release --example rst --features test-certs
```

The `rst` example is a CLI for driving the engine over a real network:

```
rst gen      --out DIR --profile mixed|large|sparse|manysmall --size 10G
rst serve    --dest DIR --addr 0.0.0.0:5555 --psk HEX
rst send     --addr HOST:5555 --cert-hex HEX --psk HEX PATH...
rst selftest --size 10G [--transport mem]
rst verify   --a DIR --b DIR
```

## License

[PolyForm Noncommercial 1.0.0](LICENSE) — use, modify and share it freely for
any noncommercial purpose. Commercial use needs a separate licence from the
copyright holder.

Chosen over a Creative Commons NonCommercial licence deliberately: CC licences
are written for prose, and their NoDerivatives variants forbid modification,
which would make a library impossible to use and would have blocked
AGPL-licensed consumers outright. PolyForm is the same intent expressed in terms
that work for software.