internity 0.2.1

Blazingly fast string interning with compact handles, compact storage, and concurrent fill support
Documentation
# internity Performance Report

Generated by [`scripts/perf_report.rs`](../scripts/perf_report.rs). Re-run it to refresh these numbers.

All timing figures are wall-clock medians measured by Criterion. They are
machine-dependent: the ratios between rows are the durable signal, not the
absolute values.

This report is a curated set of customer-facing scenarios. The crate also carries
internal Callgrind instruction-count benches in the consolidated metabench target
(`benches/internity.rs`) used for optimization work, which are not published here;
run them with `cargo bench --bench internity -- --gungraun`.

**Workload:** a corpus of ≈6000 identifier-like strings, exercised through three customer operations — `insert` (interning a string for the first time), `reuse` (interning a string that is already interned) and `lookup` (resolving a handle back to its string).

**Methodology:** every timed region measures only insert/dedupe or lookup — benchmark setup and result destruction are kept outside the elapsed-time boundary. Refcounted `string_cache` atoms are retained across dedupe/lookup rounds so hits are measured against populated dynamic entries. `lookup` uses the same random order for all crates. The concurrent flavors run the operation on `n` threads and are barrier-timed, so thread spawn/join is excluded and only the parallel work is counted.  
**Δ vs internity:** `+x%` = that row is x% slower than internity on the same workload; `-x%` = faster.  
The process-global, cache-backed rows (`ustr`, `string_cache`) keep their table alive for the whole process, so they have no repeatable first-time `insert` timing and are absent from the single-threaded `insert` table.

## Single-threaded interning

One table per operation, comparing internity against every other interner measured for that operation on a single thread.

### `insert` — single-threaded

| Interner | Time | Δ vs internity |
|---|---:|---:|
| `internity` | 509.43 µs | ref |
| `internity-threaded` | 598.58 µs | +17.5% |
| `lasso` | 803.13 µs | +57.7% |
| `string-interner` | 394.88 µs | -22.5% |
| `symbol_table` | 667.99 µs | +31.1% |

### `reuse` — single-threaded

| Interner | Time | Δ vs internity |
|---|---:|---:|
| `internity` | 172.99 µs | ref |
| `internity-threaded` | 300.31 µs | +73.6% |
| `lasso` | 307.40 µs | +77.7% |
| `string-interner` | 112.97 µs | -34.7% |
| `symbol_table` | 180.79 µs | +4.5% |
| `ustr` | 257.11 µs | +48.6% |
| `string_cache` | 288.79 µs | +66.9% |

### `lookup` — single-threaded

| Interner | Time | Δ vs internity |
|---|---:|---:|
| `internity` | 10.17 µs | ref |
| `internity-frozen` | 9.97 µs | -2.0% |
| `lasso` | 9.58 µs | -5.8% |
| `string-interner` | 11.47 µs | +12.8% |
| `symbol_table` | 49.37 µs | +385.5% |
| `ustr` | 8.19 µs | -19.4% |
| `string_cache` | 9.35 µs | -8.0% |

## Concurrent scaling

One table per operation, with a column per thread count. Each cell is the wall-clock median of the whole parallel phase at that thread count, so the total work grows with the thread count and the numbers show how each interner holds up under contention rather than a per-thread speedup. The trailing column is the delta against internity at the highest thread count, which is where the designs separate most; per-thread-count deltas are omitted to keep the tables readable.

### `insert` — concurrent

| Interner | 1 thr | 2 thr | 4 thr | 8 thr | Δ vs internity @ 8 thr |
|---|---:|---:|---:|---:|---:|
| `internity` | 598.07 µs | 1.59 ms | 2.10 ms | 2.45 ms | ref |
| `lasso-threaded` | 1.60 ms | 3.07 ms | 2.75 ms | 2.74 ms | +12.0% |
| `symbol_table` | 699.97 µs | 2.03 ms | 2.04 ms | 2.25 ms | -8.3% |

### `reuse` — concurrent

| Interner | 1 thr | 2 thr | 4 thr | 8 thr | Δ vs internity @ 8 thr |
|---|---:|---:|---:|---:|---:|
| `internity` | 291.91 µs | 890.36 µs | 1.27 ms | 2.87 ms | ref |
| `lasso-threaded` | 490.32 µs | 893.87 µs | 1.31 ms | 1.93 ms | -32.9% |
| `symbol_table` | 311.90 µs | 757.93 µs | 1.40 ms | 2.77 ms | -3.3% |
| `ustr` | 460.31 µs | 690.44 µs | 1.67 ms | 3.04 ms | +5.9% |
| `string_cache` | 538.54 µs | 1.30 ms | 1.95 ms | 4.30 ms | +49.8% |

### `lookup` — concurrent

| Interner | 1 thr | 2 thr | 4 thr | 8 thr | Δ vs internity @ 8 thr |
|---|---:|---:|---:|---:|---:|
| `internity` | 88.38 µs | 313.37 µs | 557.64 µs | 1.35 ms | ref |
| `lasso-resolver` | 97.14 µs | 253.20 µs | 409.29 µs | 1.11 ms | -17.8% |
| `symbol_table` | 181.39 µs | 528.90 µs | 1.32 ms | 2.74 ms | +102.4% |
| `ustr` | 137.81 µs | 158.25 µs | 532.71 µs | 1.13 ms | -16.6% |
| `string_cache` | 188.22 µs | 192.22 µs | 745.06 µs | 1.26 ms | -7.1% |

## Memory footprint

Live heap bytes held by each interner, measured with a tracking global allocator over the same corpus (`cargo bench --bench internity_mem`). `insert` is the filled interner; `lookup` is the read structure the lookup benchmark resolves against (the frozen read form where a crate has one). Lower is better.

```text
Corpus: 6000 distinct strings,     73.1 KiB of UTF-8 (avg 12.5 B/string)

== Owned interners — live heap held ==
  Interner                             insert       lookup
  internity  LocalLexicon           172.0 KiB     96.5 KiB
  internity  ThreadedLexicon        181.1 KiB     98.8 KiB
  lasso (live Rodeo)                264.2 KiB    264.2 KiB
  lasso (RodeoResolver)             264.2 KiB    224.2 KiB
  string-interner                   204.0 KiB    204.0 KiB
  symbol_table                      241.3 KiB    241.3 KiB

== Global caches — live heap held (persistent; insert == lookup) ==
  ustr                             8232.0 KiB
  string_cache                      351.7 KiB

Note: internity `lookup` is the frozen read form, which drops the string→handle map.
      lasso is reported as two rows: `live Rodeo` (matches the single-threaded `lookup`
      benchmark) and `RodeoResolver` (matches the `lookup-concurrent` `lasso-resolver`
      benchmark). `string-interner`/`symbol_table` have no frozen form so `lookup` == `insert`.
```