internity 0.1.0

Blazingly fast string interning with compact handles, compact storage, and concurrent fill support
Documentation
# internity Performance Report

Generated by `scripts/perf_report.rs`. Head-to-head against the main Rust
interners.

- `cargo bench --bench internity_compare` — criterion wall-clock timings for three operations (`insert`, `reuse`, `lookup`), each in a **single-threaded** flavor and **multi-threaded** flavors at 1/2/4/8 threads.
- `cargo bench --bench internity_compare_cg` — gungraun / Callgrind instruction-precise counts for a **single** interning/lookup operation against a table of 1000 filler entries plus the fixed key (1001 preloaded), reported separately as microprobes (see below).
- `cargo bench --bench internity_mem` — live heap footprint of each interner (tracking global allocator).

**Methodology:** every timed region measures only insert/dedupe or lookup — setup and result destruction are kept outside the elapsed-time boundary. Refcounted `string_cache` atoms are retained across dedupe/lookup rounds so hits are measured against populated dynamic entries. `lookup` uses the same random order for all crates. Multi-threaded flavors run the op on `n` threads, barrier-timed so only the parallel work is counted (no thread spawn/join). Corpus ≈ 6000 identifier-like strings.  
**Δ vs internity:** `+x%` = that row is x% slower than internity on the same workload; `-x%` = faster.  
The process-global/cache-backed rows (`ustr`, `string_cache`) have no repeatable single-threaded Criterion `insert` timing in this suite and are therefore omitted from that table; their deterministic single-insert cost appears in the Callgrind microprobe section below.

## `internity_compare/insert` — single-threaded

| Variant | Time | Δ vs internity |
|---|---:|---:|
| `internity` | 259.40 µs | ref |
| `internity-threaded` | 459.03 µs | +77.0% |
| `lasso` | 623.22 µs | +140.3% |
| `string-interner` | 323.61 µs | +24.8% |
| `symbol_table` | 504.68 µs | +94.6% |

## `internity_compare/insert` — multi-threaded, 1 thread

| Interner | Time | Δ vs internity |
|---|---:|---:|
| `internity` | 700.70 µs | ref |
| `lasso-threaded` | 1.69 ms | +140.9% |
| `symbol_table` | 731.92 µs | +4.5% |

## `internity_compare/insert` — multi-threaded, 2 threads

| Interner | Time | Δ vs internity |
|---|---:|---:|
| `internity` | 1.44 ms | ref |
| `lasso-threaded` | 2.60 ms | +80.3% |
| `symbol_table` | 1.96 ms | +35.5% |

## `internity_compare/insert` — multi-threaded, 4 threads

| Interner | Time | Δ vs internity |
|---|---:|---:|
| `internity` | 2.03 ms | ref |
| `lasso-threaded` | 2.38 ms | +16.9% |
| `symbol_table` | 1.77 ms | -13.1% |

## `internity_compare/insert` — multi-threaded, 8 threads

| Interner | Time | Δ vs internity |
|---|---:|---:|
| `internity` | 2.23 ms | ref |
| `lasso-threaded` | 2.42 ms | +8.6% |
| `symbol_table` | 1.96 ms | -12.3% |

## `internity_compare/reuse` — single-threaded

| Variant | Time | Δ vs internity |
|---|---:|---:|
| `internity` | 110.33 µs | ref |
| `internity-threaded` | 181.17 µs | +64.2% |
| `lasso` | 287.61 µs | +160.7% |
| `string-interner` | 136.56 µs | +23.8% |
| `symbol_table` | 174.09 µs | +57.8% |
| `ustr` | 312.15 µs | +182.9% |
| `string_cache` | 279.44 µs | +153.3% |

## `internity_compare/reuse` — multi-threaded, 1 thread

| Interner | Time | Δ vs internity |
|---|---:|---:|
| `internity` | 296.39 µs | ref |
| `lasso-threaded` | 451.18 µs | +52.2% |
| `symbol_table` | 280.85 µs | -5.2% |
| `ustr` | 532.33 µs | +79.6% |
| `string_cache` | 461.37 µs | +55.7% |

## `internity_compare/reuse` — multi-threaded, 2 threads

| Interner | Time | Δ vs internity |
|---|---:|---:|
| `internity` | 682.57 µs | ref |
| `lasso-threaded` | 912.67 µs | +33.7% |
| `symbol_table` | 766.16 µs | +12.2% |
| `ustr` | 969.95 µs | +42.1% |
| `string_cache` | 1.21 ms | +77.3% |

## `internity_compare/reuse` — multi-threaded, 4 threads

| Interner | Time | Δ vs internity |
|---|---:|---:|
| `internity` | 1.35 ms | ref |
| `lasso-threaded` | 1.36 ms | +0.4% |
| `symbol_table` | 1.27 ms | -5.6% |
| `ustr` | 1.47 ms | +9.1% |
| `string_cache` | 1.86 ms | +37.9% |

## `internity_compare/reuse` — multi-threaded, 8 threads

| Interner | Time | Δ vs internity |
|---|---:|---:|
| `internity` | 2.68 ms | ref |
| `lasso-threaded` | 2.17 ms | -19.2% |
| `symbol_table` | 2.86 ms | +6.4% |
| `ustr` | 3.08 ms | +14.8% |
| `string_cache` | 2.76 ms | +2.8% |

## `internity_compare/lookup` — single-threaded

| Variant | Time | Δ vs internity |
|---|---:|---:|
| `internity` | 12.00 µs | ref |
| `internity-frozen` | 11.44 µs | -4.7% |
| `lasso` | 9.92 µs | -17.4% |
| `string-interner` | 14.53 µs | +21.1% |
| `symbol_table` | 49.70 µs | +314.1% |
| `ustr` | 7.49 µs | -37.6% |
| `string_cache` | 10.22 µs | -14.8% |

## `internity_compare/lookup` — multi-threaded, 1 thread

| Interner | Time | Δ vs internity |
|---|---:|---:|
| `internity` | 40.71 µs | ref |
| `lasso-resolver` | 39.81 µs | -2.2% |
| `symbol_table` | 86.92 µs | +113.5% |
| `ustr` | 42.23 µs | +3.7% |
| `string_cache` | 45.19 µs | +11.0% |

## `internity_compare/lookup` — multi-threaded, 2 threads

| Interner | Time | Δ vs internity |
|---|---:|---:|
| `internity` | 174.86 µs | ref |
| `lasso-resolver` | 158.30 µs | -9.5% |
| `symbol_table` | 325.80 µs | +86.3% |
| `ustr` | 159.20 µs | -9.0% |
| `string_cache` | 161.93 µs | -7.4% |

## `internity_compare/lookup` — multi-threaded, 4 threads

| Interner | Time | Δ vs internity |
|---|---:|---:|
| `internity` | 376.99 µs | ref |
| `lasso-resolver` | 376.40 µs | -0.2% |
| `symbol_table` | 630.72 µs | +67.3% |
| `ustr` | 379.67 µs | +0.7% |
| `string_cache` | 399.32 µs | +5.9% |

## `internity_compare/lookup` — multi-threaded, 8 threads

| Interner | Time | Δ vs internity |
|---|---:|---:|
| `internity` | 701.36 µs | ref |
| `lasso-resolver` | 733.77 µs | +4.6% |
| `symbol_table` | 1.19 ms | +69.2% |
| `ustr` | 748.70 µs | +6.7% |
| `string_cache` | 873.98 µs | +24.6% |

## Instruction-count microprobes (Callgrind)

These are **standalone** deterministic microprobes, **not** counterparts to the Criterion wall-clock tables above. Each measures a **single** interning or lookup operation against a table pre-populated with 1000 filler entries plus the fixed key (1001 preloaded), with `lasso`/`string-interner` pinned to fixed-seed hashers for reproducibility. Occupancy, key distribution, hasher configuration, and operation granularity all differ from the ≈6000-operation Criterion corpus loop, so the instruction counts here must **not** be read as a per-row explanation of the wall-clock timings.  
Mem accesses = L1 + LL + RAM hits (Callgrind D-cache references).

### `insert` — single operation

| Interner | Instructions | Mem accesses |
|---|---:|---:|
| `internity` | 220 | 307 |
| `internity-threaded` | 815 | 1,119 |
| `lasso` | 272 | 404 |
| `string-interner` | 198 | 277 |
| `symbol_table` | 490 | 723 |
| `ustr` | 164 | 230 |
| `string_cache` | 654 | 821 |

### `reuse` — single operation

| Interner | Instructions | Mem accesses |
|---|---:|---:|
| `internity` | 150 | 209 |
| `internity-threaded` | 162 | 224 |
| `lasso` | 172 | 251 |
| `string-interner` | 178 | 252 |
| `symbol_table` | 441 | 656 |
| `ustr` | 130 | 181 |
| `string_cache` | 292 | 356 |

### `lookup` — single operation

| Interner | Instructions | Mem accesses |
|---|---:|---:|
| `internity` | 30 | 51 |
| `internity-frozen` | 26 | 43 |
| `lasso` | 31 | 53 |
| `string-interner` | 40 | 62 |
| `symbol_table` | 510 | 806 |
| `ustr` | 4 | 7 |
| `string_cache` | 20 | 33 |

## Memory footprint

Live heap bytes held by each interner, measured with a tracking global allocator over the same corpus (`cargo bench --bench internity_mem`). `insert` is the filled interner; `lookup` is the read structure the lookup benchmark resolves against (the frozen read form where a crate has one).

```text
Corpus: 6000 distinct strings,     73.1 KiB of UTF-8 (avg 12.5 B/string)

== Owned interners — live heap held ==
  Interner                             insert       lookup
  internity  LocalLexicon           172.0 KiB     96.5 KiB
  internity  ThreadedLexicon        181.1 KiB     98.8 KiB
  lasso (live Rodeo)                264.2 KiB    264.2 KiB
  lasso (RodeoResolver)             264.2 KiB    224.2 KiB
  string-interner                   204.0 KiB    204.0 KiB
  symbol_table                      241.3 KiB    241.3 KiB

== Global caches — live heap held (persistent; insert == lookup) ==
  ustr                             8232.0 KiB
  string_cache                      351.7 KiB

Note: internity `lookup` is the frozen read form, which drops the string→handle map.
      lasso is reported as two rows: `live Rodeo` (matches the single-threaded `lookup`
      benchmark) and `RodeoResolver` (matches the `lookup-concurrent` `lasso-resolver`
      benchmark). `string-interner`/`symbol_table` have no frozen form so `lookup` == `insert`.
```