wai-quantum
A deterministic quantum stack in pure Rust. Byte-exact circuit simulation, the operational layer around it, and signed energy-accounted receipts binding every stage.
No QPU, no cloud service, no vendor SDK, and no system libraries — so the same source runs natively, in the browser via wasm, and as a WASI component at the edge, producing byte-identical results and receipts on all three.
[]
= "0.3"
use Circuit;
use quantum_toolchain as qt;
let mut c = new;
c.h.cx.cx; // GHZ
let sv = c.simulate.unwrap;
println!; // |000> 0.5, |111> 0.5
// a portable identity for the reconstruction — same on every machine
let h: = sv.statevector_hash;
What's in it
| Feature | What it gives you |
|---|
Cross-architecture reproducibility is tested, not asserted. The full test
suite passes under wasm32-wasip2 on wasmtime as well as natively, and the
floating-point results — entropy, the Holevo bound — are pinned to their exact
IEEE-754 bits, so a platform that computes something different fails the suite
instead of quietly returning a different number:
CARGO_TARGET_WASM32_WASIP2_RUNNER=wasmtime cargo test --release --target wasm32-wasip2 --features full
The same binary answers what a state carries, not only what a circuit does:
wai-quantum entropy bell --keep 0 # half a Bell pair: 1 bit, invisible to a state vector
wai-quantum channel depolarizing --prob 0.5
wai-quantum bb84 --eavesdrop # interception shows up in the error rate
wai-quantum chsh # 2 sqrt(2), exactly, not sampled
wai-quantum qutrit # d = 3, not just qubits
channel prints the Holevo bound before and after, and it can only fall: noise
never creates information.
| quantum | byte-exact statevector simulation (dyadic Clifford+T+P(k)) |
| quantum_toolchain | algorithm library, backend recommendation, OpenQASM 2/3 ingest of what the mainstream toolchains actually export — both register and both measurement spellings, rx/ry/rz/sx/u/u2/u3/crz, and exact rational-multiple-of-pi angles (non-dyadic angles are refused, never rounded) — Bloch / entanglement / purity analysis |
| quantum_stabilizer | stabilizer (CHP) tableau — O(n²), scales far past statevector |
| quantum_mps | tensor-network (matrix-product-state) backend |
| quantum_pauli | sparse-Pauli / Heisenberg observable propagation over the dyadic gate set (up to 64 qubits) |
| quantum_spd | sparse Pauli dynamics at utility scale: arbitrary-angle Pauli rotations on up to 1024 qubits, exact Clifford and quarter-turn rotations, a rigorous bound on what truncation can move a result, and the 127-qubit heavy-hex kicked-Ising circuits — whose exactly computed 5-step magnetization it reproduces on all 127 qubits |
| quantum_spd_receipt | a signed, energy-accounted receipt over a sparse-Pauli-dynamics run, re-checkable bit for bit on any machine |
| quantum_bptn | tensor-network states on the hardware graph: simple-update evolution with belief-propagation environments and re-gauging, the device's couplings applied a layer at a time in parallel (bit-identical to one at a time). Exact on trees; on the 127-qubit lattice it gives the exactly computed 5-step magnetization to 4e-7 at bond dimension 32 — an independent check on sparse Pauli dynamics, whose errors it does not share |
| quantum_resource | fault-tolerant resource estimation: from a program's logical counts (qubits, T, Toffoli and rotation gates, measurements) to code distance, T-state factories found by exhaustive search, physical qubits and runtime, on six reference qubits and three codes, with the qubits-versus-time frontier. Reproduces the published worked examples to their printed precision |
| quantum_resource_receipt | a signed, energy-accounted receipt over an estimate, re-checkable bit for bit on any machine |
| quantum_qio | quantum-inspired optimisation: Ising and Max-Cut problems (G-set reader, spin-glass generator), simulated bifurcation in its ballistic, discrete and heated forms (batched eight trials per pass, bit-identical to one at a time), simulated annealing, single-flip descent, an exact Gray-code referee up to 34 spins, and step-to-solution. Reaches the best-known cut on seven G-set instances |
| quantum_ldpc | qLDPC memory at circuit level: bivariate-bicycle codes from their two polynomials (k by rank, logical bases), the depth-8 syndrome cycle with circuit-level noise, its detector error model, and BP-OSD decoding (min-sum with adaptive scaling, ordered statistics with a combination sweep). The error model has the same columns as the source's own decoding matrices, and on identical syndromes the decoder makes the source's decisions shot for shot |
| quantum_arch | architecture-level estimation: yoked surface-code storage (row and grid parity checks over patches, hot and cold), magic-state cultivation read off the published simulation curves, CCZ factories, and a clock set by whichever is slowest of lattice surgery, the magic-state supply and the control loop. Assembles whole-machine estimates; its storage search matches the source's footprint tool layout for layout |
| quantum_receipt | signed, energy-accounted receipt over a reconstruction |
| quantum_ops | the operations/attestation layer |
| quantum_cal, quantum_control, quantum_noise | calibration, filter-function robust control, DD noise spectroscopy, Cycle-Benchmarking noise learning |
| quantum_mitigate | zero-noise extrapolation, readout M3, classical shadows |
| quantum_qec | CSS codes (distance verified by exhaustive search, not asserted) — rotated surface [[d²,1,d]], toric, bivariate-bicycle, and the [[7,1,3]] colour code, whose self-duality buys a transversal logical Hadamard — one round, no ancillas, no surgery — verified on the simulator, with the surface code shown failing the same move — with two decoders: union-find (weighted matching, single-shot, space-time with faulty measurement, and circuit-level with hook errors; imports and exports Stim-syntax detector error models, so other tools' noise models decode here and ours decode there) and Relay-BP (for qLDPC) |
| quantum_frame | noisy Clifford circuits at QEC scale: reads the ecosystem's circuit text; samples detection events by bit-packed Pauli-frame simulation (64 shots per word); derives the detector error model by backward sensitivity, refusing non-deterministic detectors; generates the rotated surface-code memory experiment byte for byte as the reference layout; and decodes by exact minimum-weight perfect matching (Edmonds' blossom on integer weights). Measures logical error rates and fits the code model resource estimates assume |
| quantum_tn | tensor-network contraction: circuits (any unitary on any qubits) become networks; a seeded hyper-optimiser finds contraction trees from randomised-greedy and multilevel-partition families, reconfigures them optimally in subtrees, and slices them to a memory cap with a second search around the slices; slices contract in a fixed order, threaded, to the same bits everywhere. Single amplitudes or open-output batches |
| quantum_gbs | Gaussian boson sampling, exactly: hafnians (power traces, Hessenberg + characteristic polynomial, O(n³ 2^{n/2})), torontonians, lossy Gaussian states from squeezers and an interferometer, exact click and photon-number probabilities, exact samples drawn mode by mode, and a check of observed samples against the exact one- and two-mode click statistics |
| quantum_sv | the dense state vector at any angle: one-, two- and k-qubit kernels with a diagonal fast path, gate fusion around multi-qubit gates, threads over independent blocks; every kernel path computes each amplitude in the same order, so threading never changes a bit. Agrees with an independent simulator to 2·10⁻¹⁶ |
| quantum_compile | Clifford routing + stabilizer-tableau equivalence proof |
| quantum_atom | neutral-atom register preparation (Hungarian / LSAP) |
| quantum_qir | QIR export + ingest — emit and parse QIR (the QIR Alliance's LLVM-based IR); emitted modules validated with llvm-as, and 9 of 10 catalog algorithms re-ingest to a bit-identical statevector |
| quantum_comm | key distribution and security: BB84 with an intercept-resend eavesdropper (a quiet channel gives a perfect key; interception costs a quarter of it), and the no-cloning theorem priced |
| quantum_qudit | qubits and qutrits: the Weyl–Heisenberg group, the Fourier gate over Z_d, the controlled sum — exact for d = 2 and 3 |
| quantum_info | mixed states: density matrices, the partial trace, purity, exact Pauli expectations, and CHSH against the classical and Tsirelson bounds |
| quantum_source | the limits on encoding and transport: von Neumann entropy, the Schumacher limit, and the Holevo bound — which says n qubits carry at most n classical bits, so quantum does not compress classical media |
| quantum_channel | transport itself: Kraus channels (bit flip, dephasing, depolarizing, amplitude damping), and linear-inversion tomography to see what actually arrived — with the law that a channel can only lose information, never create it |
| quantum_qbom | Quantum Bill of Materials — stage receipts into one manifest |
| quantum_vml, quantum_phasor, quantum_qfhrr, quantum_kernel, quantum_qdata, quantum_phasor_meter | the interference / phasor ML layer |
| full | everything |
Default features are quantum, quantum_toolchain, quantum_receipt.
Speed
Byte-exactness is the contract, so the simulator is optimised only in ways that
cannot move a single amplitude: identity does nothing, a diagonal gate phases half
the state with one multiply rather than four, and X — so CX — is a swap with no
arithmetic at all. fxmul(ONE, x) == x exactly in this fixed point, which is what
makes those paths provably identical rather than merely close, and they are checked
against the previous implementation amplitude by amplitude.
An all-real matrix — H, and any real rotation — halves its multiplies too, since
the imaginary cross terms are multiplications by exactly zero.
At 20 qubits that is 8.9x on Z, 7.1x on X, 4.6x on CZ, 4.3x on CX, 3.4x on
T, 1.9x on Y and 1.8x on H.
Utility scale
quantum_spd evolves an observable instead of a state, so a circuit costs what
its non-Clifford rotations make it cost rather than 2ⁿ:
computes ⟨Z₆₂⟩ after 20 Trotter steps of the 127-qubit heavy-hex kicked-Ising
circuit in a fraction of a second, with the bound on how far truncation can have moved
it. --example utility_mz checks the 5-step magnetization against the exactly
computed values for every angle, and tests/utility_reference.rs holds the
published references. quantum_bptn computes the same circuits as tensor
networks (--example utility_bptn), so a result can be checked by a second
method that does not share the first one's approximation.
Fault-tolerant resources
What a program would need on a machine that does not exist yet is a resource
estimate, and claims about that machine stand on one. quantum_resource makes
the estimate a deterministic function of its inputs:
|
The model is the planar-ISA one of arXiv:2211.07629: surface and Floquet-code patches, lattice surgery, and 15-to-1 distillation factories chosen by exhaustive search over unit types, distances and copy counts. The tests reproduce that source's own tables:
- every code distance and factory of its per-qubit factoring table, down to the copy counts;
- the distance, factory count, physical qubits and runtime of all twelve factoring and chemistry rows of its summary table;
- all twelve factory counts of its dynamics rows. Factoring on the
(ns, 10⁻⁴)qubit comes out atd = 13, 18 factories, 8,716,258 physical qubits and 17 h 43 min.
Holding the model to those tables exposed places where the source's prose and its numbers disagree, and one factory that is not the minimum of its own model. Each is documented in the module and settled by the tables.
A sealed estimate names its whole job (counts, qubit, code, budget, slowdown and search limits, to the bit). Re-running it anywhere must reproduce every figure, down to each factory round, and the bytes are the same natively and as a WASI component.
Optimisation
Optimisation is the claim made most often for quantum hardware, and every
such claim is a comparison against a classical solver. quantum_qio is a
re-runnable one. It offers:
- simulated bifurcation (ballistic, discrete, and both with heating);
- simulated annealing;
- single-flip descent;
- an exact referee that settles small instances.
A result is a spin configuration, so anyone can re-check its cut or energy in one pass. On the G-set Max-Cut benchmark the best-known cut is reached on G1, G2, G3, G11, G14, G22 and G43. On G55 the best found is 10,292 against 10,299. The module documents each setting and success rate, and the bound that sparse graphs put on the time step.
qLDPC memory
High-rate codes are where most of the projected qubit savings in fault
tolerance come from, and their headline numbers are logical failure rates per
syndrome cycle under circuit noise. quantum_ldpc re-derives them for the
bivariate-bicycle family (arXiv:2308.07915):
- the codes are built from their polynomials;
- the depth-8 syndrome cycle is written out gate by gate;
- the circuit is sampled with the Pauli-frame simulator;
- its error model is decoded with belief propagation plus ordered statistics.
The source published its simulation software, and running it gives the oracle. On the 72-qubit code the derived error model has the source's columns, detector for detector, and on identical syndromes the decoder makes the source's decision shot for shot. Logical failure rates track the source's; the one open residual, 2.3σ on one half at the highest noise, is documented in the module. On the 144-qubit code the rates land on the source's published curve.
Architecture
Most of the recent fall in projected fault-tolerant cost comes from the
architecture around the code, not from better qubits. Idle qubits sit in yoked
storage (arXiv:2312.04522). T states are cultivated inside one patch rather
than distilled (arXiv:2409.17595). The clock runs at the pace of the slowest of
lattice surgery, the magic-state supply and the control loop. quantum_arch
prices each of these.
It re-runs the 2025 estimate that 2048-bit factoring takes under a week on under a million noisy qubits (arXiv:2505.15917).
- Reproduced. Its 897,864 physical qubits, its storage density of 429.03 qubits per logical qubit, and its Toffoli counts for every tabulated key size, each to the reference implementation's own rounding.
- Corrected. Re-running it surfaced six places where the source disagrees with itself, each documented in the module. Its hot-storage count is below its own peak formula. Its cold storage assumes fractional blocks. Its printed runtime is 14% above its own tallies. With every correction applied, the machine needs 958,480 qubits and 4.3 days. The conclusion holds.
Error correction at circuit level
An error-correction experiment is a noisy Clifford circuit, its detectors and
its logical observable. quantum_frame samples it, derives its detector error
model, and decodes it exactly:
The model is checked against an independent implementation:
- the generated circuits are identical, byte for byte, across distances 2–15 and 1–8 rounds;
- the derived error models match mechanism for mechanism, to 10⁻¹² in probability, across surface codes in both bases (rotated and unrotated) and repetition codes;
- decoded logical error rates agree point by point within sampling error.
tests/frame_reference.rs keeps two of those fixtures in the repository. The
fitted code model comes out at a ≈ 0.023, p* ≈ 0.0100. That is a measured
value for the constants a resource estimate (quantum_resource) otherwise takes
on trust, and it can be fed straight into one.
State vector
quantum_sv runs any gate at any angle on up to 32 qubits, in double
precision. gates provides the standard gates by name (H, S, T, √X, √Y, √W,
rotations, CX, CZ, iSWAP, fSim, …) for it and for the tensor networks. A
26-qubit, depth-20 random circuit runs in about nine seconds on one machine,
and the bits do not depend on how many threads ran it.
Tensor networks
An amplitude of a circuit is a tensor network, and the order it is contracted
in decides whether it costs 2³⁰ operations or 2⁶⁰. quantum_tn searches for
that order, slices it to fit memory, and contracts it:
Circuits are JSON lines, one gate per line: {"q":[a,b],"m":[re,im, …]}.
On single amplitudes of random circuits on the 54-qubit advantage-experiment
grid, the search finds trees of 10¹⁰·⁵ operations at 10 cycles and 10¹⁸·⁸ at
20, in seconds. An independently developed hyper-optimiser, given the same
networks, lands within a factor of 2.5 of these at every depth, in either
direction. The 54-qubit, 10-cycle amplitude, contracted here in 16 slices,
agrees with an independent contraction to 12 significant figures. The module
documentation states what slicing costs on deep circuits and what the scalar
executor can and cannot contract.
Gaussian boson sampling
A photonic sampler's outcome probabilities are hafnians (photon-number
detection) or torontonians (threshold detection) of its Gaussian state.
quantum_gbs computes both exactly. They agree with an independent
implementation:
- 10⁻¹² for probabilities and for hafnians up to
20 × 20; - 10⁻⁷ at
36 × 36.
--check holds a file of observed click patterns to the state's exact one-
and two-mode statistics. Those are computable for any number of modes, and
every sampler of the state must reproduce them.
Determinism
Simulation is byte-exact: amplitudes are dyadic fixed-point, so a circuit
reconstructs to the same statevector_hash on every machine and every target.
The approximate methods (MPS truncation, sparse Pauli dynamics, the phasor/ML
layer) are reproducible f64: IEEE-754-strict, and identical given the same
inputs and seed on every platform. They are documented as such rather than
claimed byte-exact. Their sin, cos, ln and atan2 are computed in-crate
from + − × ÷, because platform math libraries differ in the last place — enough
that four of these layers once gave different bits natively and as a WASI
component. tests/bit_identity.rs pins one result from each, so any platform
whose arithmetic differs fails there.
A receipt separates the two halves of a cost honestly: work is
portable-exact (n_ops · 2^n amplitude updates, recomputed by the verifier, so a
sink cannot inflate it) and energy is measured-attested (only the signer can
vouch for its own silicon). Signatures travel between substrates because the
computation is reproducible — not because two runtimes agreed to trust each other.
An energy figure is only comparable with another when it says how it was
acquired — read from the silicon's own counter, from a calibrated instrument,
from a model, or a constant. quantum_energy::Labelled seals that acquisition
class, with its declared uncertainty, inside the receipt's signature, and
LabelledQbom does the same per entry of a bill of materials. A receipt sealed
without a label keeps exactly the bytes it always had. MaybeLabelled parses
and verifies a receipt in whichever form it arrives, and checks a job's link to
its calibration in either form. A label's declared uncertainty must be at
least the 0.5/√3 µJ (≈ 0.289 µJ) that rounding the figure to whole
microjoules adds: joules_micro × relative_ppm of at least 288 676.
Also available
- A CLI (
wai-quantumin thewaicrate) — run circuits, emit OpenQASM, seal and verify receipts, native or underwasmtime. - A WASI-HTTP component serving the stack as an ordinary HTTP handler.
- Browser demos and a course: https://wai.transaction.science/quantum-ide
License
Apache-2.0