gpu snark: g16
- Snarkjs compatible .zkey and .wtns format
- CPU (pure rust)
- Apple Metal
- NVIDIA CUDA
- WebGPU (demo)
- BN254 support
- Backends behind a rust feature flags
- Library support
Backend support and validation covers work modes, ceremony capabilities, failure policies and the hardware actually tested.
benchmarks
apple-m2-max
note: our CPU benchmarks appear faster than rapidsnark (CPU only) on apple silicon
proving (warm):
blank = not measurable on this box: snarkjs has no warm mode, its CLI starts a fresh process per proof. Its cold numbers are in
bench/results/machines/.
trusted setup (faster of cpu/metal in bold):
phase 1, powers of tau, run once for every circuit that follows:
| command | input | g16 cpu (s) | g16 metal (s) | snarkjs (s) |
|---|---|---|---|---|
ptau new |
power 14 | 0.01 | — | 0.5 |
ptau contribute |
power 19 | 36.3 | 5.2 | 162.3 |
ptau beacon |
power 19 | 36.2 | 5.2 | 162.0 |
ptau prepare |
power 20 | 929.8 | 83.8 | ~9,300 |
ptau verify |
power 20 | 16.9 | — | 116.4 |
phase 2, the proving key, run once per circuit:
| command | domain | g16 cpu (s) | g16 metal (s) | snarkjs (s) |
|---|---|---|---|---|
setup circom |
2^20 | 17.4 | 17.1 | 86.4 |
zkey contribute |
2^20 | 16.4 | 1.8 | 72.6 |
zkey beacon |
2^20 | 16.3 | 1.8 | 73.6 |
g4dn.2xlarge
warm:
blank = not measurable on this box: snarkjs has no warm mode, its CLI starts a fresh process per proof. Its cold numbers are in
bench/results/machines/.
note: constraint count is a poor predictor of proving time. e.g.
railgun-13x01has fewer constraints thankeccak256but takes more to prove.
to reproduce the benchmarks you can use the script on machine of interest:
npm --prefix bench test checks these warm proving tables against
bench/results/machines/ and runs the table-check regressions, without running proofs.
The separate trusted setup tables are not covered. Suspect runs cannot support a table;
multiple records for one machine require an explicit source selection rather than silently
picking a newer run.
using it from the command line
The package and the binary are both snarkrs, a drop-in for the snarkjs 0.7.6 command line on Groth16 and
BN254: the same commands, aliases, positional file names and defaults, -e=/-n=/-v
options, exit codes (0 ok, 1 failed or invalid, 99 bad usage) and [INFO] snarkJS: OK!
log lines. snarkrs --help lists what it runs; a snarkjs command it does not run yet
fails with exit 1 and says so.
The short aliases work too (ptn, ptc, pt2, g16s, zkc, zkev, g16p, g16v, ...),
and the ceremony commands generally take --backend cpu|metal. CUDA supports only the
ptau prepare group FFT, not ceremony MSM or key scaling; WGPU supports none of those
three ceremony seams. The output is what snarkjs expects, so the two are interchangeable
in either direction:
witnesses
wtns calculate and groth16 fullprove take circom's circuit_js/circuit.wasm or the
native binary circom --c builds. The file's first bytes decide which, not its name:
The wasm runs in wasmtime, reads input.json the way snarkjs does, prints the circuit's
log() lines and its errors the way snarkjs does, and writes the same .wtns byte for
byte. fullprove keeps that witness in memory. A native binary runs as a subprocess and
fullprove hands it a private temp file that is zeroed and deleted once read. wtns debug
is wasm only. --no-default-features --features cpu,wgpu builds without the
witness-wasm feature: no wasmtime, native binaries only.
On Apple Silicon the C++ from circom --c does not build as it comes: fr.cpp passes
uint64_t* where GMP takes mp_limb_t*, which is unsigned long there, and main.cpp
includes nlohmann/json.hpp, which the generated Makefile never points at. --no_asm
leaves out fr.asm and the nasm it needs.
--stage-timings prints where the time went, which is the fastest way to find out whether
a circuit is MSM-bound or transform-bound:
using it as a library
= { = "https://github.com/zemse/snarkrs", = ["metal"] }
Proving needs an optimised build. In a debug build prove returns ProveError::Unoptimized
rather than run tens of times slower, so optimise dependencies there too, which keeps your
own code quick to compile and debug:
[]
= 3
The prover, verifier, key formats and CPU backend are always in. The rest is opt in, so a build compiles only the backend it runs on:
| feature | adds |
|---|---|
metal |
snarkrs_lib::metal, Apple GPUs |
cuda |
snarkrs_lib::cuda, NVIDIA GPUs |
wgpu |
snarkrs_lib::wgpu, WebGPU |
ceremony |
snarkrs_lib::ceremony, powers of tau and phase 2 |
witness |
snarkrs_lib::witness, circom's native witness binary |
witness-wasm |
witness plus circom's circuit.wasm on wasmtime |
use ;
// warm pk once for repeated proving the same circuit
let pk = load?;
let n_public = pk.n_public;
let circuit = new?.prepare?;
let w = load?.0;
let mut t = default;
let proof = prove?;
let public = &w;
verify?;
write_proof?;
write_public?;
Swap MetalBackend for snarkrs_lib::CpuBackend or snarkrs_lib::cuda::CudaBackend to change where
it runs; nothing else in the snippet changes, which is the point of the trait.
circuit.key() preserves the dimensions, full verification key (including IC), and the
five assembly headers (alpha_g1, beta_g1, beta_g2, delta_g1, delta_g2). Bulk
coefficients and query vectors are backend-dependent and may be released after upload.
It is not guaranteed to be a complete key for preparing another backend; retain or reload
the original key for that.
The blinders come from the OS CSPRNG. There is no seed override on this path on purpose: a
reused (r, s) across two proofs of different witnesses leaks the witness, so the
deterministic entry point stays test-only.
witness from memory
prove takes the witness as &[Fr], so it never has to be a file:
use ;
use ;
let pk = load?;
let circuit = new?.prepare?;
// compile the wasm once, then one witness per input (feature `witness-wasm`)
let calc = from_file?;
let w = calc.calculate?;
let mut t = default;
let proof = prove?;
w is just 1, the public signals, then the other wires in circom's order (the second
column of the .sym file). Anything that computes it can feed prove, so the fastest
witness generator is one written in optimised Rust for your circuit, with no wasm and no
.wtns written and read back in between. snarkrs prints a tip saying so on stderr after
wtns calculate and fullprove; SNARKRS_NO_TIPS=1 turns it off.
what it checks
These APIs have checked defaults and explicit _unchecked alternatives for input you
already trust. This is not a claim that every entry point has an unchecked twin:
| checked | what it adds | unchecked |
|---|---|---|
ProvingKey::load, from_bytes |
every point on the curve; refuses a key with no phase-2 contribution | load_unchecked, from_bytes_unchecked |
VerifyingKey::from_json |
refuses a key anyone can forge against | from_json_unchecked |
verify |
proof points on the curve, in the subgroup, not infinity | verify_unchecked |
prove |
verifies its own proof before returning it | prove_unchecked |
The check in prove is also what stops a hostile zkey from reading the witness out of the
proof, so a key from someone else should only ever meet prove, and should still be checked
with snarkrs zkey verify against the circuit and the ptau. Proving time depends on how many
witness entries are zero or one unless you pass --constant-work (cpu, metal, wgpu and cuda
backends, with circuit-dependent overhead). This fixes the scalar-dependent MSM sizing
and disables zero/one fast paths, not all witness-dependent work: GPU bucket occupancy,
atomics contention and accumulation still depend on the witness. Constant-work is not
constant-time. CUDA's bounded Tesla T4 validation and measured overhead are in the
backend matrix. The audit and its current status are in
security/README.md.