Skip to main content

Crate snarkrs_gpu_layout

Crate snarkrs_gpu_layout 

Source
Expand description

The packed layouts that cross the Rust/GPU boundary, defined once for every backend.

§One definition, four languages

Rust here, MSL in crates/metal/src/shaders/bn254_fr.metal, CUDA C in crates/gpu-kernels/src/kernels/bn254_fr.cuh, and the WGSL that crates/wgpu/src/gen emits. Every #[repr(C)] type below has a byte-for-byte twin on each of those sides. Metal and CUDA retype the constants as kernel text, so each kernel side carries a static_assert on sizeof and each of those two backends keeps a test that greps its own source for the exact constant lines; this side carries a const assertion on size_of. snarkrs-wgpu reads the constants out of this crate at run time and generates its WGSL from them, so it has no copy that can drift, but it does depend on the struct strides and on the infinity sentinel. If you change a struct here you are changing a wire format three backends read; change the MSL and the CUDA header in the same commit.

This crate exists so that the three backends cannot disagree about what a point is. They run on different hardware with different compilers, and a benchmark that compares them is only meaningful if they are fed bit-identical inputs.

§Why repacking is not optional

Measured on this workspace’s arkworks version: size_of::<ark_ec::G1Affine>() is 72, not 64, because the affine type carries an infinity: bool plus padding, and size_of::<ark_ec::G2Affine>() is 136, not 128. Neither ark_ff::Fp nor the affine types are repr(C), so their field order is not even guaranteed. Byte-casting a Vec<G1Affine> into a device buffer hands the GPU a 72-byte stride while the kernel reads 64, and every point after the first is garbage. There is no way around an explicit repack, so this module makes the repack the only path.

§Montgomery form, and the asymmetry that silently breaks proofs

arkworks stores an Fp internally in Montgomery form with R = 2^256, which is exactly the radix both GPU CIOS routines use. So:

  • PackedFr and PackedFq hold Montgomery limbs, copied straight out of Fp::0.0 with no conversion on either side. This is what field kernels want: the NTT, the coset shift and H = A*B - C are all multiplications and additions, and Montgomery form is closed under both.
  • PackedScalar holds standard limbs, via into_bigint(). This is what an MSM wants, because Pippenger slices a scalar into window digits and a window digit of a Montgomery representative is a digit of a*R mod n, which is a different number.

Getting this backwards produces a proof that is wrong by a factor of R and fails verification with nothing else to go on, so the two types are deliberately distinct and neither converts into the other.

Re-exports§

pub use glv::PackedGlv;

Modules§

glv
The GLV twiddle table behind ptau prepare’s group FFT: the 36-byte decomposed-scalar layout and the host-side lattice work that fills it, shared by the Metal and CUDA FFT drivers so the two cannot disagree about what a twiddle is. The GLV twiddle table behind the group inverse FFT, decomposed once for every backend.
testrng
A deterministic PRNG for tests in the GPU backend crates.

Structs§

PackedFq
BN254 base field element, 32 bytes, Montgomery form, little-endian 32-bit limbs.
PackedFq2
Fq2 = Fq[u]/(u^2 + 1), 64 bytes, c0 then c1, each Montgomery.
PackedFr
BN254 scalar field element, 32 bytes, Montgomery form, little-endian 32-bit limbs.
PackedG1Affine
G1 affine point, exactly 64 bytes: x then y, no flag word.
PackedG2Affine
G2 affine point, exactly 128 bytes: x then y, each an PackedFq2. Infinity is the all-zero encoding, for the same reason as PackedG1Affine (BN254 G2 is y^2 = x^3 + 3/(9 + u), whose constant term is nonzero, so (0, 0) is off-curve).
PackedScalar
BN254 scalar as an integer in [0, r), 32 bytes, standard form, little-endian 32-bit limbs. For MSM window decomposition only. See the module docs.

Constants§

FQ_MODULUS
BN254 base field modulus q, little-endian 32-bit limbs.
FQ_N0
-q^{-1} mod 2^32. Equals 3834012553, the same value zkmopro and zkonduit ship for BN254 Fq, which is a cheap independent cross-check on the derivation.
FR_MODULUS
BN254 scalar field modulus r, little-endian 32-bit limbs.
FR_N0
-r^{-1} mod 2^32, the CIOS per-limb reduction multiplier for Fr.
LIMBS
Limbs per field element. 32-bit limbs, not 64.

Traits§

Packed
Marker for a type whose in-memory bytes are exactly what the GPU should see.

Functions§

as_bytes
Byte view of a packed slice, for a device upload.