1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
//! SM4 SIMD backends.
//!
//! v0.5 W4 phase 2 landed the AVX2 8-way packed bitsliced S-box
//! [`sbox_x8::sbox_x8`] — 8 bytes packed into the low lanes of
//! `__m256i` (7 of 8 lanes wasted in the phase-2 `tau` consumer).
//!
//! v0.6 W6 (phase 3) added:
//! - [`sbox_x32::sbox_x32`] — AVX2 32-byte full-width packed S-box,
//! the throughput-favorable shape for an 8-block CBC-decrypt
//! batch (8 SM4 blocks × 4 `tau` bytes per round = 32 bytes).
//! - [`sbox_x16::sbox_x16`] — NEON 16-byte packed S-box on
//! `aarch64` (4 SM4 blocks × 4 `tau` bytes per round = 16 bytes).
//! Compile-time baseline; no runtime CPU detect.
//!
//! Issue #163 added:
//! - [`sbox_x4::sbox_x4`] — four-byte serial-`tau` adapter. `AArch64`
//! reuses one NEON x16 invocation; `x86_64` production is four
//! scalar calls (AVX2 not selected: 10% rule unmeasured);
//! [`sbox_x8`] is kept as an internal candidate / test surface.
//!
//! The scalar primitives (Boyar-Peralta Itoh-Tsujii gate sequence)
//! live in `scalar` and serve as the fallback path for every SIMD
//! entry point on targets without the relevant intrinsics. The
//! AVX2 byte-parallel primitives live in `avx2` and are shared
//! between [`sbox_x8`] (low-lanes staged) and [`sbox_x32`]
//! (full-width). The NEON byte-parallel primitives live in `neon`.
pub
pub
pub