Expand description
Portable-SIMD kernels for FrankenTUI’s two lane-parallel hot loops.
§Role in FrankenTUI
Two places in the render path do the same narrow thing over a long run of bytes: the diff compares one row of 16-byte cells against the previous frame’s row, and the width fast path asks whether a string is entirely ASCII. Both are a wide equality test followed by “where was the first difference”, which is what SIMD lanes are for.
§How it fits in the system
Every kernel here is safe std::simd and carries a _scalar twin with
identical semantics. The twins are not dead weight: they are what the
parity tests compare against, what the benches measure against, and what
callers use when the simd feature is off. Nothing in the workspace
depends on this crate unless that feature is enabled, so the scalar paths
stay the default until benches justify otherwise.
§Why u128 and not Cell
ftui_render::Cell is #[repr(C, align(16))] over four u32 fields, so a
cell is bit-for-bit a u128. Keeping the kernels on u128 leaves this
crate free of a dependency on the render crate, and leaves the conversion
(which must stay safe, so it composes the fields rather than transmuting)
on the caller’s side.
§Lane widths
The kernels use 512-bit vectors (u64x8, u8x64). That is wider than most
targets execute natively; std::simd splits them into whatever the target
has, which keeps one code path across x86-64 and aarch64 and gives the
compiler a full unrolled chunk to work with.
§Which of these are worth calling
Measured 2026-09-18, full numbers in
docs/perf/simd_kernels_2026-09-18.md:
all_asciiandascii_widthbeat their twins by 5x at 64 bytes and by 20-44x from a kilobyte up. Call these.first_mismatch_u128androws_equal_u128are 3-4x slower than their twins and are deliberately not wired into the diff.#![forbid(unsafe_code)]leaves no way to view&[u128]as lanes, so the kernel must build each vector with shifts and masks, while the scalara[i] != b[i]overu128is already one 128-bit compare after LLVM is done with it. They are kept because the measurement is worth keeping, and because the parity tests over them are what prove the lane indexing is right.
Prefer the _scalar twin for cell comparison. That is not a placeholder.
Functions§
- all_
ascii - Whether every byte is ASCII, i.e. has its high bit clear.
- all_
ascii_ scalar - Scalar twin of
all_ascii. - ascii_
width - Display width of a run that is entirely printable ASCII, or
None. - ascii_
width_ scalar - Scalar twin of
ascii_width. - first_
mismatch_ u128 - Index of the first element where
aandbdiffer, orNonewhen the common prefix runs to the end of the shorter slice. - first_
mismatch_ u128_ scalar - Scalar twin of
first_mismatch_u128, and the reference its parity tests compare against. - rows_
equal_ u128 - Whether two equal-length runs of cells are bitwise identical.
- rows_
equal_ u128_ scalar - Scalar twin of
rows_equal_u128.