Skip to main content

Crate ftui_simd

Crate ftui_simd 

Source
Expand description

Portable-SIMD kernels for FrankenTUI’s two lane-parallel hot loops.

§Role in FrankenTUI

Two places in the render path do the same narrow thing over a long run of bytes: the diff compares one row of 16-byte cells against the previous frame’s row, and the width fast path asks whether a string is entirely ASCII. Both are a wide equality test followed by “where was the first difference”, which is what SIMD lanes are for.

§How it fits in the system

Every kernel here is safe std::simd and carries a _scalar twin with identical semantics. The twins are not dead weight: they are what the parity tests compare against, what the benches measure against, and what callers use when the simd feature is off. Nothing in the workspace depends on this crate unless that feature is enabled, so the scalar paths stay the default until benches justify otherwise.

§Why u128 and not Cell

ftui_render::Cell is #[repr(C, align(16))] over four u32 fields, so a cell is bit-for-bit a u128. Keeping the kernels on u128 leaves this crate free of a dependency on the render crate, and leaves the conversion (which must stay safe, so it composes the fields rather than transmuting) on the caller’s side.

§Lane widths

The kernels use 512-bit vectors (u64x8, u8x64). That is wider than most targets execute natively; std::simd splits them into whatever the target has, which keeps one code path across x86-64 and aarch64 and gives the compiler a full unrolled chunk to work with.

§Which of these are worth calling

Measured 2026-09-18, full numbers in docs/perf/simd_kernels_2026-09-18.md:

  • all_ascii and ascii_width beat their twins by 5x at 64 bytes and by 20-44x from a kilobyte up. Call these.
  • first_mismatch_u128 and rows_equal_u128 are 3-4x slower than their twins and are deliberately not wired into the diff. #![forbid(unsafe_code)] leaves no way to view &[u128] as lanes, so the kernel must build each vector with shifts and masks, while the scalar a[i] != b[i] over u128 is already one 128-bit compare after LLVM is done with it. They are kept because the measurement is worth keeping, and because the parity tests over them are what prove the lane indexing is right.

Prefer the _scalar twin for cell comparison. That is not a placeholder.

Functions§

all_ascii
Whether every byte is ASCII, i.e. has its high bit clear.
all_ascii_scalar
Scalar twin of all_ascii.
ascii_width
Display width of a run that is entirely printable ASCII, or None.
ascii_width_scalar
Scalar twin of ascii_width.
first_mismatch_u128
Index of the first element where a and b differ, or None when the common prefix runs to the end of the shorter slice.
first_mismatch_u128_scalar
Scalar twin of first_mismatch_u128, and the reference its parity tests compare against.
rows_equal_u128
Whether two equal-length runs of cells are bitwise identical.
rows_equal_u128_scalar
Scalar twin of rows_equal_u128.