Expand description
Where some bytes are in a block of sixty four, one bit a byte.
This is the step a structural scan starts from, the one simdjson and simdcsv take: compare a
block against a few bytes of interest and get a u64 per byte, first byte in the lowest bit.
The compares vectorize by themselves. Turning sixteen compare results into sixteen bits does
not, because the portable way to gather them is a multiply per eight bytes and the compiler does
not know that pmovmskb does it in one instruction. On a lineitem load from CSV the gathering
was more than half of the splitter’s time.
So on x86_64 the gather is SSE2’s movemask, which every x86_64 processor has, and there is
nothing to detect at run time. Everywhere else it is the portable multiply, which is also what
the tests hold the SSE2 version to.
Functions§
- masks
- One mask per byte in
needles, with bitiof masknset whenblock[i] == needles[n].