macro_rules! gemv_kernel {
(
4,
$is_fused:expr,
$out_c:expr,
$out_len:expr,
$in_frame:expr,
$weights:expr,
$bias:expr,
$out_frame:expr,
$do_bias:expr,
$setzero:expr,
$load_out:expr,
$load_bias:expr,
$add_ps:expr,
$load_weight:expr,
$fmadd_ps:expr,
$store_ps:expr
) => { ... };
(
8,
$is_fused:expr,
$out_c:expr,
$out_len:expr,
$in_frame:expr,
$weights:expr,
$bias:expr,
$out_frame:expr,
$do_bias:expr,
$setzero:expr,
$load_out:expr,
$load_bias:expr,
$add_ps:expr,
$load_weight:expr,
$fmadd_ps:expr,
$store_ps:expr
) => { ... };
}Expand description
GEMV kernel macro — generates platform-specific FMADD accumulate loops.
Parameterized by SIMD width (4 = AVX2/256-bit, 8 = AVX-512/512-bit) and all relevant SIMD operations as inline closures. Both variants use 8 independent accumulators to maximize FMA throughput.
§Safety
Caller must ensure valid pointer arithmetic and slice bounds.