et-k-rs
Device-side library for writing ET-SoC-1 compute kernels in pure no_std
Rust — the device counterpart to the et_soc1 host crate, with no C
dependency. The library (et_kernel, src/lib.rs) provides hart identity, the
U-mode trace write, a hardware fence, scratchpad addressing, and the safe
Grid partitioning abstraction. Launch-argument structs are shared with the host
launcher through the et-abi crate, so the two sides cannot drift
on layout. Three demo kernels build on the library and double as worked examples.
Kernels
hello-rs(src/bin/hello.rs) -- every hart writes"Hello World from hart N"to its trace buffer. A drop-in Rust replacement for the SDK's Chello.c; reimplementsget_hart_id(thehartidCSR0xCD0) and theTrace_Stringwrite directly.spsc-rs(src/bin/spsc.rs) -- a single-producer/single-consumer, lock-free, non-atomic queue across two harts (plain volatile loads/stores +fence rw,rw, no atomics, no locks). A coherence probe: it showed the ET-SoC-1 is software-coherent -- fence-only cross-hart sharing does not propagate (even within one minion), so this needs explicit cache management or genuinely shared memory. See the crate root README.reduce-rs(src/bin/reduce.rs) -- a data-parallel reduction (sum) over a DRAM array across a shire's 64 harts. Each hart reduces its disjoint slice (Grid::my_slice) and writes its own cache-line-padded partial cell (no false sharing); the host combines. No cross-hart sharing during the kernel, so it is coherence-clean and validated on hardware.sgemm-rs(src/bin/sgemm.rs) -- single-precision GEMM (C = A*B) using the ET-SoC-1 tensor extension. Tile assignment is shire-blocked: each shire handles a contiguous slice of the tile grid, improving A-row reuse in the shire-shared L2 (+26% at N=4096 over global-cyclic). v0.1 supports alpha=1.0 and beta=0.0; N may be any positive integer (partial last-column tile handled via stride-aligned padding). Verified on hardware: 64x64x64 and 32x20x32 (partial-N) both produce correct results with zero floating-point error. See the host-sidesgemmandsgemm_partialexamples inet-rs/.
Tensor extension (et_kernel::tensor)
All tensor operations on the ET-SoC-1 are encoded as standard RISC-V
csrrw xd, <csr>, xs writes (PRM Chapter 9). No custom opcode or target
feature is required; riscv64imac suffices. The tensor module exposes
typed, inline-asm wrappers for each instruction:
| Function | CSR | Role |
|---|---|---|
tensor_load |
0x83F |
Async load from DRAM into L1 scratchpad (ID=0 or ID=1). |
tensor_load_b |
0x83F |
Async load from DRAM into the TenB register file (bit 52 set). |
tensor_load_l2 |
0x85F |
Async prefetch from DRAM to shire L2 cache (no L1 fill). |
tensor_fma32 |
0x801 |
Async FMA32: C += A * B (or C = A * B when mul_only). |
tensor_store |
0x87F |
Async store from FP register file to DRAM. |
tensor_wait |
0x830 |
Stall hart until Load0, Load1, Fma, or Store event fires. |
tensor_error |
0x808 |
Read latched co-processor error flags (#[must_use]). |
check_tensor_error |
0x808 |
Returns Ok(()) or Err(TensorError) (typed wrapper). |
set_tensor_mask |
0x805 |
Write per-row FMA enable bits. |
fma32_xs constructs the xs bit field for tensor_fma32, encoding BCOLS,
AROWS, ACOLS, AOFFSET, TENB, BSTART, ASTART, MUL, and MSK from PRM Table 9-4.
x31 (t6) carries the row stride for tensor_load, tensor_load_b, and
tensor_store; each function sets it atomically with the CSRRW inside the
same asm block.
Usage pattern
use ;
use fence;
// Load A tile into L1 scratchpad (lines 0..arows); ID=false -> Load0.
unsafe
unsafe
// Load B tile into TenB register file; ID=true -> Load1, keeping B's
// event independent of the A Load0 above.
unsafe
// Issue FMA: C = A * B (first k-tile) or C += A * B (subsequent).
let xs = fma32_xs;
unsafe
// Optional: check for co-processor faults before storing.
check_tensor_error.expect;
// Store C from FP registers to DRAM, then drain the store and fence.
unsafe
unsafe
fence;
PMU counters (et_kernel::pmu)
U-mode-accessible hardware performance counters for characterising kernel
behaviour. Reads hpmcounterN (CSR 0xC03 + (N-3)) without any privilege
escalation.
RTLMIN-6496 workaround. All read functions emit four consecutive csrrs
instructions for the same CSR in a 16-byte-aligned block (.align 4) and
return the fourth value. This satisfies the hardware erratum requirement; a
single read returns an unreliable value on affected silicon. The timestamp()
function in the library root applies the same workaround.
use ;
// Read counter 4 before and after the k-loop; the delta is the number of
// TFMA_WAIT_TENB stall cycles (assuming firmware assigned PmuEvent::TfmaWaitTenb
// to counter 4 via mhpmevent4).
let before = pmu_read;
// ... tensor k-loop ...
let after = pmu_read;
let stalls = after.wrapping_sub;
Counter availability. Only hpmcounter3-hpmcounter8 accumulate events;
counters 9-31 are tied to zero by the hardware. mcycle (CSR 0xC00) and
minstret (CSR 0xC02) are also permanently zero -- use PmuEvent::Cycles
in hpmcounter3-hpmcounter6 and PmuEvent::RetiredInst0/1 instead
(PRM section 1.3.2).
| Item | Description |
|---|---|
pmu_read(counter: u8) -> u64 |
Read hpmcounterN for N in 3..=8; counters 9-31 return 0 (RTLMIN-6496 workaround applied). |
pmu_read_cycle() -> u64 |
Always returns 0 (mcycle is tied to zero on ET-SoC-1). |
pmu_read_instret() -> u64 |
Always returns 0 (minstret is tied to zero on ET-SoC-1). |
PmuEvent |
29 Minion-level event codes for mhpmevent3-mhpmevent6 (PRM section 1.3.2, Table 1-3). |
NeighborhoodEvent |
22 neighbourhood-level event codes for mhpmevent7-mhpmevent8 (PRM section 1.3.2, Table 1-4). |
PS SIMD (et_kernel::simd)
The simd module provides wrappers for the ET-SoC-1 packed-single (PS) SIMD
extension, gated on cfg(target_feature = "f"). Encodings are sourced from
esperanto-opc.h in the ET-SoC-1 binutils fork (present on aifoundry3 at
/home/rich/riscv-gnu-toolchain/gdb/include/opcode/esperanto-opc.h).
| Function | PS instruction | Notes |
|---|---|---|
broadcast_ps(scalar) -> f32 |
FBCX.PS f28, tmp |
Broadcasts scalar to all 8 lanes of f28 (scratch). fmv.x.w moves bit pattern to integer register first. |
scale_c_row(row, alpha) |
FBCX.PS + FMUL.PS |
Broadcasts alpha to f28, then element-wise multiplies f[2*row] and f[2*row+1] by f28. row must be 0..=13. |
Constraint. Both functions use f28 (ft8) as a broadcast scratch register.
Rows 14 and 15 place C-tile data in f28..=f31, conflicting with this scratch.
For a full 16-row GEMM tile (GEMM_TILE_M = 16) rows 14 and 15 cannot be
scaled with this API without an additional save slot; the current sgemm kernel
uses alpha = 1.0 and never calls scale_c_row.
The module remains #[doc(hidden)] pending hardware verification on
aifoundry3. Do not depend on it in production code.
The safety story (reduce-rs)
The kernel body is safe Rust over Grid: a hart can obtain only its own input
slice and its own output cell, so an out-of-partition access or a cross-hart data
race is unrepresentable. The only unsafe is a thin, commented boundary that
turns launch arguments and device addresses into typed slices.
Build
Cross-compiles to the compute harts (RV64IMAC); target, code model
(medium = medany, for the fixed high link address) and linker script are in
.cargo/config.toml:
# -> target/riscv64imac-unknown-none-elf/release/{hello-rs,spsc-rs,reduce-rs}
The library (--lib) also compiles on the host target for IDE type-checking and
rust-analyzer support. All items that contain RISC-V inline assembly
(hart_id, fence, timestamp, trace_str, kernel_entry!, Grid, and the
tensor/pmu/cache modules) are gated on #[cfg(target_arch = "riscv64")]
and are absent on non-RISC-V targets. MsgBuf, device_slice, scp_shire_base,
and CACHE_LINE remain available everywhere. The binary kernels (hello-rs etc.)
contain device asm and still require the RISC-V target.
When consuming et-k-rs as a registry dependency from a workspace whose
.cargo/config.toml sets [build] target = "riscv64imac-unknown-none-elf",
build from inside the kernel crate directory so Cargo finds the config:
# From workspace root via --manifest-path: Cargo may resolve to host target.
# Safer: cd into the crate first.
&&
Run
Load and launch with the host crate's examples (from the repository root), emulator or hardware:
&& &&
K=et-k-rs/target/riscv64imac-unknown-none-elf/release
Kernel facts (reference)
- Entry/exit:
_startsetsgp, callsentry_point, thenecallwithSYSCALL_RETURN_FROM_KERNEL(8) /KERNEL_RETURN_SUCCESS(0). Firmware sets the stack pointer. - Launch args: the launch command's
pointer_to_argsis delivered ina0(notra, despite the SDK docs — verified on device);a0flows through_startintoentry_point's first parameter. - No
.bss(the linker script asserts it), no heap, no unwinding (panic = "abort").
Publishing
This crate is a separate cargo package from the host crate (and excluded from the host workspace) because it targets RISC-V bare metal. Its library and demo bins use RISC-V inline assembly and a linker script, so they cannot be built for the host target; a crates.io release must therefore verify against the device target:
Thanks
Thanks to AiNEKKO https://nekko.ai/ and AI Foundry https://aifoundry.org/ for allowing me time on their community ET-SoC-1 servers to develop this code.
The ET-SoC-1 ET Platform SDK and software emulator can be found on their GitHub: https://github.com/aifoundry-org/et-platform
Licence
Apache-2.0, matching the ET Platform SDK headers this crate binds to.
ET-SoC-1 ET Platform API is under the Apache 2 License.