cera
Rust-native LLM inference engine. Load a GGUF, generate text, make it fast.
Note: In version 0.6.0, Cera will introduce breaking API changes to simplify usage and consolidate several APIs across the engine and language bindings. Follow updates in Releases.
See the project README for benchmarks and design notes.
cera is the core library: GGUF loading, a quantized CPU kernel stack
(AVX2/AVX-512, NEON dotprod/i8mm) with optional wgpu GPU and BLAS backends, a
stateful session API with prefix caching, and a streaming token sink. It powers
the cera-cli CLI, the
cera-ffi mobile
bindings, and cera-wasm.
Install
[]
= "0.5"
Highlights in 0.5.0
- FreeToken: Semantic Anchor Caching (arXiv:2406.14588): Two-tier prefix caching (
cera::kv_cache::KvPrefixCache) with semantic anchor points, TurboQuant cold storage compression, and FlatBuffers v2 disk persistence. - DSpark: Neural Speculative Decoding (arXiv:2407.08608): Neural speculative drafting via lightweight sidecars (
cera::spec::dspark), parallel multi-token GPU verification on Metal and WebGPU, and batched LM-head verification. - TurboQuant KV-Cache Compression (arXiv:2504.19874): Pure-Rust PolarQuant + QJL compression achieving ~12x KV-cache memory reduction across CPU, Metal, and WebGPU backends.
- Pure-Rust Silero VAD v5 (
cera::vad): Native ONNX-free voice activity detection engine (SileroVad,VadIterator,VadConfig,VadSampleRate) operating on 512-sample streaming audio frames with automatic speech segment timestamping. - Hugging Face Model Repositories & Streaming Quantization (
cera::bundle::hf,cera::convert): Direct download and loading of Hugging Face repositories, with on-the-fly zero-disk streaming quantization of remote SafeTensors models directly to GGUF in memory. - WebGPU Depthformer Acceleration & Voice Modes: High-performance compute shaders for Depthformer audio decoder, unified web runtime, and 4 dedicated voice interaction modes.
- Multimodal Vision ViT Optimization: High-resolution image encoding improvements and async WebGPU readbacks.
Breaking changes in 0.4.0
0.4.0 adds public fields and enum variants to public types, so it is a minor
(not patch) release; a cargo update from 0.3.x will not pull it in
automatically. No type in cera is #[non_exhaustive], so these break any code
that writes exhaustive struct literals or exhaustive matches. Code that keeps
the default settings sees no behavior change; the one exception is the new
KvCompressionConflict below, which turns a previously-silent mismatch into an
error.
GenerateOptsgainedspec: Option<SpecDecode>: opt-in greedy speculative decoding (see Speculative decoding). Code that constructsGenerateOptswith an exhaustive struct literal must add the field; prefer functional-update syntax,GenerateOpts { max_tokens: 256, ..Default::default() }, which stays source-compatible across field additions. It defaults toNone, preserving prior behavior.CeraErrorgainedKvCompressionConflict: returned when a second session asks a model for a different KV-compression mode than the one it was built with, instead of handing back a cache the kernels do not match. Exhaustivematches overCeraErrorneed a new arm. This is the one item here that changes runtime behavior: that call used to succeed silently and corrupt the prefix cache.CeraErroralso gainedLoraUnsupportedByBackend: an adapter that is well-formed and dimensionally valid, but adapts a target the active backend has no hook for. Today that is exactly the routed mixture-of-experts targets (the router and the per-expert projections) on the two GPU backends, which the CPU backend applies. Another new arm for exhaustivematches. Note the FFI mirrorFfiErrorappends its counterpart at the end of the enum: UniFFI serializes by declaration ordinal, so a mid-enum insertion would renumber every later variant for a prebuilt binding.ModelConfiggainedmoe: Option<MoeConfig>,Nonefor every dense architecture, alongside the new publicMoeConfigtype. Struct literals ofModelConfigneed the field;..Default::default()is not available here, so this is a compile error rather than a silent behavior change.LoraTargetgained four routed-FFN variants (FfnGateInpfor the router, andFfnGateExps/FfnUpExps/FfnDownExpsfor the per-expert projections). Exhaustivematches need the arms, andLoraTarget::ALLis now[LoraTarget; 13]rather than[LoraTarget; 9]: prefer the publicLORA_TARGET_COUNTover a hard-coded length in any array typed by it.- The f16 KV cache (added for decode-at-depth) widened four public items in
the
cera::kv_cachemodule:KvCompressiongainedF16,LayerSnapshotgainedAttentionF16,LayerState::Attentiongainedkey_cache_f16andvalue_cache_f16, andInferenceStategainedkv_f16: bool. Exhaustivematches and literals over any of them need updating; matchingLayerState::Attention { .. }with a rest pattern is unaffected. - Behind the non-default
gpufeature,backend::wgpu::io_stats::GpuIoStatsgained apasses: u64counter, andGpuIoStats::per_tokenreturns(f64, f64, f64, f64)instead of(f64, f64, f64). That one is a signature change rather than an exhaustiveness break, so callers must destructure the extra element. These are debug counters; nothing in the inference path uses them.
Also new (non-breaking): SpecDecode and the cera::spec module; the Model
trait gained a defaulted truncate_kv method, so a backend can override the
speculative-decoding KV rewind while existing implementors inherit the prior
behavior unchanged; TurboQuant KV-cache compression now runs on the wgpu and
native Metal backends, not just CPU. The rest of the release is CPU and GPU
performance and correctness work, including a batched LM-head projection that
amortizes the output matrix across all verified positions in speculative decode,
and a Q4_1 decode path that now matches the batched prefill GEMM bit-for-bit. See
the benchmarks.
Changes in 0.3.1
A patch release: CPU and GPU performance work, Q5_K/Q4_1 quantization support,
and a native wgpu flash-attention decode path. No changes to CeraEngine,
Session, GenerateOpts, or any other type in the public prelude.
One caveat for anyone reaching into backend internals: the dead WGSL kernels
backend::wgpu::shaders::{GEMM_Q4_0, GEMM_Q8_0, ATTENTION} were removed once
the register-tiled GEMM and flash-attention kernels superseded them. They were
shader source text behind the gpu feature, never part of the intended
API, so this ships as a patch rather than a minor bump, but a ^0.3 consumer
that named them will need to stop.
Breaking changes in 0.3.0
0.3.0 adds public fields to two public structs, so it is a minor (not patch)
release; a cargo update from 0.2.x will not pull it in automatically.
GenerateOptsgainedignore_eos: bool(run decode to exactlymax_tokens, ignoring EOS/stop tokens, thellama.cpp --ignore-eosanalog). Code that constructsGenerateOptswith an exhaustive struct literal must add the field; prefer functional-update syntax,GenerateOpts { max_tokens: 256, ..Default::default() }, which stays source-compatible across field additions. It defaults tofalse, preserving prior behavior.ModelMetadatagainedadd_eos_token: bool(mirrors GGUFtokenizer.ggml.add_eos_token, alongside the existingadd_bos_token). This is an engine output type, so it only affects code that exhaustively pattern-matches or constructs it.
Also new (non-breaking): BpeTokenizer::encode_special (and the FFI
encode_text_special / wasm encodeSpecial wrappers) apply BOS/EOS to match
llama.cpp's llama_tokenize, and GenerateSummary::prompt_eval_ms now
reports real prefill wall time paired with prompt_eval_tokens.
Supported models
cera loads GGUF weights, either a raw .gguf file or a
LeapBundles manifest that points
at one. Dispatch is on the GGUF general.architecture string:
| Architecture | Examples |
|---|---|
lfm2 |
Liquid LFM2 / LFM2.5 (the canonical LeapBundles family) |
lfm2moe |
Liquid LFM2.5-8B-A1B (routed mixture-of-experts) |
qwen2, qwen3 |
Qwen2 / Qwen2.5 / Qwen3 |
llama |
LLaMA 2/3, and classic Mistral 7B (ships as GGUF arch llama) |
granite |
IBM Granite 3.x, and the dense Granite 4.1 line (3b / 8b / 30b) |
Any other architecture errors out with unsupported architecture: <name> (this
includes the newer mistral3/mistral4 layouts). No Granite 4.0 model loads
today: the 4.0-H hybrids convert to the separate arch granitehybrid; the
non-hybrid ones (granite-4.0-micro, -1b, -350m) do convert to granite,
but write attention.head_count_kv as a per-layer array the loader does not yet
accept. Granite 4.1 is unaffected.
Modalities: text-to-text is fully supported for every architecture above.
LFM2-Audio (lfm2-audio-v1, text+audio in/out) also loads. Vision (VL,
image-to-text) is wired up end-to-end: CeraEngine auto-attaches the vision
mmproj encoder for VL bundles, and Session::append_image (or
append_chat_with_images) runs image → ViT → projector → soft-token prefill.
Verified against LFM2.5-VL-450M. The ViT encode runs on the GPU (native Metal or
wgpu, selected by BackendPreference) with a CPU fallback.
Quick start
Load a local GGUF and stream tokens to stdout as they decode:
use ;
use BpeTokenizer;
/// A `ModalitySink` receives decoded tokens as generation streams. Only
/// `on_done` is required; `on_text_tokens` defaults to a no-op.
Session keeps the KV cache alive across append_text / generate calls, so a
chat loop reuses the prefix cache instead of re-prefilling each turn. Render a
model's chat template with cera::tokenizer::apply_chat_template.
Auto-downloading LeapBundles
With the remote feature, load a model straight from
huggingface.co/LiquidAI/LeapBundles
by id and quant (cached locally, SHA-256 verified):
use ;
use BundleRepo;
// `BundleRepo` caches downloaded manifests + model files under this directory.
let cfg = EngineConfig ;
let engine = from_bundle_id?;
Sampling
GenerateOpts exposes the usual knobs: temperature, top_p, top_k,
min_p, repetition_penalty, plus stop_tokens and an optional GBNF
grammar for constrained / JSON-shaped output. temperature <= 0 (or
top_k == 1) selects deterministic greedy decoding; otherwise sampling is
stochastic. Min-p and repetition penalty apply on the stochastic path only.
Speculative decoding
GenerateOpts::spec opts into greedy speculative decoding with prompt-lookup
(n-gram) drafting: no draft model, so no extra weight memory. The drafter
guesses the next tokens from the most recent earlier occurrence of the last
ngram tokens, and the target verifies up to k of them in a single forward.
A target forward reads every weight once, so verifying K drafted tokens in one
pass amortizes that read over the accepted run, which is why this targets the
memory-bandwidth wall in CPU decode-at-depth. As of #327 the verify path projects
all 1 + k positions' logits in a single batched GEMM, so the LM head (the
largest matrix in the model) is read once per round rather than once per position;
a per-row fallback remains for head dtypes without a batched kernel. The 1.49x
measured on a repetitive prompt predates that change and did not include its gain.
use SpecDecode;
let opts = GenerateOpts ;
Every emitted token is the target's own argmax, a valid greedy decode, so
a poor draft lowers the acceptance rate without affecting correctness. It is not
guaranteed bit-identical to a sequential greedy run: the verifier forwards a
batch where a sequential loop forwards one token at a time, and the two
reduction orders can pick opposite sides of a near-tie. It engages only on the
plain greedy path (temperature <= 0 or top_k == 1, no grammar), with a model
that reports
supports_all_logits() and an uncompressed (f32/f16) KV cache. In practice
that means the CPU dense (llama-family) path only: LlamaModel is the
one implementor, and the trait default is false, so LFM2 and every GPU model
fall through. Any other configuration falls back to normal decode transparently
rather than erroring, so setting spec unconditionally is safe; it is a
no-op where unsupported.
The CLI exposes it on bench (--spec, --spec-ngram, --spec-k) for
measuring the win; it is not wired into run or chat.
Tool calling
cera::tools renders tool schemas into the chat template and parses tool calls
back out, format-aware: ToolFormat::detect(arch) picks Pythonic (LFM2) vs
Hermes JSON (Qwen2.5/Qwen3) from the GGUF architecture.
Continuing from the Quick start (which sets up engine, session, and the
chat messages, and produces the decoded reply_text), the schema below uses
the serde_json crate, which cera does not re-export, so add it to your
Cargo.toml:
use Arc;
use Grammar;
use ;
use apply_chat_template_with_tools;
let tools = vec!;
let format = detect
.unwrap_or;
// Render tools into the prompt.
let prompt = apply_chat_template_with_tools?;
session.append_text?;
// Optional: constrain to a valid call via grammar + lazy start-marker trigger.
let mut opts = default;
if let Some = engine.tokenizer.special_token_id
// After generating, parse the reply. `ToolCall { name, arguments }`.
let calls = parse_tool_calls?; // empty vec == answered in prose
The constrained path guarantees a well-formed call (valid function name, valid argument names, correctly-typed values via JSON-Schema → GBNF); without it the model decides freely whether and how to call a tool.
LoRA adapters & hidden states
Load a LoRA adapter, a llama.cpp GGUF (from convert_lora_to_gguf) or a PEFT
.safetensors, and attach it to a Session. The delta is applied at inference
time (y += scale·B·(A·x)), never merged into the weights, so the base model
stays quantized and adapters hot-swap / unload per request. Runs on CPU, Metal,
and wgpu (batched-GEMM prefill + decode) and is dimension-checked at attach. The
one gap is lfm2moe's routed-FFN targets (the router and the per-expert
projections), which apply on CPU only; on a GPU backend such an adapter is
refused with CeraError::LoraUnsupportedByBackend rather than half-applied.
use LoraAdapterWeights;
let adapters = from_safetensors?; // or ::from_gguf(path)
session.attach_lora_adapters?; // hot-swap-able; applies to every forward
// ... generate / extract hidden states with the adapter active ...
session.remove_lora_adapters;
Pull the per-token last-layer hidden state (post-final-RMSNorm, the llama.cpp
--pooling none vector) straight out of the engine, reflecting the active
adapter. This is the classifier / embedding path (e.g. a section router: LFM2.5
- a
route_sectionLoRA + a small linear head over the mean-pooled state):
let hs = session.hidden_states_for_tokens?; // [T * hidden_size], row-major
let pooled = session.hidden_states_mean_pooled?; // [hidden_size]
Both are also exposed over the FFI (LoraAdapters / attachLora /
hiddenStatesMeanPooled) and WASM bindings.
Feature flags
Default-on features keep desktop/CLI builds full-featured; turn them off to
shrink the crate for wasm32-unknown-unknown or embedded targets
(--no-default-features).
| Feature | Default | What it adds |
|---|---|---|
parallel |
✅ | Multi-threaded CPU kernels (persistent affinity-pinned threadpool on native; rayon on wasm) |
std-fs |
✅ | Filesystem access (paths, caches) |
mmap |
✅ | Memory-mapped GGUF loading (⇒ std-fs) |
disk-cache |
✅ | Cold KV-cache tier on disk (⇒ std-fs) |
vl-preprocess |
✅ | Image input decode/resize for VL models |
avx512 |
✅ | x86-64 AVX-512 Q8_0/Q4_0 tier (needs Rust 1.89+) |
gpu |
- | wgpu compute backend |
metal |
- | Apple Metal backend (⇒ mmap) |
blas |
- | Opt-in GEMM accelerator |
remote |
- | BundleRepo HTTP download + SHA-256 (⇒ std-fs) |
MSRV: Rust 1.94 (edition 2024; the NEON f16 vcvt_f32_f16 KV-cache widen needs
1.94). The default-on avx512 feature enables the AVX-512 tier on x86;
disabling it caps that tier at AVX2.
CPU threading & tuning
On native targets the CPU backend dispatches GEMV/GEMM rows through a
persistent, affinity-pinned worker pool (not a per-call fork-join). Rows are
handed out by dynamic chunk-stealing, so a faster core simply claims more
chunks. On a part whose cores differ in speed each worker's chunk is also
sized to the cpu_capacity of the core it is pinned to, so that every
worker's chunk costs roughly the same wall-clock time and one slow core cannot
hold the rest at the dispatch barrier for a multiple of what its chunk cost.
Both mechanisms are inert on homogeneous hosts, and CERA_PIN=0 turns the
sizing off along with the pinning it reads placement from.
On heterogeneous big.LITTLE parts (Linux/Android) detection separates the
performance cores from the efficiency ones and sizes both pools to the former,
which fixes the multi-core decode collapse there. The efficiency cores are
still recorded, so a deliberately widened pool can give every worker its own
core rather than falling off a cliff, but nothing is that wide by default.
Capacity-sized chunks make such a widened pool considerably less costly
(measured on a Tensor G5, CERA_THREADS=8 prefill 116.5 to 142.0 tok/s) but
not free, so it remains an override rather than the default.
Elsewhere, desktop/server (where sysfs detection is skipped and every
logical CPU counts as a "perf core") and macOS (where the P-core count comes
from hw.perflevel0), prefill
uses all of them while decode is sized from the loaded model (see "How the
decode thread count is chosen" in the top-level README): small models that
spread a token across many small pool dispatches run narrow, large ones that
move more bytes per dispatch run wide. Where that sizing does not apply,
heterogeneous parts, or a host whose physical core count cannot be detected
(Windows, BSD, Intel macOS), the flat cap applies instead: the detected
perf-core count, capped at 12. Both pools are process-wide singletons, so the
decode width is sized from the first model loaded into a process and stays
there for any loaded after it; it does not re-size per load. Everything else is
auto-detected per device; the
environment variables below only override for tuning (CERA_THREADS moves the
detected performance-core count, which both RowPools and rayon's global pool
size from):
| Variable | Default | Effect |
|---|---|---|
CERA_DECODE_THREADS |
auto |
Decode worker count. A fixed <n> pins the width and overrides the automatic sizing (clamped to the detected performance cores); auto selects the model-based sizing below. |
CERA_DECODE_SIZING |
on | 0 / false / off disables model-aware decode sizing, falling back to the flat cap (detected perf cores, capped at 12). |
CERA_DECODE_NARROW |
physical / 2, capped at 12 |
Decode width for barrier-bound models (below the bytes-per-dispatch threshold); never exceeds the wide arm. Setting it also forces sizing on where it would otherwise be declined (on a host whose physical core count is undetectable, both arms must be pinned). |
CERA_DECODE_WIDE |
physical + physical / 4, capped at 24 |
Decode width for bandwidth-bound models, clamped to the detected cores. Setting it also forces sizing on where it would otherwise be declined (on a host whose physical core count is undetectable, both arms must be pinned). |
CERA_DECODE_BPD_KB |
2500 | Bytes-per-dispatch threshold (decimal KB) separating the two arms above. Unlike the two widths, this does not force sizing on where it is declined; it moves the threshold, it does not pin a width. |
CERA_THREADS |
detected perf-core count | Override the detected performance-core count (moves the auto width for both pools). Clamped to the number of pinnable cores on hosts that have any, with a warning: past that the surplus workers run unpinned and contend with pinned ones that are spin-waiting, measured at 35x slower on a Tensor G4. Not clamped where nothing gets pinned anyway: hosts with no affinity, or CERA_PIN=0. |
CERA_PREFILL_THREADS |
detected perf-core count | Prefill pool width on its own, without moving decode. May reach past the performance cores up to every pinnable core, for sweeping a part where widening might pay. It does not on the parts measured so far, though it is no longer costly: the prefill pool stops pinning once it is wider than the performance cores, which took widening on a Tensor G5 from 211 to 148 tok/s down to 211 to about 200. Still a small loss, so the default stays narrow. Same no-clamp-without-pins rule as CERA_THREADS, including the CERA_PIN=0 case. Note the two interact: CERA_THREADS truncates the pinnable-core list, so setting it as well lowers the ceiling this is clamped against. To sweep prefill past the perf cores, leave CERA_THREADS unset. |
CERA_MIN_ROWS |
128 | Minimum output rows a decode-GEMV worker takes before another joins. |
CERA_PAR_THRESHOLD |
256 | Minimum output dimension before a GEMV parallelizes; smaller GEMVs stay serial. |
CERA_SPIN |
100000 | Spin iterations before an idle worker parks. |
CERA_PIN |
on | 0 / false / off disables affinity pinning (for hosts that manage thread placement themselves). |
RAYON_NUM_THREADS |
detected perf-core count (moved by CERA_THREADS) |
Width of rayon's global pool, which covers the parallel sites outside the RowPools: dequantization (so, model load), the ViT patch embed, and the LFM2-Audio conv stem. It does not move text prefill or decode width; every GEMM and GEMV on that path runs on a RowPool, sized by CERA_PREFILL_THREADS for prefill and CERA_DECODE_THREADS for decode (both defaulting from CERA_THREADS). It can still move a VL or audio prefill, whose encoders fan out on rayon. Read by cera itself rather than left to rayon, so the pool is built eagerly with a known CPU mask instead of lazily inheriting the mask of whichever thread reached it first. |
CERA_RAYON_GLOBAL |
on | 0 / false / off stops cera claiming rayon's process-global pool, for a Rust host that wants to build it itself. Such a host can also just call rayon::ThreadPoolBuilder::new().build_global() before loading a model; cera then logs a warning and leaves it alone. |
CERA_CPU_TIER |
auto | Force a lower CPU SIMD tier (downgrade only), for parity testing on capable hardware. |
CERA_POOL_STATS |
off | 1 annotates each cera bench run with the pool's fan-out health: how many dispatches wanted more than one worker, and how many of those silently ran serially because the pool was already busy. The counts are exact; the accompanying work percentage mixes units across dispatch kinds, so read the counts. |
CERA_LM_HEAD_NO_GEMM |
unset | 1 puts the LM-head projection in forward_prefill_logits_all back on the per-row loop the batched GEMM replaced, so both halves of a speculative-decoding A/B run from one binary. Measurement lever only; both paths compute the same projection, to within f32 accumulation order. |
Affinity pinning applies on Linux/Android with a detected heterogeneous topology; homogeneous hosts and macOS run unpinned.
None of this needs configuring from the host. Loading a model builds rayon's
global pool; the RowPools build themselves on first use (prefill on the first
GEMM, decode on the first decode GEMV, which is what lets decode size itself
from the loaded model). Two functions let a host move that work earlier if it
wants to: cera::backend::cpu::ensure_rayon_global_pool() builds just the
rayon pool, and configure_thread_pool() also warms the prefill RowPool and
returns its width, so a CLI can report a thread count before a model exists.
Both are optional and idempotent.
License
Apache-2.0 OR MIT.