cera
Rust-native LLM inference engine. Load a GGUF, generate text, make it fast.
See the project README for benchmarks and design notes.
cera is the core library: GGUF loading, a quantized CPU kernel stack
(AVX2/AVX-512, NEON dotprod/i8mm) with optional wgpu GPU and BLAS backends, a
stateful session API with prefix caching, and a streaming token sink. It powers
the cera-cli CLI, the
cera-ffi mobile
bindings, and cera-wasm.
Install
[]
= "0.4"
Breaking changes in 0.4.0
0.4.0 adds public fields and enum variants to public types, so it is a minor
(not patch) release — a cargo update from 0.3.x will not pull it in
automatically. No type in cera is #[non_exhaustive], so these break any code
that writes exhaustive struct literals or exhaustive matches. Code that keeps
the default settings sees no behavior change; the one exception is the new
KvCompressionConflict below, which turns a previously-silent mismatch into an
error.
GenerateOptsgainedspec: Option<SpecDecode>— opt-in greedy speculative decoding (see Speculative decoding). Code that constructsGenerateOptswith an exhaustive struct literal must add the field; prefer functional-update syntax —GenerateOpts { max_tokens: 256, ..Default::default() }— which stays source-compatible across field additions. It defaults toNone, preserving prior behavior.CeraErrorgainedKvCompressionConflict— returned when a second session asks a model for a different KV-compression mode than the one it was built with, instead of handing back a cache the kernels do not match. Exhaustivematches overCeraErrorneed a new arm. This is the one item here that changes runtime behavior: that call used to succeed silently and corrupt the prefix cache.- The f16 KV cache (added for decode-at-depth) widened four public items in
the
cera::kv_cachemodule:KvCompressiongainedF16,LayerSnapshotgainedAttentionF16,LayerState::Attentiongainedkey_cache_f16andvalue_cache_f16, andInferenceStategainedkv_f16: bool. Exhaustivematches and literals over any of them need updating; matchingLayerState::Attention { .. }with a rest pattern is unaffected. - Behind the non-default
gpufeature,backend::wgpu::io_stats::GpuIoStatsgained apasses: u64counter, andGpuIoStats::per_tokenreturns(f64, f64, f64, f64)instead of(f64, f64, f64). That one is a signature change rather than an exhaustiveness break, so callers must destructure the extra element. These are debug counters; nothing in the inference path uses them.
Also new (non-breaking): SpecDecode and the cera::spec module; the Model
trait gained a defaulted truncate_kv method, so a backend can override the
speculative-decoding KV rewind while existing implementors inherit the prior
behavior unchanged; TurboQuant KV-cache compression now runs on the wgpu and
native Metal backends, not just CPU. The rest of the release is CPU and GPU
performance and correctness work, including a batched LM-head projection that
amortizes the output matrix across all verified positions in speculative decode,
and a Q4_1 decode path that now matches the batched prefill GEMM bit-for-bit. See
the benchmarks.
Changes in 0.3.1
A patch release: CPU and GPU performance work, Q5_K/Q4_1 quantization support,
and a native wgpu flash-attention decode path. No changes to CeraEngine,
Session, GenerateOpts, or any other type in the public prelude.
One caveat for anyone reaching into backend internals: the dead WGSL kernels
backend::wgpu::shaders::{GEMM_Q4_0, GEMM_Q8_0, ATTENTION} were removed once
the register-tiled GEMM and flash-attention kernels superseded them. They were
shader source text behind the gpu feature, never part of the intended
API, so this ships as a patch rather than a minor bump — but a ^0.3 consumer
that named them will need to stop.
Breaking changes in 0.3.0
0.3.0 adds public fields to two public structs, so it is a minor (not patch)
release — a cargo update from 0.2.x will not pull it in automatically.
GenerateOptsgainedignore_eos: bool(run decode to exactlymax_tokens, ignoring EOS/stop tokens — thellama.cpp --ignore-eosanalog). Code that constructsGenerateOptswith an exhaustive struct literal must add the field; prefer functional-update syntax —GenerateOpts { max_tokens: 256, ..Default::default() }— which stays source-compatible across field additions. It defaults tofalse, preserving prior behavior.ModelMetadatagainedadd_eos_token: bool(mirrors GGUFtokenizer.ggml.add_eos_token, alongside the existingadd_bos_token). This is an engine output type, so it only affects code that exhaustively pattern-matches or constructs it.
Also new (non-breaking): BpeTokenizer::encode_special (and the FFI
encode_text_special / wasm encodeSpecial wrappers) apply BOS/EOS to match
llama.cpp's llama_tokenize, and GenerateSummary::prompt_eval_ms now
reports real prefill wall time paired with prompt_eval_tokens.
Supported models
cera loads GGUF weights — either a raw .gguf file or a
LeapBundles manifest that points
at one. Dispatch is on the GGUF general.architecture string:
| Architecture | Examples |
|---|---|
lfm2 |
Liquid LFM2 / LFM2.5 (the canonical LeapBundles family) |
qwen2, qwen3 |
Qwen2 / Qwen2.5 / Qwen3 |
llama |
LLaMA 2/3, and classic Mistral 7B (ships as GGUF arch llama) |
granite |
IBM Granite 3.x |
Any other architecture errors out with unsupported architecture: <name> (this
includes the newer mistral3/mistral4 layouts).
Modalities: text-to-text is fully supported for every architecture above.
LFM2-Audio (lfm2-audio-v1, text+audio in/out) also loads. Vision (VL,
image-to-text) is wired up end-to-end: CeraEngine auto-attaches the vision
mmproj encoder for VL bundles, and Session::append_image (or
append_chat_with_images) runs image → ViT → projector → soft-token prefill.
Verified against LFM2.5-VL-450M. The ViT encode runs on the GPU (native Metal or
wgpu, selected by BackendPreference) with a CPU fallback.
Quick start
Load a local GGUF and stream tokens to stdout as they decode:
use ;
use BpeTokenizer;
/// A `ModalitySink` receives decoded tokens as generation streams. Only
/// `on_done` is required; `on_text_tokens` defaults to a no-op.
Session keeps the KV cache alive across append_text / generate calls, so a
chat loop reuses the prefix cache instead of re-prefilling each turn. Render a
model's chat template with cera::tokenizer::apply_chat_template.
Auto-downloading LeapBundles
With the remote feature, load a model straight from
huggingface.co/LiquidAI/LeapBundles
by id and quant (cached locally, SHA-256 verified):
use ;
use BundleRepo;
// `BundleRepo` caches downloaded manifests + model files under this directory.
let cfg = EngineConfig ;
let engine = from_bundle_id?;
Sampling
GenerateOpts exposes the usual knobs: temperature, top_p, top_k,
min_p, repetition_penalty, plus stop_tokens and an optional GBNF
grammar for constrained / JSON-shaped output. temperature <= 0 (or
top_k == 1) selects deterministic greedy decoding; otherwise sampling is
stochastic. Min-p and repetition penalty apply on the stochastic path only.
Speculative decoding
GenerateOpts::spec opts into greedy speculative decoding with prompt-lookup
(n-gram) drafting — no draft model, so no extra weight memory. The drafter
guesses the next tokens from the most recent earlier occurrence of the last
ngram tokens, and the target verifies up to k of them in a single forward.
A target forward reads every weight once, so verifying K drafted tokens in one
pass amortizes that read over the accepted run — which is why this targets the
memory-bandwidth wall in CPU decode-at-depth. As of #327 the verify path projects
all 1 + k positions' logits in a single batched GEMM, so the LM head (the
largest matrix in the model) is read once per round rather than once per position;
a per-row fallback remains for head dtypes without a batched kernel. The 1.49x
measured on a repetitive prompt predates that change and did not include its gain.
use SpecDecode;
let opts = GenerateOpts ;
Every emitted token is the target's own argmax — a valid greedy decode, so
a poor draft lowers the acceptance rate without affecting correctness. It is not
guaranteed bit-identical to a sequential greedy run: the verifier forwards a
batch where a sequential loop forwards one token at a time, and the two
reduction orders can pick opposite sides of a near-tie. It engages only on the
plain greedy path (temperature <= 0 or top_k == 1, no grammar), with a model
that reports
supports_all_logits() and an uncompressed (f32/f16) KV cache. In practice
that means the CPU dense (llama-family) path only: LlamaModel is the
one implementor, and the trait default is false, so LFM2 and every GPU model
fall through. Any other configuration falls back to normal decode transparently
rather than erroring, so setting spec unconditionally is safe — it is a
no-op where unsupported.
The CLI exposes it on bench (--spec, --spec-ngram, --spec-k) for
measuring the win; it is not wired into run or chat.
Tool calling
cera::tools renders tool schemas into the chat template and parses tool calls
back out, format-aware: ToolFormat::detect(arch) picks Pythonic (LFM2) vs
Hermes JSON (Qwen2.5/Qwen3) from the GGUF architecture.
Continuing from the Quick start (which sets up engine, session, and the
chat messages, and produces the decoded reply_text) — the schema below uses
the serde_json crate, which cera does not re-export, so add it to your
Cargo.toml:
use Arc;
use Grammar;
use ;
use apply_chat_template_with_tools;
let tools = vec!;
let format = detect
.unwrap_or;
// Render tools into the prompt.
let prompt = apply_chat_template_with_tools?;
session.append_text?;
// Optional: constrain to a valid call via grammar + lazy start-marker trigger.
let mut opts = default;
if let Some = engine.tokenizer.special_token_id
// After generating, parse the reply. `ToolCall { name, arguments }`.
let calls = parse_tool_calls?; // empty vec == answered in prose
The constrained path guarantees a well-formed call (valid function name, valid argument names, correctly-typed values via JSON-Schema → GBNF); without it the model decides freely whether and how to call a tool.
LoRA adapters & hidden states
Load a LoRA adapter — a llama.cpp GGUF (from convert_lora_to_gguf) or a PEFT
.safetensors — and attach it to a Session. The delta is applied at inference
time (y += scale·B·(A·x)), never merged into the weights, so the base model
stays quantized and adapters hot-swap / unload per request. Runs on CPU, Metal,
and wgpu (batched-GEMM prefill + decode) and is dimension-checked at attach.
use LoraAdapterWeights;
let adapters = from_safetensors?; // or ::from_gguf(path)
session.attach_lora_adapters?; // hot-swap-able; applies to every forward
// ... generate / extract hidden states with the adapter active ...
session.remove_lora_adapters;
Pull the per-token last-layer hidden state (post-final-RMSNorm — the llama.cpp
--pooling none vector) straight out of the engine, reflecting the active
adapter. This is the classifier / embedding path (e.g. a section router: LFM2.5
- a
route_sectionLoRA + a small linear head over the mean-pooled state):
let hs = session.hidden_states_for_tokens?; // [T * hidden_size], row-major
let pooled = session.hidden_states_mean_pooled?; // [hidden_size]
Both are also exposed over the FFI (LoraAdapters / attachLora /
hiddenStatesMeanPooled) and WASM bindings.
Feature flags
Default-on features keep desktop/CLI builds full-featured; turn them off to
shrink the crate for wasm32-unknown-unknown or embedded targets
(--no-default-features).
| Feature | Default | What it adds |
|---|---|---|
parallel |
✅ | Multi-threaded CPU kernels (persistent affinity-pinned threadpool on native; rayon on wasm) |
std-fs |
✅ | Filesystem access (paths, caches) |
mmap |
✅ | Memory-mapped GGUF loading (⇒ std-fs) |
disk-cache |
✅ | Cold KV-cache tier on disk (⇒ std-fs) |
vl-preprocess |
✅ | Image input decode/resize for VL models |
avx512 |
✅ | x86-64 AVX-512 Q8_0/Q4_0 tier (needs Rust 1.89+) |
gpu |
— | wgpu compute backend |
metal |
— | Apple Metal backend (⇒ mmap) |
blas |
— | Opt-in GEMM accelerator |
remote |
— | BundleRepo HTTP download + SHA-256 (⇒ std-fs) |
MSRV: Rust 1.94 (edition 2024; the NEON f16 vcvt_f32_f16 KV-cache widen needs
1.94). The default-on avx512 feature enables the AVX-512 tier on x86;
disabling it caps that tier at AVX2.
CPU threading & tuning
On native targets the CPU backend dispatches GEMV/GEMM rows through a
persistent, affinity-pinned worker pool (not a per-call fork-join), with dynamic
chunk-stealing so faster cores absorb more work on heterogeneous big.LITTLE
mobile. On heterogeneous big.LITTLE parts (Linux/Android) detection keeps at
most 6 big cores for both pools, which fixes the multi-core decode collapse
there. Elsewhere — desktop/server (where sysfs detection is skipped and every
logical CPU counts as a "perf core") and macOS (where the P-core count comes
from hw.perflevel0) — prefill
uses all of them while decode is sized from the loaded model (see "How the
decode thread count is chosen" in the top-level README): small models that
spread a token across many small pool dispatches run narrow, large ones that
move more bytes per dispatch run wide. Where that sizing does not apply —
heterogeneous parts, or a host whose physical core count cannot be detected
(Windows, BSD, Intel macOS) — the previous flat cap applies as before (≤12
homogeneous, ≤6 on big.LITTLE). Both pools are process-wide singletons, so the
decode width is sized from the first model loaded into a process and stays
there for any loaded after it — it does not re-size per load. Everything else is
auto-detected per device; the
environment variables below only override for tuning (CERA_THREADS moves the
detected count, which both pools size from):
| Variable | Default | Effect |
|---|---|---|
CERA_DECODE_THREADS |
auto |
Decode worker count. A fixed <n> pins the width and overrides the automatic sizing (clamped to the detected performance cores); auto selects the model-based sizing below. |
CERA_DECODE_SIZING |
on | 0 / false / off disables model-aware decode sizing, falling back to the flat cap (detected perf cores, ≤6 heterogeneous / ≤12 homogeneous). |
CERA_DECODE_NARROW |
physical / 2, capped at 12 |
Decode width for barrier-bound models (below the bytes-per-dispatch threshold); never exceeds the wide arm. Setting it also forces sizing on where it would otherwise be declined (on a host whose physical core count is undetectable, both arms must be pinned). |
CERA_DECODE_WIDE |
physical + physical / 4, capped at 24 |
Decode width for bandwidth-bound models, clamped to the detected cores. Setting it also forces sizing on where it would otherwise be declined (on a host whose physical core count is undetectable, both arms must be pinned). |
CERA_DECODE_BPD_KB |
2500 | Bytes-per-dispatch threshold (decimal KB) separating the two arms above. Unlike the two widths, this does not force sizing on where it is declined — it moves the threshold, it does not pin a width. |
CERA_THREADS |
detected perf-core count | Override the detected performance-core count (moves the auto width for both pools). |
CERA_MIN_ROWS |
128 | Minimum output rows a decode-GEMV worker takes before another joins. |
CERA_PAR_THRESHOLD |
256 | Minimum output dimension before a GEMV parallelizes; smaller GEMVs stay serial. |
CERA_SPIN |
100000 | Spin iterations before an idle worker parks. |
CERA_PIN |
on | 0 / false / off disables affinity pinning (for hosts that manage thread placement themselves). |
CERA_CPU_TIER |
auto | Force a lower CPU SIMD tier (downgrade only) — for parity testing on capable hardware. |
CERA_LM_HEAD_NO_GEMM |
unset | 1 puts the LM-head projection in forward_prefill_logits_all back on the per-row loop the batched GEMM replaced, so both halves of a speculative-decoding A/B run from one binary. Measurement lever only — both paths compute the same projection, to within f32 accumulation order. |
Affinity pinning applies on Linux/Android with a detected heterogeneous topology; homogeneous hosts and macOS run unpinned.
License
Apache-2.0 OR MIT.