cera-cli
Command-line interface for the cera LLM inference engine. Installs
a cera binary for running, chatting with, inspecting, and benchmarking GGUF /
LeapBundles models locally.
Note: Part of a learning-experiment project exploring LLM inference internals in Rust — see the project README. Not intended for production use.
Install
This builds the cera binary. For an Apple Metal or wgpu GPU build:
Usage
Point at a local model — a .gguf file, a .json LeapBundles manifest, or a
directory containing one — or let it auto-download a bundle by id/quant from
huggingface.co/LiquidAI/LeapBundles
(cached under $HOME/.cache/cera). Supported architectures: lfm2,
qwen2/qwen3, llama (incl. classic Mistral), and granite — see the
cera README
for the full list and modality support.
# Generate from a local GGUF
# Auto-download a bundle and generate
# Constrain output to valid JSON (bundled grammar) or a custom GBNF
# Interactive multi-turn chat REPL (keeps the prefix cache warm across turns)
# Attach a LoRA adapter (llama.cpp .gguf or PEFT .safetensors) for the session
# Extract hidden-state embeddings for a prompt (mean-pooled [hidden_size] vector)
Commands
| Command | Purpose |
|---|---|
run |
Run inference on a prompt — text, optional grammar/JSON, plus audio input for LFM2-Audio bundles. Optional --lora adapter. |
chat |
Interactive multi-turn REPL with /help, /clear, /exit slash commands. Optional --lora adapter. |
embed |
Extract last-layer hidden-state embeddings for a prompt — mean-pooled by default, --per-token for the full matrix, --json for array output. |
inspect |
Inspect a GGUF file's metadata and resolved CPU backend tier. |
cpu |
Print the host's CPU backend tier + detected SIMD features (no model needed). |
tokenize |
Tokenize text and print token IDs (e.g. to compare against HuggingFace). |
bench |
Measure decode throughput (tok/s) with p10/p50/p90/mean/stddev over N runs. |
list-bundles |
List bundles on LiquidAI/LeapBundles (add --quants for per-bundle quants). |
download-bundles |
Download bundle manifests + model files without loading them. |
Run cera <command> --help for the full flag list. Common run flags:
--max-tokens (default 256), --temperature (default 0.7), --device
(cpu / gpu / metal / auto, default auto), --grammar / --json, and
--lora to attach a LoRA adapter. run, chat, and embed all accept --lora <PATH> (a llama.cpp .gguf or PEFT .safetensors adapter); it applies to every
forward pass — generation and hidden-state extraction alike. For a PEFT
.safetensors adapter whose alpha differs from its rank, pass --lora-alpha <ALPHA> (scale = alpha / rank; .gguf adapters carry alpha in their metadata).
License
Apache-2.0 OR MIT.