cera-cli 0.2.2

CLI for the Cera LLM inference engine
cera-cli-0.2.2 is not a library.

cera-cli

Command-line interface for the cera LLM inference engine. Installs a cera binary for running, chatting with, inspecting, and benchmarking GGUF / LeapBundles models locally.

Note: Part of a learning-experiment project exploring LLM inference internals in Rust — see the project README. Not intended for production use.

Install

cargo install cera-cli

This builds the cera binary. For an Apple Metal or wgpu GPU build:

cargo install cera-cli --features metal   # or gpu

Usage

Point at a local model — a .gguf file, a .json LeapBundles manifest, or a directory containing one — or let it auto-download a bundle by id/quant from huggingface.co/LiquidAI/LeapBundles (cached under $HOME/.cache/cera). Supported architectures: lfm2, qwen2/qwen3, llama (incl. classic Mistral), and granite — see the cera README for the full list and modality support.

# Generate from a local GGUF
cera run --model model.gguf --prompt "Explain quantization in one sentence."

# Auto-download a bundle and generate
cera run --bundle-id LFM2.5-1.2B-Instruct --quant Q4_0 --prompt "Hello"

# Constrain output to valid JSON (bundled grammar) or a custom GBNF
cera run -m model.gguf -p "List 3 colors as JSON" --json
cera run -m model.gguf -p "..." --grammar @schema.gbnf

# Interactive multi-turn chat REPL (keeps the prefix cache warm across turns)
cera chat --bundle-id LFM2.5-1.2B-Instruct --quant Q4_0

# Attach a LoRA adapter (llama.cpp .gguf or PEFT .safetensors) for the session
cera run -m model.gguf -p "..." --lora adapter.safetensors
cera chat -m model.gguf --lora adapter.gguf

# Extract hidden-state embeddings for a prompt (mean-pooled [hidden_size] vector)
cera embed -m model.gguf -p "a chunk of text to embed"
cera embed -m model.gguf -p "a chunk" --per-token   # full [T*hidden_size] matrix, one row per token
cera embed -m model.gguf -p "a chunk" --json        # JSON array output instead of space-separated floats

Commands

Command Purpose
run Run inference on a prompt — text, optional grammar/JSON, plus audio input for LFM2-Audio bundles. Optional --lora adapter.
chat Interactive multi-turn REPL with /help, /clear, /exit slash commands. Optional --lora adapter.
embed Extract last-layer hidden-state embeddings for a prompt — mean-pooled by default, --per-token for the full matrix, --json for array output.
inspect Inspect a GGUF file's metadata and resolved CPU backend tier.
cpu Print the host's CPU backend tier + detected SIMD features (no model needed).
tokenize Tokenize text and print token IDs (e.g. to compare against HuggingFace).
bench Measure decode throughput (tok/s) with p10/p50/p90/mean/stddev over N runs.
list-bundles List bundles on LiquidAI/LeapBundles (add --quants for per-bundle quants).
download-bundles Download bundle manifests + model files without loading them.

Run cera <command> --help for the full flag list. Common run flags: --max-tokens (default 256), --temperature (default 0.7), --device (cpu / gpu / metal / auto, default auto), --grammar / --json, and --lora to attach a LoRA adapter. run, chat, and embed all accept --lora <PATH> (a llama.cpp .gguf or PEFT .safetensors adapter); it applies to every forward pass — generation and hidden-state extraction alike. For a PEFT .safetensors adapter whose alpha differs from its rank, pass --lora-alpha <ALPHA> (scale = alpha / rank; .gguf adapters carry alpha in their metadata).

License

Apache-2.0 OR MIT.