# cera-cli
Command-line interface for the [`cera`](https://github.com/hyeons-lab/cera/tree/main/cera) LLM inference engine. Installs
a `cera` binary for running, chatting with, inspecting, and benchmarking GGUF /
LeapBundles models locally.
> **Note:** Part of a learning-experiment project exploring LLM inference
> internals in Rust — see the [project README](https://github.com/hyeons-lab/cera).
> Not intended for production use.
## Install
```sh
cargo install cera-cli
```
This builds the `cera` binary. For an Apple Metal or wgpu GPU build:
```sh
cargo install cera-cli --features metal # or gpu
```
## Usage
Point at a local model — a `.gguf` file, a `.json` LeapBundles manifest, or a
directory containing one — or let it auto-download a bundle by id/quant from
[`huggingface.co/LiquidAI/LeapBundles`](https://huggingface.co/LiquidAI/LeapBundles)
(cached under `$HOME/.cache/cera`). Supported architectures: `lfm2`,
`qwen2`/`qwen3`, `llama` (incl. classic Mistral), and `granite` — see the
[`cera` README](https://github.com/hyeons-lab/cera/tree/main/cera#supported-models)
for the full list and modality support.
```sh
# Generate from a local GGUF
cera run --model model.gguf --prompt "Explain quantization in one sentence."
# Auto-download a bundle and generate
cera run --bundle-id LFM2.5-1.2B-Instruct --quant Q4_0 --prompt "Hello"
# Constrain output to valid JSON (bundled grammar) or a custom GBNF
cera run -m model.gguf -p "List 3 colors as JSON" --json
cera run -m model.gguf -p "..." --grammar @schema.gbnf
# Tool calling: pass tool schemas (inline JSON or @file). stdout gets ONLY the
# JSON array of calls (the assistant reply + timing stream to stderr), so it
# pipes cleanly. Add --constrain-tools to force a valid call.
# Interactive multi-turn chat REPL (keeps the prefix cache warm across turns)
cera chat --bundle-id LFM2.5-1.2B-Instruct --quant Q4_0
# Attach a LoRA adapter (llama.cpp .gguf or PEFT .safetensors) for the session
cera run -m model.gguf -p "..." --lora adapter.safetensors
cera chat -m model.gguf --lora adapter.gguf
# Extract hidden-state embeddings for a prompt (mean-pooled [hidden_size] vector)
cera embed -m model.gguf -p "a chunk of text to embed"
cera embed -m model.gguf -p "a chunk" --per-token # full [T*hidden_size] matrix, one row per token
cera embed -m model.gguf -p "a chunk" --json # JSON array output instead of space-separated floats
```
### Commands
| `run` | Run inference on a prompt — text, optional grammar/JSON or tool calling (`--tools` / `--constrain-tools`), plus audio input for LFM2-Audio bundles. Optional `--lora` adapter. |
| `chat` | Interactive multi-turn REPL with `/help`, `/clear`, `/exit` slash commands. Optional `--lora` adapter. |
| `embed` | Extract last-layer hidden-state embeddings for a prompt — mean-pooled by default, `--per-token` for the full matrix, `--json` for array output. |
| `logits` | Dump the next-token logits over the full vocabulary for a prompt (single prefill) — `--top-k` for the K highest `(token_id, logit)` pairs, `--json` for array output. Handy for cross-backend parity checks. |
| `inspect` | Inspect a GGUF file's metadata and resolved CPU backend tier. |
| `cpu` | Print the host's CPU backend tier + detected SIMD features (no model needed). |
| `tokenize` | Tokenize text and print token IDs (e.g. to compare against HuggingFace). |
| `bench` | Measure decode throughput (tok/s) with p10/p50/p90/mean/stddev over N runs. |
| `list-bundles` | List bundles on `LiquidAI/LeapBundles` (add `--quants` for per-bundle quants). |
| `download-bundles` | Download bundle manifests + model files without loading them. |
Run `cera <command> --help` for the full flag list. Common `run` flags:
`--max-tokens` (default 256), `--temperature` (default 0.7), `--device`
(`cpu` / `gpu` / `metal` / `auto`, default `auto`), `--grammar` / `--json`, and
`--lora` to attach a LoRA adapter. For tool calling, `--tools <JSON|@file>`
passes an array of OpenAI-style function schemas (rendered into the chat
template; the reply's tool calls are parsed to a JSON array on stdout), and
`--constrain-tools` (requires `--tools`) forces a well-formed, correctly-typed
call via a grammar + lazy trigger. `run`, `chat`, and `embed` all accept
`--lora <PATH>` (a llama.cpp `.gguf` or PEFT `.safetensors` adapter); it applies
to every forward pass — generation and hidden-state extraction alike. For a PEFT
`.safetensors` adapter whose `alpha` differs from its rank, pass
`--lora-alpha <ALPHA>` (`scale = alpha / rank`; `.gguf` adapters carry alpha in
their metadata).
## License
Apache-2.0 OR MIT.