TypedLM
Typed, testable LLM programs for Rust.
A language model call becomes a function with a typed contract: the input is a Rust struct, the output is a Rust type, and the answer is validated against it before your code sees it. The same contract lets you measure the program on a labelled dataset.
use JsonSchema;
use ;
use OpenAiCompatible;
use *;
/// Classify a customer support ticket by urgency and category.
let classify = new;
let ticket = classify.run.await?;
match ticket.urgency
No prompt to write and no JSON to parse: the doc comment becomes the instruction, the output type becomes the JSON Schema.
Install
Your types derive Serialize, Deserialize and JsonSchema. Any async runtime works;
the examples use tokio. Documentation: https://casoon.github.io/typedlm/, API:
https://docs.rs/typedlm.
What it does
- Typed contract.
#[derive(TypedLm)]on the input struct, anyDeserialize + JsonSchematype as output. Doc comments on types and fields reach the model. - Validation with repair. The answer is extracted (code fences, surrounding text
and trailing commas are tolerated), checked against the schema — every violation
with its path, e.g.
$.items[2].quantity— and against your own rules (#[lm(validate = fn)]). Invalid answers go back to the model with that list, up to a configurable number of times. - Four strategies. Native structured output, forced tool call, JSON mode or prompt only — the strongest one the provider supports is chosen, with the schema rewritten for the provider (e.g. OpenAI strict mode).
- Evaluation. Datasets in JSON Lines with partial labels,
ExactMatchandFieldAccuracy, and a report with a 95 % confidence interval, per-field accuracy, repair and failure rates, latency percentiles and token usage. - Regression tests. Store a report as baseline; later runs are compared example by example and fail only on a statistically significant drop, naming the examples that got worse. Several runs per example show how stable the answers are.
- Testing without a model. Fixed answers, closures, schema-valid answers, and
recording real interactions to replay them offline in
cargo test. - Compiled programs. Worked examples, instructions and settings stored as a reviewable JSON artefact with the evaluation it was accepted on; loading checks it against the signature.
- Optimization. The optimizer chooses the worked examples that score best on a validation set and has a teacher model rewrite the instructions from the program's mistakes; the result is a compiled program.
- Typed tools. Actions are enum variants; the model chooses one, your policy authorizes it and confirms writes, then your code executes it — one action per call, native tool calls where the provider supports them.
- Command line.
typedlm-cliruns programs across several models as a matrix, keeps every run as a snapshot and compares runs; thetypedlmbinary shows and compares stored reports, with exit codes for CI. - Tracing. One
tracingspan per call with OpenTelemetry GenAI field names. Inputs and answers are only recorded on request.
Providers
OpenAiCompatible talks to any OpenAI-compatible chat completions API: OpenAI,
local servers such as Ollama, and routers that translate to other vendors. Other
providers implement the Provider trait.
let openai = new.api_key;
let local = ollama;
Transport errors, 429 and 5xx are retried with backoff and Retry-After;
repairing invalid answers is counted separately.
Evaluation
{"input": {"text": "Our production database is down."}, "expected": {"urgency": "High", "category": "Technical"}}
{"input": {"text": "I was charged twice this month."}, "expected": {"category": "Billing"}}
let dataset = from_jsonl?;
let report = evaluate.await;
println!;
The interval is the point: eight examples say little, and the report shows it.
// In a test: fails only on a significant drop against the stored run.
at.check?;
Examples
TYPEDLM_MODEL=qwen3:32b
TYPEDLM_MODEL=qwen3:32b
TYPEDLM_MODEL=qwen3:32b
Without TYPEDLM_API_KEY the examples use a local Ollama; with it, the OpenAI API
(or TYPEDLM_BASE_URL).
Features
| Feature | Default | Adds |
|---|---|---|
derive |
yes | #[derive(TypedLm)] |
http |
yes | OpenAiCompatible (reqwest with rustls, tokio timer) |
eval |
yes | datasets, metrics, evaluate |
optimize |
no | optimizing worked examples and instructions (pathwise; Rust 1.88) |
Without default features the crate depends on serde, serde_json, schemars and
tracing only, and on no async runtime.
Status
0.1 — the API may still change before 1.0. Tested live against Ollama (qwen3:32b) with
all four strategies and native tools; OpenAI has not been tested live yet. Not in scope:
multi-step agents. MSRV 1.85 (1.88 with optimize).
License
MIT