laya
Rust inference for Laya, a non-autoregressive typed-decision model. You give it a state (text or JSON) and a set of typed questions; it returns typed answers with calibrated probabilities in a single forward pass. It never generates text, so there is nothing to parse and nothing to hallucinate.
Pure Rust on candle — no Python, no torch, no ONNX export step. Runs on CPU, Metal, or CUDA.
Getting the weights
The root checkpoint is Apache-2.0 and ungated, so no token is needed:
B=https://huggingface.co/convaiinnovations/laya/resolve/main
for; do
done
That is ~847 MB. The repo also carries two sibling variants under multilingual/ and
typed-decisions/; this crate loads any of them, since the layout is identical — point
--model at the directory you downloaded.
CLI
Web console
A local single-page console: the input on the left, the decision on the right. Each answer
shows the verdict, its probability, the confidence, and one line naming the runner-up — enough
to see a close call without reading a full distribution. Table view opens the rest: every
option's probability, the score legend, and P(answer).
Two things it shows that the CLI cannot:
- A prompt inspector, collapsed under the answers. The exact token sequence the encoder
saw, coloured by segment (instructions / options / state / structure) with each scored
[MASK]marker ringed and numbered, plus how much of themax_lenbudget the state used — and a red state truncated warning on the collapsed header when it overflowed. - Presets for support triage, content moderation, agent-trace review, and the four banking sessions below.
⌘/Ctrl + ↵ runs. Works in light and dark, and down to phone width. It binds to loopback only
and has no auth — it is a dev tool, not a deployment.
Banking demo
banking/ is a worked end-to-end example: a Session / Event / User / IP / Tag data model for
internet banking, four sample sessions (routine, account takeover, genuine travel, mule
pass-through), four typed questions, and the measured results — including which two of the four
questions actually carry signal on the base checkpoint and which do not. See
banking/README.md.
Library
use ;
use json;
let agent = from_dir?;
let questions = vec!;
let response = agent.system_one?;
println!;
# Ok::
All questions for one state go through a single batched forward pass, so asking five questions costs about as much as asking one.
The three question types
| Type | criteria |
Answer |
|---|---|---|
choice |
list of labels, or {label: description} |
the argmax label plus a distribution over labels |
score |
list of level descriptions | the expectation over level indices, plus the distribution |
noul |
optional {"false": ..., "true": ...} |
P(the statement holds) |
choice criteria keep the order you wrote them in, and that order is what the label
probabilities are reported against.
Every answer also carries rl_agent.act_probability: the act head's probability of answering
rather than escalating to a stronger model.
How it works
flowchart LR
S["state<br/>(text or JSON)"] --> B
Q["typed question"] --> B
B["[CLS] type question: instructions [SEP]<br/>[MASK] opt0 [MASK] opt1 … [SEP]<br/>state [SEP]"] --> E
E["ModernBERT-large<br/>28 layers, 1024 hidden"] --> T["+ question-type embedding"]
T --> H["decision head<br/>2 pre-norm transformer layers"]
H --> G["gather the [MASK] marker<br/>at each option"]
G --> SC["scorer → one logit per option"]
SC --> TS["÷ temperature<br/>(per type and option count)"] --> P["calibrated distribution"]
H --> AP["pooled [CLS]"] --> AH["act head"] --> AC["P(answer) vs P(escalate)"]
SC -.->|"top-1, margin, entropy, k"| AH
Each option gets a [MASK] marker in the prompt; the head scores those marker positions and
softmaxes over them. That is why the whole question set resolves in one pass and why the model
cannot emit anything outside your option list.
Configuration
rl_agent_config.json sets the budgets: max_len 512 and head_max_len 192 on the root
checkpoint. Options share the head budget, so a question with many long options will fail with
an explicit head_max_len error rather than silently dropping options. For more than ~20
options, raise both budgets (the encoder supports up to 8192 positions) or split the question
into a coarse-to-fine pair.
Backends
f32 only. Weights are f16 on disk and are upcast when loaded, so peak memory is around 2.4 GB; candle's ModernBERT builds its attention mask as f32 unconditionally, so an f16 backbone fails inside the first attention block.
Calibration
The checkpoint ships over-confident. The bundled temperatures move mean ECE from 0.466 to
0.081 on the authors' data, but those were fitted on their distribution. Refit one
temperature per (question type, option count) bucket on your own data before trusting the
probabilities as probabilities; temperature_by_options in rl_agent_config.json is the knob,
keyed choice:3-5, noul:2, and so on.
Also worth knowing before you build on it: the base checkpoint is near chance on the authors'
own typed-decisions benchmark (0.362 against a 0.461 majority-class baseline). The strong
numbers belong to the fine-tuned typed-decisions variant. Treat this as a fast base to
specialise, not a zero-shot decision engine.
Tests
The rendering tests run anywhere. The inference tests skip themselves unless a checkpoint is
present at models/laya-base (override with LAYA_MODEL_DIR).
License
Apache-2.0, matching the upstream model. The console's header image is supplied by the repository owner and is not covered by that license.