llm-verify
English · 简体中文
Verify the LLM endpoint you are actually using: model authenticity, billing inflation, relay provenance, performance and silent downgrades.
A single binary with no runtime dependencies. Results come out as an HTML report you open in a browser.
Install
From crates.io, with a Rust toolchain (1.82+):
Without a toolchain, the install script fetches a prebuilt binary:
|
On Windows: irm https://raw.githubusercontent.com/asale-ai/llm-verify/main/install.ps1 | iex
Use
It writes an HTML report and opens it.
Credentials can live in .env or the environment instead, leaving only the model on the command line:
# .env
LLM_VERIFY_BASE_URL=https://api.anthropic.com
LLM_VERIFY_API_KEY=sk-ant-...
LLM_VERIFY_MODEL=claude-opus-4-5
Example run
Probing gpt-5.6-sol through OpenRouter at the default balanced depth — 40 probes, 47 requests, 131.2s:

The summary that follows carries the verdict, the two axes, and the per-group scores:

Relayed / Relay: a real model, reached through one hop. Channel provenance and performance are what drag the groups down, not identity or billing.
Full HTML reports from two runs against the same model:
| Report | Verdict | Origin | Score |
|---|---|---|---|
| openrouter.ai (preview) | Relayed | Relay | 92 / 100 |
| gw.asale.ai (preview) | Relayed | Undetermined | 93 / 100 |
Same model, same verdict, two points apart — and the reports still differ. OpenRouter names itself in the response headers, so the origin is pinned to a single relay hop; the second path carries no channel markers at all, which leaves the origin undetermined rather than proven clean. Origin is a separate reading from both score and verdict.
Options
| Flag | Meaning |
|---|---|
--protocol anthropic|openai |
Inferred from the URL and model name if omitted |
--depth fast|balanced|forensic |
Default balanced. forensic samples more — slower and costlier, but firmer |
--claimed-model <ID> |
Use when the vendor's advertised name differs from the ID you request. This is how you check for a downgrade |
--lang en|zh |
Report language. Follows the system locale, then falls back to English |
-o <path> |
HTML report path; pass a directory to auto-name the file |
--json <path> |
Also emit machine-readable JSON |
--no-open |
Do not open a browser |
Exit codes
| Code | Meaning |
|---|---|
| 0 | Clean |
| 1 | Failing score, or a suspicious / counterfeit / inconclusive verdict |
| 2 | A hard gate tripped |
Usable as a CI gate as-is.
Reading the report
The verdict has two independent axes:
- Authenticity — genuine / genuine-with-defects / relayed / suspicious / counterfeit / inconclusive
- Origin — direct from vendor / cloud platform / subscription-derived / relay / reconstructed channel / undetermined
A real model behind a relay is relayed, not counterfeit. Longer path, same model.
Hard gates are facts no weighted score can excuse. Any one of them forces a suspicious verdict and exit code 2:
silent fallback · shared-pool forwarding · tier downgrade · third-party wrapper injection · cache replay · hidden prompt injection · response replay
Use it from an AI coding tool
The skill lives in this repository at skills/llm-verify/SKILL.md. Install it with skills, which supports Claude Code, Codex, Cursor, OpenCode and 70-odd other agents:
It is also published on ClawHub:
Then just ask: "is this API actually giving me what I paid for?"
The skill is only the usage guide — the llm-verify binary does the work, so you need both.
What it checks
40 probes across seven groups:
| Group | Question it answers |
|---|---|
| Protocol contract | Is this a genuine API channel? |
| Streaming | Does streaming follow the protocol, or arrive empty? |
| Metering & billing | Are the token counts honest, or are you overcharged? |
| Channel provenance | What relays sit on this path? |
| Performance | First-token latency, throughput and jitter |
| Model identity | Is the model behind this the one that was sold? |
| Cross-request consistency | Does the endpoint behave the same way every time? |
Use it as a library
The binary is a thin CLI over the crate, so an embedder gets the same probes, the same order and the same verdict — which is what lets it claim its own results and the published tool agree.
[]
= { = "0.4", = false }
default-features = false drops the CLI half: clap, the terminal writer and the HTML renderer are dead weight in a service.
use ;
let cfg = new;
let report = run.await?;
println!;
Probing through a relay
Half the suite asks about the endpoint — its error envelopes, its response headers, the token counts it reports. Behind a relay those describe the relay, not the model, and several of them deliberately send malformed or unauthenticated requests, which is not traffic you want attributed to somebody else's account.
Selection::model_only() keeps only the steps whose evidence is the generated text itself, which survives any number of hops:
let cfg = new
.model_only
.depth
// Chosen by the caller and recorded in the report, so a contested verdict
// can be replayed probe for probe.
.seed
// Spread the run out instead of issuing it as one recognisable burst.
.pace;
Progress, cancellation, and your own probes
engine::run reports each step as it starts and finishes, and takes a Cancel that is honoured between steps. Selection::with appends your own [Probe] implementations, which share the run's client, seed and observation buffers — so a private probe is indistinguishable from a built-in one both in the report and on the wire. That matters if you are policing a marketplace: everything in this repository is readable by the endpoint you are probing.
What it costs
Measured, not estimated — tests/request_budget.rs pins these so that adding a probe cannot quietly multiply anyone's bill:
| Selection | Depth | Requests |
|---|---|---|
model_only |
fast |
21 |
model_only |
forensic |
46 |
| everything | balanced |
49 |
Limits
One false accusation against an honest provider costs far more than one miss. This tool abstains when the evidence is thin rather than guessing. Please know the following:
- Resolution stops at tier granularity (flagship / mid / light). Adjacent versions inside one tier — say 4.5 and 4.6 of the same line — cannot be separated.
- The tier call depends on sampling. The default depth asks 9 capability questions, and adjacent tiers can still swing on a single one. The tool abstains when the margin is narrow, which costs it some genuine downgrades. Use
--depth forensic(15 questions) when the answer has to hold up. The report states how many questions the call rests on and by how much the winner beat the runner-up. - Delivering above the claimed tier is not fraud and carries no risk weight. Only measuring below the claim counts.
- Middle layers contaminate identity fingerprints, which is why the contract layer runs first; where injection is found, identity confidence is reduced automatically.
- Server-side weights cannot be proven — only whether behaviour matches expectations.
- One run describes one moment. Gradual degradation needs periodic re-runs and comparison.
- The probes are public. That is the point of an auditable tool, and it is a ceiling: an endpoint can read this repository too, and answer these questions honestly while serving everything else from somewhere cheaper. Timing can be disguised (
RunConfig::pace); the questions cannot be. Anyone using this to police supply they do not control needs probes of their own on top — seeSelection::with— and should treat the result as a probability rather than a proof.