# llm-browser-testkit
Describe browser tests in plain English. The LLM figures out which elements to
click and whether the page looks right — plus A2A agents, MCP tool-calling, cost
tracking, and budgets.
```
llm-browser-testkit run smoke.toml
```
## Contents
- [Quick start](#quick-start)
- [Write your first test](#write-your-first-test)
- [Step reference](#step-reference)
- [Assertion presets](#assertion-presets)
- [CLI reference](#cli-reference)
- [Endpoints](#endpoints)
- [A2A agents](#a2a-agents)
- [Run as an A2A agent](#run-as-an-a2a-agent)
- [MCP tools](#mcp-tools)
- [MCP server](#mcp-server-exposure)
- [Cost tracking & budgets](#cost-tracking--budgets)
- [How it works](#how-it-works)
- [Use as a library](#use-as-a-library)
- [LLM authentication](#llm-authentication)
- [License](#license)
## Quick start
Everything you need for a first green run — copy-paste, no prior setup:
```bash
# 1. Install the CLI
cargo install llm-browser-testkit
# 2. Point it at any OpenAI-compatible API
export HARNESS_LLM_TEST_URL=https://api.openai.com
export HARNESS_LLM_TEST_MODEL=gpt-4o-mini
export HARNESS_LLM_API_KEY=sk-...
# 3. Describe one test in a tiny TOML file (example.com — no account needed)
cat > hello.toml <<'EOF'
[config]
base_url = "https://example.com"
start_url = "/"
[[test]]
name = "Homepage loads"
[[test.steps]]
kind = "navigate"
url = "/"
[[test.steps]]
kind = "assert"
preset = "no_error_on_page"
EOF
# 4. Run it — Chrome runs headless, the LLM checks the page
llm-browser-testkit run hello.toml
```
Example output:
```
Test: Homepage loads — passed (6.2s, $0.0005, 138 tokens, 1 calls, 2+0+0 steps)
## Write your first test
The quick-start `hello.toml` is the smallest useful scenario. Its three
blocks:
- `[config]` — `base_url` is the app under test; `start_url` is where Chrome
loads first.
- `[[test]]` — one named test, built from `[[test.steps]]` that run top to
bottom.
- Steps — every step has a `kind`: `navigate` opens a URL relative to
`base_url`; `assert` sends the page to the LLM and expects `PASS`/`FAIL`.
`preset` picks a built-in check (`no_error_on_page` = no errors, stack
traces, or broken UI).
Reusable checks go into `[[definitions]]` — name a preset or prompt once and
reference it from any assertion:
```toml
[[definitions]]
name = "no_errors"
preset = "no_error_on_page"
[[definitions]]
name = "example_domain_visible"
preset = "text_visible"
assert_text = "Example Domain"
[[test]]
name = "Homepage loads"
[[test.steps]]
kind = "navigate"
url = "/"
[[test.steps]]
kind = "assert"
definition = "no_errors"
[[test.steps]]
kind = "assert"
definition = "example_domain_visible"
```
Run it:
```bash
llm-browser-testkit run hello.toml
```
## Step reference
Every step has a `kind`. Required fields depend on the kind.
| `navigate` | Open a URL | `url` | `wait_after_ms` |
| `click` | Click an element | `target` | `selector`, `wait_after_ms`, `endpoint`, `idempotent` |
| `type` | Type into a field | `target`, `text` | `selector`, `wait_after_ms`, `endpoint`, `idempotent` |
| `wait` | Wait for an element and/or visible text | `target` | `selector`, `text`, `timeout_ms`, `endpoint`, `idempotent` |
| `assert` | Check the page | one of `definition`, `preset`, or `prompt` | `assert_text`, `endpoint` |
| `screenshot` | Save a .png | — | `path` |
| `agent` | Call an A2A agent | `agent`, `task` | `definition` |
| `mcp` | Call an MCP tool | `server`, `tool` | `args` |
**`target`** is natural language ("the submit button", "the search input"). The
LLM looks at the page DOM and picks the right CSS selector at runtime. Skip the
LLM with an explicit `selector`.
**`idempotent`** — an optional flag on `click`, `type` and `wait` steps.
When the step's target is absent, the step is reported **skipped** instead
of failed: the action was already done or not applicable. This is the
generic building block for flows that repeat in one browser session —
e.g. logging in on every viewport-matrix variant of the same test:
```toml
[[test.steps]]
kind = "navigate"
url = "/auth/login"
# Already authenticated? The form is gone, so these steps skip
# instead of failing.
[[test.steps]]
kind = "type"
selector = "#email"
target = "the email input"
text = "admin@example.com"
idempotent = true
[[test.steps]]
kind = "type"
selector = "#password"
target = "the password input"
text = "correct horse battery staple"
idempotent = true
[[test.steps]]
kind = "wait"
selector = "input[name="cf-turnstile-response"][value]:not([value=""])"
target = "the bot-protection token"
timeout_ms = 15000
idempotent = true
[[test.steps]]
kind = "click"
selector = "button.btn--landing.btn--primary"
target = "the sign-in button"
idempotent = true
# The final check stays strict: a real login attempt that never
# reaches the authenticated shell still fails the test.
[[test.steps]]
kind = "wait"
selector = "app-account-shell"
target = "the authenticated shell"
timeout_ms = 30000
```
Semantics:
- `idempotent` `click` / `type`: probe for up to 5s; absent target → skipped.
- `idempotent` `wait`: run the wait as normal; a timeout → skipped instead
of failed.
- Skipped steps do **not** fail the test and do **not** trigger fail-fast.
- Keep the final verification step strict (no `idempotent`) so real
failures in the middle of an idempotent flow still surface.
**`endpoint`** routes this step to a specific [endpoint](#endpoints). Use it to
send element targeting to one model and assertions to another.
**`wait` with `text`** waits until the page's visible text contains a
substring — no selector or LLM needed:
```toml
[[test.steps]]
kind = "wait"
target = "the success message"
text = "Welcome back"
timeout_ms = 5000
```
Set both `selector` and `text` to require both conditions. The combined wait
shares one `timeout_ms` budget.
## Failure diagnostics & artifacts
When a step fails, the runner captures the current page state and writes a
screenshot, so CI logs answer *why* the step failed instead of printing a bare
timeout:
```
✗ [wait] the authenticated shell — wait for app-account-shell timed out after
30000ms: The event waited for never came — page: http://127.0.0.1:8082/auth/login
(Immosai) — visible: "Email address ⏎ Password ⏎ Sign in" (30.1s)
│ url: http://127.0.0.1:8082/auth/login
│ title: Immosai — Anmeldung
│ content: Email address Password ... Invalid credentials.
screenshot: artifacts/account__login-and-open-account__006-wait.png
```
- **Page state** — URL, title, visible text, and any alert/error elements
(`[role="alert"]`, `.error-message`, snackbars, …) are appended to the step
message and printed in full to stderr.
- **Screenshots** — one PNG per failed step, written under
`--artifacts-dir` (default `artifacts/`, env `HARNESS_ARTIFACTS_DIR`).
- **Fail fast** — by default the first failed step ends the test and the
remaining steps are reported as skipped (no LLM budget is burned asserting
against a page that is already known broken). Set
`continue_on_failure = true` in `[config]` or pass `--continue-on-failure`
to keep executing every step.
- **LLM element targeting is verified** — a response that is not a selector
(`:not(*)`, `null`, explanations, …) fails immediately with the raw LLM
output; a selector that matches nothing triggers one retry with feedback.
- **LLM errors are specific** — HTTP status, a truncated response-body
snippet, and the attempt count are included, and deterministic client
errors (401/403/404) fail fast instead of burning the retry budget.
- **Assertions always see the page** — custom `system`+`user_template`
definitions that omit the `{content}` placeholder automatically get the
page URL/title/content appended, so the LLM never answers "I can't
determine that without seeing the page".
## Reporting: human- and machine-readable runs
All output flows through a single event stream. Every event (test/step
started + finished, LLM call with duration/tokens/cost, budget warning) is
rendered for humans and serialized for machines:
- **Console** — level-filtered, ASCII-safe, colors only on a TTY
(respects `NO_COLOR`). Default shows config, per-step results with
durations, the run summary and the cost report. Use `-q`/`-qq` to hide
step results (then warnings too), or `-v`/`-vv` to add LLM call details
(endpoint, model, duration, tokens, cost) and step starts.
- **NDJSON log** (`--log-file run.jsonl`) — one JSON object per event with
a `type` discriminator and an epoch-ms `ts` field. Lossless apart from
secret redaction: untruncated messages, ideal for CI artifact analysis:
```
jq '. | select(.type == "step_finished" and .status == "failed")' run.jsonl
jq '. | select(.type == "llm_call_finished") | {endpoint, ok, duration_ms, cost}'
```
- **JUnit XML** (`--junit report.xml`) — one `<testcase>` per test with a
`<failure>` per failed step, for Jenkins/GitLab/Azure/TeamCity.
- **Perfetto trace** (`--trace run.json`) — test/step/LLM spans in Chrome
Trace Event Format, viewable at https://ui.perfetto.dev.
- **GitHub Actions** — in CI the reporter automatically emits `::error::`
annotations for failed steps (with the screenshot as `file=`) and appends
a run summary to `GITHUB_STEP_SUMMARY`.
Truncation (`<truncated N chars>`) is always boundary-safe (multi-byte
UTF-8 can never panic it) and reports how much was cut; full text is
preserved in the NDJSON log.
## Secret redaction
Every sink is redacted through one chokepoint, so a leaked secret can never
make it into a log or CI report:
- **API keys** — `llm_api_key` and per-endpoint `api_key`
- **Static credentials** — Entra `auth.client_secret`, AWS
`secret_access_key` / `session_token`
- **Sensitive header values** — values of `authorization`-family headers
(matched case-insensitively: `authorization`, `api-key`, `x-api-key`,
`x-auth-token`, `token`, `cookie`, …) in `llm_headers` and endpoint
`headers`
- **Runtime-obtained tokens** — token-command and header-command output,
Entra client-credentials and managed-identity access tokens (registered
the moment they are fetched, so the `LlmCallFinished` event that echoes
them is redacted too)
- **Explicit extras** — `--redact <SECRET>` (repeatable) or the
`HARNESS_REDACT` env var (comma-separated), for secrets not present in
the config (URL query tokens, scenario-embedded test data):
```
llm-browser-testkit run scenario.toml --redact 'abc123tokenxyz'
HARNESS_REDACT='abc123tokenxyz,xyz789tokenabc' llm-browser-testkit run scenario.toml
```
Secrets shorter than 6 characters are skipped when derived from config or
runtime sources so short values (e.g. `"dev"`) do not destroy log
readability; explicit `--redact` values always apply. Replacement is
exact-match and case-sensitive. Raw `eprintln!` sites that bypass the
reporter (MCP/A2A server startup banners, the `#[browser_test]` run-report
strings, the cost report) are not redacted.
## Vision assertions (screenshots)
Text/DOM evaluation cannot see *how* the page renders — overlapping elements,
clipped text, or a cookie banner covering the content are invisible to
`innerText`. Mark an endpoint as vision-capable and attach a screenshot to an
assert step to let the LLM evaluate the actual pixels:
```toml
[config] # optional: cap the screenshot size
screenshot_max_height = "20x" # full page, tiled up to 20× the viewport height
screenshot_max_dimension = 1400
[config.endpoints.vision] # MUST declare vision = true
type = "llm"
url = "https://api.openai.com"
model = "gpt-4o"
api_key = "sk-..."
vision = true # ← the flag
pricing = { input_per_1m_tokens = 2.50, output_per_1m_tokens = 10.00 }
[[definitions]]
name = "no_overlaps"
preset = "visual_no_overlaps"
[[test.steps]]
kind = "assert"
definition = "no_overlaps"
endpoint = "vision"
screenshot = true # ← attach the viewport screenshot
```
How it works:
- The **full scrollable page** is captured in a single CDP call (via the
`captureBeyondViewport` flag — nothing below the fold is skipped) and
split into viewport-tall tiles from the top, covering at most
`screenshot_max_height` (default `"20x"` = twenty viewports).
- Each tile is downscaled in Rust (Lanczos) so its longest edge is at most
`screenshot_max_dimension` (default 1400) and re-encoded as quality-85
JPEG — no page JS, deterministic, and every tile keeps 1:1 detail at its
own depth (a single downscaled composite would lose all detail past ~4
viewport heights). A 14400px page at 720px viewport and `"20x"` sends
twenty 1280×720 tiles covering the whole page; `"4x"` sends four tiles
covering 2880px. Tile count is hard-capped at 30.
- The tiles are sent as OpenAI-compatible `image_url` content parts next
to the text prompt (which still includes the page text for context),
ordered from the top of the page down. Token cost is effectively
coverage ÷ viewport height — the cap bounds it. Tile count is hard-capped
at 30 so no assertion can ever produce an unbounded request.
- `screenshot_max_height` accepts an absolute pixel count (`2880`, useful
when you know exactly how far down a page is dynamic) or a viewport
multiple (`"2x"`, `"20x"` — auto-follows viewport-matrix and per-test
viewport overrides). Values below the viewport height are raised to it,
so the visible viewport is always fully included (`0` / `"0x"` = exactly
the viewport, the pre-full-page behavior).
- Built-in presets: `visual_no_issues`, `visual_no_overlaps`,
`visual_text_visible` (uses `assert_text`). Custom `screenshot = true`
prompts work too.
- A `screenshot = true` step that resolves to an endpoint without
`vision = true` fails immediately with a clear configuration error.
- Text-only workflows are untouched: without `screenshot = true` the
request keeps the plain string `content` shape.
See [`examples/visual-overlays.toml`](examples/visual-overlays.toml) + the
bundled [`examples/visual-test-page.html`](examples/visual-test-page.html)
fixture for a runnable demo that passes on a clean page and detects a cookie
banner overlay. The prompts of the visual presets are intentionally strict
("only fail on clearly visible, user-impacting defects") — tune them per app
if your overlay detection needs to be more or less sensitive.
## DOM layout assertions (`layout_no_issues`)
Vision models see pixels but cost money per page × viewport. For cheap,
deterministic layout coverage there is a DOM-only preset that never calls
the LLM:
```toml
[[test.steps]]
kind = "assert"
preset = "layout_no_issues" # no endpoint, no screenshot, no tokens
```
It evaluates a geometry scan in the page and fails with the detected issues:
- **page-overflow-x** — the document is wider than the viewport
(horizontal scrolling or a runaway element);
- **element-out-of-viewport** — a visible element that scrolling cannot
reveal: a fixed element outside the viewport, left/negative overflow,
right-edge overflow beyond the scrollable content, or bottom overflow
on a page that cannot scroll down. Below-the-fold content on a tall
scrollable page is normal flow and is NOT reported;
- **text-clipped** — content inside an `overflow: hidden` container is
measurably larger than the box (cut-off text);
- **element-overlap** — an interactive element's center point is covered
by a different element that would intercept the click.
Intentional stacking (off-canvas drawers, dropdowns, badges, sticky
headers) is excluded by position/relation filters. Elements whose class
matches a prefix in `[config] layout_ignore_classes` are skipped by the
fixed-element, text-clipped, and overlap checks — the default covers the
Angular CDK screen-reader helpers (`.cdk-visually-hidden`,
`.cdk-describedby-message-container`, `.cdk-overlay-container`), which
are intentionally 1x1 / off-screen:
```toml
[config]
layout_ignore_classes = ["cdk-visually-hidden", "cdk-describedby-message-container", "cdk-overlay-container"]
```
Run it after every page load — it is free, so it is also the perfect
companion for the viewport matrix below.
## Viewport matrix (mobile / tablet / desktop)
`[config.viewport_matrix]` expands **every test** in a scenario into one
variant per named viewport. Each variant overrides the browser viewport via
CDP device-metrics emulation and gets a ` — <name>` suffix on the test name;
per-test budgets apply per variant.
```toml
[config.viewport_matrix]
viewports = [
{ name = "mobile", width = 390, height = 844 },
{ name = "tablet", width = 768, height = 1024 },
{ name = "desktop", width = 1280, height = 720 },
]
[[test]]
name = "Dashboard renders"
steps = [
{ kind = "navigate", url = "/dashboard", wait_after_ms = 2000 },
{ kind = "assert", preset = "layout_no_issues" },
]
```
The above runs "Dashboard renders — mobile", "— tablet" and "— desktop",
each at its viewport, and the layout scan flags sticky overlays, off-screen
text, or covered controls per size. Use it with `screenshot = true` +
`visual_no_issues` on a vision endpoint for pixel-level checks on top.
Single-test overrides work too — any `[[test]]` may set
`viewport_width` / `viewport_height` directly, which also switches the
browser viewport via CDP for just that test:
```toml
[[test]]
name = "Narrow phone layout"
viewport_width = 320
viewport_height = 568
```
## Assertion presets
Built-in presets you can use inline or from `[[definitions]]`.
| `no_error_on_page` | No errors, stack traces, or broken UI on the page |
| `text_visible` | Specific text appears on the page (`assert_text`) |
| `element_exists` | A described UI element is present |
| `layout_no_issues` | **DOM scan, no LLM**: page overflow, elements out of viewport, clipped text, covered controls |
| `visual_no_issues` | **Screenshot**: no layout/rendering defects (overlaps, clipping, cut-off content, broken images, blank panels) |
| `visual_no_overlaps` | **Screenshot**: no elements covering other content or intercepting clicks |
| `visual_text_visible` | **Screenshot**: `assert_text` is fully visible and readable (not clipped or covered) |
Custom assertions with `prompt` send any question to the LLM:
```toml
[[test.steps]]
kind = "assert"
prompt = "Does the page have a heading that says 'Example Domain'?"
```
Custom presets with `system` + `user_template` let you define reusable assertion
logic with template variables `{url}`, `{title}`, `{content}`, `{expected_text}`,
and `{description}`. Forgetting `{content}` is no longer a problem — the page
context is appended automatically whenever the template does not reference it:
```toml
[[definitions]]
name = "text_matches"
system = "You are a QA tester."
user_template = "Does the page at {url} contain the text: {expected_text}?"
assert_text = "Welcome back"
```
## CLI reference
```
llm-browser-testkit run <scenario.toml> [OPTIONS]
```
| `--llm-url` | `$HARNESS_LLM_TEST_URL` or `http://localhost:8080` | OpenAI-compatible endpoint |
| `--llm-model` | `$HARNESS_LLM_TEST_MODEL` or `deepseek` | Model name |
| `--llm-api-key` | `$HARNESS_LLM_API_KEY` | API key (Bearer token) |
| `--llm-fallback-url` | `$HARNESS_LLM_FALLBACK_URL` | Fallback endpoint tried when the primary exhausts its attempts |
| `--llm-fallback-model` | `$HARNESS_LLM_FALLBACK_MODEL` | Fallback model name |
| `--llm-fallback-api-key` | `$HARNESS_LLM_FALLBACK_API_KEY` | Fallback API key |
| `--llm-header` | — | Custom header `Name:Value` (repeatable) |
| `--redact` | `$HARNESS_REDACT` (comma-separated) | Literal value redacted from all rows/sinks (repeatable) |
| `--model-param` | — | Provider param `key=value` (repeatable) |
| `--base-url` | `$HARNESS_BROWSER_BASE_URL` or `http://localhost:4200` | App under test |
| `--headless` | `true` | Run Chrome headlessly |
| `--timeout` | `60` | Seconds per action |
| `--viewport-width` | `1280` | Browser width |
| `--viewport-height` | `720` | Browser height |
| `--start-url` | `/dashboard` | First page to load |
| `--max-cost` | — | Global budget: max USD across all tests |
| `--max-tokens` | — | Global budget: max tokens across all tests |
| `--budget-enforcement` | `hard` | Budget mode: `hard` (abort) or `soft` (warn) |
| `--artifacts-dir` | `$HARNESS_ARTIFACTS_DIR` or `artifacts` | Directory for failure screenshots |
| `--continue-on-failure` | off | Keep running remaining steps after a step failure (default: fail fast) |
| `-q`, `--quiet` | off | Repeatable: hide step results (`-q`), then warnings too (`-qq`). Failures and the run summary always show |
| `-v`, `--verbose` | off | Repeatable: show LLM call details and step starts (`-v`), then everything (`-vv`) |
| `--log-file` | — | Write a machine-readable NDJSON event log (one JSON object per event, `type` + `ts` fields) |
| `--junit` | — | Write a JUnit XML report for CI systems |
| `--trace` | — | Write a Perfetto-format trace (test/step/LLM spans) |
| `--color` | `auto` | Console colors: `auto` (TTY + `NO_COLOR` aware), `always`, `never` |
CLI flags override the scenario `[config]`.
**Retry + fallback behavior:** every LLM call is retried up to
`HARNESS_LLM_CALL_ATTEMPTS` times (default 3) on transient failures
(network errors, HTTP 429/5xx, invalid JSON, and HTTP 200 with an empty
body — the gateway warm-up signature). When an endpoint still fails, the
fallback chain is tried: `--llm-fallback-url`/`--llm-fallback-model`/
`--llm-fallback-api-key` (or `$HARNESS_LLM_FALLBACK_*`) configure a single
fallback endpoint for the implicit default endpoint. Pair a cheap primary
with a more expensive, more powerful fallback — the fallback is only billed
when the primary fails. Scenarios that declare `[config.endpoints]` use
per-endpoint `fallbacks = [...]` instead (see below).
## Endpoints
Define multiple named endpoints — LLM providers, MCP servers, and A2A agents —
each with their own pricing, and route test steps to them automatically or
explicitly.
```toml
[config.endpoints.default]
type = "llm"
url = "https://api.openai.com"
model = "gpt-4o-mini"
api_key = "sk-..."
pricing = { input_per_1m_tokens = 0.15, output_per_1m_tokens = 0.60 }
default_for = ["targeting", "assertion"]
[config.endpoints.vision]
type = "llm"
url = "https://api.openai.com"
model = "gpt-4o"
api_key = "sk-..."
pricing = { input_per_1m_tokens = 2.50, output_per_1m_tokens = 10.00 }
default_for = []
[config.endpoints.db_mcp]
type = "mcp"
command = "npx"
args = ["-y", "@modelcontextprotocol/server-postgres", "postgresql://localhost/mydb"]
pricing = { per_call = 0.001 }
[config.endpoints.audit_agent]
type = "a2a"
url = "http://localhost:9090"
pricing = { per_call = 0.01 }
[[test]]
name = "Dashboard with vision"
[[test.steps]]
kind = "navigate"
url = "/dashboard"
# Use the vision endpoint just for this assertion
[[test.steps]]
kind = "assert"
preset = "no_error_on_page"
endpoint = "vision"
```
**Endpoint types:**
- `llm` — LLM chat API. Defaults to the OpenAI-compatible chat completions
endpoint; `provider = "azure"` (Azure OpenAI) and `provider = "bedrock"`
(AWS Bedrock, [see below](#llm-providers-azure--aws-bedrock)) are
supported. Pricing is per-token (`input_per_1m_tokens`,
`output_per_1m_tokens`).
- `mcp` — [Model Context Protocol](https://modelcontextprotocol.io) server.
Launched as a subprocess via `command` + `args`. Pricing is `per_call`.
- `a2a` — [Agent-to-Agent Protocol](https://a2aprotocol.org) agent. Communicates
via JSON-RPC over HTTP at the given `url`. Pricing is `per_call`.
**Routing:**
- `default_for` lists which task types an endpoint serves automatically
(`targeting` for element resolution, `assertion` for assertions) — this is
the per-task model specification: give each task type its own endpoint
(e.g. a cheap model for targeting, a stronger one for assertions) by
splitting `default_for` across endpoints.
- Add `endpoint = "name"` on any step or `[[test]]` group to override routing.
**Retries + fallback chains:**
- `max_attempts` (per endpoint, default 3; global env override
`HARNESS_LLM_CALL_ATTEMPTS`) — how often a single chat completion is
retried on transient failures before the endpoint is considered failed.
- `fallbacks = ["other_endpoint", ...]` (LLM endpoints only) — an ordered
chain: when an endpoint exhausts its attempts, the next fallback is tried,
and so on, until one answers. The answering endpoint is the one charged
(per-usage cost reporting attributes the call correctly). Practical
pattern: a cheap primary model with a more powerful, more expensive
fallback that is only billed when the primary fails.
```toml
# Cheap by default; escalate to a stronger model when the gateway is
# down or returns garbage. Both endpoints serve both task types.
[config.endpoints.default]
type = "llm"
url = "$LLM_URL" # home gateway / cheap model
model = "deepseek-v3"
default_for = ["targeting", "assertion"]
max_attempts = 5 # be patient with the local gateway
fallbacks = ["pro"] # escalate only after 5 attempts
[config.endpoints.pro]
type = "llm"
url = "https://api.openai.com"
model = "gpt-4.1"
api_key = "sk-..."
pricing = { input_per_1m_tokens = 2.00, output_per_1m_tokens = 8.00 }
default_for = []
```
Chains are cycle-guarded and deduplicated; non-LLM endpoints in a
`fallbacks` list are skipped. When all endpoints fail, the error message
names every endpoint and its failure.
## LLM providers: Azure & AWS Bedrock
Besides the default OpenAI-compatible API, LLM endpoints can target Azure
OpenAI and AWS Bedrock.
### Azure OpenAI
```toml
[config.endpoints.azure]
type = "llm"
provider = "azure"
url = "https://my-resource.openai.azure.com" # resource endpoint, no path
deployment = "gpt-4o" # defaults to `model` when unset
api_version = "2024-10-21" # default when unset
api_key = "..." # sent as the `api-key` header
model = "gpt-4o" # unused by Azure; kept for pricing model names
pricing = { input_per_1m_tokens = 2.50, output_per_1m_tokens = 10.00 }
default_for = ["targeting", "assertion"]
```
Requests go to
`<url>/openai/deployments/<deployment>/chat/completions?api-version=<v>`.
With the default `api-key` auth mode the key is sent in the `api-key` header
(Azure's classic convention). Use `auth.api_key_header` to send an API key in
any custom header on any provider.
Instead of a static key you can authenticate with Entra ID — client
credentials or a managed identity (see
[LLM authentication](#llm-authentication) below).
### AWS Bedrock
```toml
[config.endpoints.bedrock]
type = "llm"
provider = "bedrock"
model = "anthropic.claude-3-5-sonnet-20241022-v2:0"
region = "eu-central-1" # optional: overrides the chain default
profile = "staging" # optional: named profile from ~/.aws
pricing = { input_per_1m_tokens = 3.00, output_per_1m_tokens = 15.00 }
default_for = ["targeting", "assertion"]
# Optional: explicit credentials instead of the credential chain
# [config.endpoints.bedrock.aws]
# access_key_id = "AKIA..."
# secret_access_key = "..."
# session_token = "..." # only for temporary credentials
```
Requests are signed with SigV4 and sent to
`https://bedrock-runtime.<region>.amazonaws.com/model/<model>/converse`; the
system prompt, `temperature`, `maxTokens`, and vision screenshots map to the
Converse API (images become `image` content blocks). When no credentials are
configured, the standard AWS credential chain is used — `AWS_ACCESS_KEY_ID` /
`AWS_SECRET_ACCESS_KEY` / `AWS_PROFILE` env vars, `~/.aws/config` and
`~/.aws/credentials`, SSO, ECS and EC2 IMDS — exactly like the AWS CLI. The
resolved credentials (and region) are cached per endpoint config for the
process lifetime. Requires building with the `aws` cargo feature
(`cargo run --features aws ...`); the Docker image and release binaries
include it.
## A2A agents
Call remote A2A agents in your test scenarios as steps, or use them inside
assertion definitions for reusable agent-backed checks.
### Agent step
```toml
[config.endpoints.audit_bot]
type = "a2a"
url = "http://localhost:9090"
pricing = { per_call = 0.01 }
[[test]]
name = "Audit trail check"
steps = [
{ kind = "navigate", url = "/admin/audit" },
{ kind = "agent", agent = "audit_bot", task = "Check if user 'admin' appears in the recent audit log" },
]
```
### Agent-backed assertions
Define reusable agent assertions with `task_template`:
```toml
[[definitions]]
name = "audit_verify"
agent = "audit_bot"
task_template = "Verify that {expected_text} is true for the page at {url}"
[[test.steps]]
kind = "assert"
definition = "audit_verify"
assert_text = "the user can see the dashboard"
```
Template variables available: `{url}`, `{title}`, `{content}`, `{expected_text}`,
`{description}`, `{task}`.
## Run as an A2A agent
Enable the `a2a-server` feature to expose the framework as an A2A agent that
other agents or orchestrators can call. The server listens on a port and accepts
`tasks/send` JSON-RPC requests.
```toml
[config.a2a_server]
enabled = true
port = 3100
```
```bash
# Build and run with the a2a-server feature
cargo run --features a2a-server -- run scenario.toml --agent-port 3100
```
Or via CLI without modifying the TOML:
```bash
llm-browser-testkit run scenario.toml --agent-port 3100
```
### Docker deployment
```bash
docker build -t llm-browser-testkit .
docker run --rm \
-p 3100:3100 \
-v "$(pwd)/scenario.toml:/scenario.toml:ro" \
-e HARNESS_LLM_TEST_URL=https://api.openai.com \
-e HARNESS_LLM_TEST_MODEL=gpt-4o-mini \
-e HARNESS_LLM_API_KEY=sk-... \
llm-browser-testkit run /scenario.toml --agent-port 3100
```
A `Dockerfile` is included in the repository — it uses a multi-stage build with
Alpine and Chromium. The image is built with `--all-features`, so all LLM
providers (including Azure and AWS Bedrock) are available out of the box.
The published image also ships the `aws` and `az` CLIs (useful for
`auth.mode = "token-command"`, e.g. `az account get-access-token`). Leaner
variants can be built locally — both are opt-in build args, off by default:
```bash
docker build -t llm-browser-testkit:aws --build-arg ENABLE_AWS_CLI=true .
docker build -t llm-browser-testkit:azure --build-arg ENABLE_AZURE_CLI=true .
```
## MCP tools
Call MCP server tools directly from test steps to query databases, read files,
or invoke any tool an MCP server exposes.
```toml
[config.endpoints.db]
type = "mcp"
command = "npx"
args = ["-y", "@modelcontextprotocol/server-postgres", "postgresql://localhost/mydb"]
pricing = { per_call = 0.001 }
[[test]]
name = "Database smoke test"
steps = [
{ kind = "navigate", url = "/dashboard" },
{ kind = "mcp", server = "db", tool = "query", args = { sql = "SELECT count(*) FROM users" } },
{ kind = "assert", preset = "no_error_on_page" },
]
```
MCP servers are launched as subprocesses via the configured `command` and `args`.
The framework handles the MCP initialize handshake, tool listing, and invocation
automatically.
## MCP server exposure
Enable the `mcp-server` feature to expose the framework as an MCP server so
other tools can invoke it remotely.
```toml
[config.mcp_server]
enabled = true
port = 3000
```
```bash
cargo run --features mcp-server -- run scenario.toml
```
When enabled, other MCP clients can call tools like `run_scenario` and
`get_page_state` on port 3000.
## Cost tracking & budgets
Every LLM call, agent invocation, and MCP tool call is tracked. After the run
completes, a cost report is printed with per-test and per-endpoint breakdowns.
### Per-test default budget
```toml
[config.budgets.per_test_default]
max_cost = 1.0
max_tokens = 100_000
max_calls = 50
enforcement = "hard"
```
### Per-test override
```toml
[[test]]
name = "Expensive test"
budget = { max_cost = 2.0, max_tokens = 200_000, enforcement = "soft" }
```
### Global budget
```toml
[config.budgets.global]
max_cost = 5.0
max_tokens = 500_000
enforcement = "hard"
```
### Enforcement
| `hard` | Abort the test or run immediately when budget is exceeded |
| `soft` | Print a warning but continue executing remaining steps |
### CLI budgets
```bash
llm-browser-testkit run scenario.toml --max-cost 10.0 --max-tokens 1000000 --budget-enforcement soft
```
### Sample report output
```
-------------------------------
COST REPORT
-------------------------------
Test: "Dashboard smoke" — $0.0891 | 4,567 tokens | 6 calls
endpoint.vision: 2 calls, 3,000 tokens, $0.0450
endpoint.default: 3 calls, 1,567 tokens, $0.0441
endpoint.audit_bot: 1 call, 0 tokens, $0.0000
-------------------------------
GLOBAL SUMMARY
Total cost: $0.1014
Total tokens: 5,801
Total calls: 10
-------------------------------
```
## How it works
Four pieces:
1. **Chrome** — launched via the [Chrome DevTools Protocol](https://chromedevtools.github.io/devtools-protocol/)
(`headless_chrome` crate). It navigates, clicks, types, and extracts page
content.
2. **LLM** — any OpenAI-compatible API. Used in two places:
- **Element targeting**: when a step says `target = "the login button"`, the
runner sends the page's interactive elements to the LLM and asks for a CSS
selector.
- **Assertions**: the runner sends page content to the LLM with a QA prompt
and expects `PASS` or `FAIL: <reason>`.
3. **A2A + MCP** — connect to remote agents via the Agent-to-Agent Protocol
and to MCP servers for tool-calling. Both are first-class step kinds.
4. **TOML scenarios** — declarative test files. No code, no CSS selectors
required. Just describe what you want in English.
```
TOML file → CLI runner → Chrome (CDP) → LLM API
→ A2A agent
→ MCP server
```
## Use as a library
```toml
[dependencies]
llm-browser-testkit = { version = "0.1", features = ["macros", "mcp-server"] }
```
```rust
use llm_browser_testkit::runner::ScenarioRunner;
use llm_browser_testkit::scenario::Scenario;
let scenario: Scenario = toml::from_str(&contents)?;
let runner = ScenarioRunner::new(scenario.config.clone(), scenario.definitions);
let report = runner.run(&scenario.test)?;
println!("Passed: {}, Failed: {}", report.tests_passed, report.tests_failed);
// Access cost/usage data
let usage = runner.usage_tracker();
let global = usage.global_snapshot();
println!("Total cost: ${:.4}", global.total_cost);
// Print the cost report
llm_browser_testkit::reporting::print_report(
&usage.per_test_snapshots(),
&global,
);
```
### Macros: `#[browser_test]` in `cargo test`
Enable the `macros` feature to write browser tests directly in your Rust test
modules:
```toml
[dev-dependencies]
llm-browser-testkit = { version = "0.1", features = ["macros"] }
```
```rust,ignore
use llm_browser_testkit::browser_test;
use llm_browser_testkit::browser_test_inline;
// Run a TOML scenario file
browser_test!(homepage => "tests/homepage.toml");
// Inline small scenarios
browser_test_inline!(hello, r#"
[config]
base_url = "https://example.com"
[[test]]
name = "hello"
[[test.steps]]
kind = "navigate"
url = "/"
[[test.steps]]
kind = "assert"
preset = "no_error_on_page"
"#);
```
Tests auto-skip when no LLM endpoint or Chrome is available — safe to include in
every CI run. They only execute with real `PASS`/`FAIL` when infrastructure is
present.
## LLM authentication
The runner supports API keys, Entra ID tokens, token commands, and custom
(static or command-produced) headers — per endpoint.
**API key** (default `api-key` mode): sent as `Authorization: Bearer <key>`
on OpenAI-compatible endpoints, as the `api-key` header on Azure. Set
`auth.api_key_header` to use a different header name on any provider:
```toml
[config.endpoints.api]
type = "llm"
url = "https://api.openai.com"
api_key = "sk-..."
model = "gpt-4o"
[config.endpoints.api.auth]
mode = "api-key" # default; can be omitted
api_key_header = "X-Api-Key" # send the key here instead of Authorization
```
**Custom headers** — static `headers` and per-call `header_commands`
(provider-agnostic; the command's first stdout line becomes the header value):
```toml
[config.endpoints.prod]
type = "llm"
url = "https://api.example.com"
model = "gpt-4o"
headers = { "X-Org-ID" = "acme" } # static
[config.endpoints.prod.header_commands]
X-M2M-Token = "kubectl exec tokenizer -- token" # dynamic, per call
```
**Token commands** (`mode = "token-command"`): run any program, its stdout
(first line) becomes the bearer token. Works with any CLI that prints a
token — Azure CLI, Vault, ...:
```toml
[config.endpoints.azure]
type = "llm"
provider = "azure"
url = "https://my-resource.openai.azure.com"
model = "gpt-4o"
[config.endpoints.azure.auth]
mode = "token-command"
token_command = "az account get-access-token --resource https://cognitiveservices.azure.com --query accessToken -o tsv"
cache_ttl_secs = 240 # reuse the token this long (default 300)
```
**Entra ID client credentials** (`mode = "entra-client-credentials"`): the
OAuth 2.0 client-credentials grant against
`https://login.microsoftonline.com/<tenant>/oauth2/v2.0/token`; the token is
cached until its server-issued expiry, then refreshed automatically:
```toml
[config.endpoints.azure.auth]
mode = "entra-client-credentials"
tenant_id = "<tenant-or-uuid>"
client_id = "<app-registration-client-id>"
client_secret = "<client-secret>"
scope = "https://cognitiveservices.azure.com/.default" # default
```
**Entra ID managed identity** (`mode = "entra-managed-identity"`): fetches a
token from the Azure IMDS endpoint — zero credentials in the config; works on
Azure VMs, App Service, and ACI with a system-assigned identity:
```toml
[config.endpoints.azure.auth]
mode = "entra-managed-identity"
```
Token acquisition is cached process-wide per auth configuration, so a test
run authenticates once instead of on every LLM call.
Via CLI / env, the runner still supports plain keys and headers:
```bash
llm-browser-testkit run tests.toml \
--llm-api-key sk-... \
--llm-header "X-Org-ID:acme" \
--llm-header "X-Project:qa"
```
```bash
export HARNESS_LLM_API_KEY=sk-...
export HARNESS_LLM_HEADERS='{"X-Org-ID":"acme","X-Project":"qa"}'
```
## License
Apache-2.0 OR MIT