# captchaforge: Production deployment guide
This is the operator's manual for running captchaforge in a long-lived
service. Library API reference is in `README.md`; module-level
architecture is in `ARCHITECTURE.md`. This document answers the
"what do I configure / monitor / on-call when X happens" questions.
## Stealth: required for production
Modern WAFs (Cloudflare, Akamai, PerimeterX/HUMAN, DataDome) detect
headless Chromium *trivially* via `navigator.webdriver === true`,
empty `navigator.plugins`, missing `window.chrome`, and headless WebGL
fingerprints. **No amount of "realistic mouse movement" recovers from
those signals.** Without stealth overrides, captchaforge's behavioural
solver against real Cloudflare interactive challenges will reliably
return `success: false`.
**Apply stealth before the first navigation:**
```rust
let page = browser.new_page("about:blank").await?;
captchaforge::stealth::apply_stealth(&page).await?;
page.goto("https://protected.example.com/login").await?;
// ... captchaforge::auto_solve(&page).await ...
```
`auto_solve()` also calls `apply_stealth` defensively, but ONLY on
the page passed to it, by then the first navigation has already
loaded and the WAF has already fingerprinted you. Always apply
stealth on the page object BEFORE any `goto()`.
What stealth covers (see `src/stealth.rs` for the full surface):
- `navigator.webdriver` returns `undefined`
- `navigator.plugins` reports a plausible PDF plugin set
- `navigator.languages` returns `["en-US", "en"]`
- `window.chrome` exposes the expected runtime/loadTimes/csi/app
- WebGL vendor/renderer report `Intel Inc. / Intel Iris OpenGL Engine`
- `Notification.permission` matches the page's actual permission state
What stealth does NOT cover, these are YOUR responsibility:
- **User-Agent string.** Set a real Chrome UA on
`chromiumoxide::BrowserConfig`. NEVER ship `HeadlessChrome/...`.
- **TLS fingerprint (JA3/JA4).** Different chromium build or a
network-layer proxy. Out of scope here.
- **IP reputation.** Use residential proxies if your traffic gets
blocked by ASN.
## Sanity-check before you ship
Before captchaforge handles a real request:
1. **Run the real-WAF bench against live sitekeys.** Proves the
production pipeline survives real Cloudflare / hCaptcha / Google
scripts, not just self-hosted fixtures:
```sh
cd bench/
cargo run --release -- --suite real-waf --instances 1
```
Pass condition: every entry in the `SUCC%` column is `100.0%` and
`SOLVER` is one of `BehavioralBypass` / `VisionLLM` / `AudioBypass` /
`ThirdPartyService`. If any row reports `FAILED` with
`"token looks fake"`, the chain reported success but didn't actually
inject a real token, likely a regression in the per-vendor JS
probe; investigate before shipping.
2. **Run the throughput suite at production concurrency.** One minute,
the same `max_concurrent` you'll see live. Throughput target is
workload-specific; a single 4090-class GPU on Ollama qwen3-vl:30b
sustains ~10 image-grid solves/min. PoW captchas are CPU-bound and
sustain hundreds/min.
3. **Set the cache TTL to match upstream token lifetime.** Default is
60s; bump to the vendor's published token-validity window if you
have real-domain traffic. A high `expired_misses : hits` ratio in
`TokenCache::stats()` means TTL is shorter than the typical
re-visit window (bump it).
## Configuration layers
Three layers, lightest → heaviest. Higher layers beat lower ones.
### 1. Compiled defaults
Whatever ships in `SolveConfig::default()` / `ChainConfig::default()` /
`TokenCache::new()`. The library is usable with zero config; defaults
target a single-process headless-browser deployment.
### 2. Tier-A operational config (`.captchaforge.toml`)
The recommended layer for production. Drop a config file and the
discovery search path is walked:
- `$CAPTCHAFORGE_CONFIG` (explicit path)
- `./.captchaforge.toml`
- `./captchaforge.toml`
- `$XDG_CONFIG_HOME/captchaforge/config.toml` (or `~/.config/...`)
See `.captchaforge.toml.example` for the schema. Unknown keys are a
parse error so typos surface instead of being silently ignored.
**Don't put secrets in this file.** The third-party API key is the
only secret captchaforge handles, and `Config::load_from_path` emits
a `tracing::warn!` if it sees `[third_party] api_key` set inline.
Use the `CAPTCHAFORGE_THIRDPARTY_API_KEY` env var or a secrets-store
sidecar instead.
### 3. Env vars (per-process overrides)
| `CAPTCHAFORGE_CONFIG` | (unset) | Explicit Tier-A config path |
| `CAPTCHAFORGE_VLM_ENDPOINT` | `http://localhost:11434` | Ollama base URL |
| `CAPTCHAFORGE_VLM_MODEL` | `qwen3-vl:30b` | VLM model name |
| `CAPTCHAFORGE_THIRDPARTY_API_KEY` | (unset) | 2captcha-compatible key |
| `CAPTCHAFORGE_THIRDPARTY_ENDPOINT` | (per service) | Override service base URL |
Builder methods on the solver structs always beat env, which always
beats Tier-A, which always beats compiled defaults.
## Telemetry: wire it on day one
The default chain ships with `NoopTelemetry`. Production deployments
**must** install a real `SolverTelemetry` impl before going live
without one you have no visibility into success rates, p95 latency,
or which solvers are actually winning.
### Recommended sinks (drop-in)
- **Prometheus**: implement `SolverTelemetry::record` to push
`captchaforge_solve_duration_seconds{outcome,solver,kind}` and
`captchaforge_solve_total{outcome,solver}` counters. Cache hits
fire as `solver: "TokenCache"` so dashboards distinguish solved-
from-scratch vs cache-served.
- **OpenTelemetry**: same shape, attach `SpanKind::Client` for the
third-party HTTP span, `SpanKind::Internal` for behavioral/VLM/
audio. Wire to your existing exporter.
- **Sentry-as-error-sink**: implement `record` to dispatch
`Outcome::Error` events to `sentry::capture_message`. Don't
forward `Outcome::Failure`: that's expected business-as-usual
(the chain falls through to the next solver).
- **SQLite-tail**: append-only event log to a local SQLite for offline
audit. Useful for the "why did this user see a captcha twice" debug
question 3 weeks later.
### What to alert on
- `Outcome::Error` rate > 1% over 10 min → the chain is hitting
unexpected exceptions. Likely chromium crashed or VLM endpoint went
down.
- `Outcome::Timeout` rate > 5% → `per_solver_timeout_ms` is too low
OR upstream is slow. Check vendor status pages.
- Cache `hit_rate` flips from steady-state to ~0 → a TTL setting
changed, or domain pattern shifted (e.g. new path includes
per-request token).
- `Success` count flat-lines but `Failure` count climbs → vendor
changed their detection logic. Time to look at `PatternStore` to
see if the EMA flipped winners; if a method has gone from
`success_rate: 1.0` to `0.1` over the last 10 samples, that's the
signal.
## Token cache, what to expect
`TokenCache` is per-process, in-memory, `Mutex<HashMap>`-backed. This
is correct for **single-process services** with steady traffic to a
small set of domains.
Worth knowing:
- **Restart loses state.** If you restart workers more than once per
minute, the cache is effectively never hot. Either bump TTL +
reduce restarts, or front captchaforge with a sticky load balancer
so the same worker handles the same domain.
- **No bounded size.** `TokenCache` will grow without limit if your
domain set grows without limit. For services that touch arbitrary
user-supplied URLs (e.g. crawlers), wrap `TokenCache` with your own
LRU eviction or call `clear()` periodically. Keep an eye on
`stats().puts` minus `stats().invalidations`.
- **Not shared across processes.** Multi-worker deployments with the
default cache will each maintain their own, duplicate solves on
cache misses. For a shared cache, write a thin wrapper around the
`TokenCache` API that proxies to Redis / memcached / your existing
cache layer; the public methods (`get`, `put_full`, `invalidate`)
are the contract.
## Pattern store, when to persist
`PatternStore` is per-process by default. For a long-running service:
- **Persist on shutdown.** `chain.with_pattern_store(store)` and call
`store.save_to_path(...)` from your shutdown handler. On next start,
`store.load_from_path(...)` rehydrates the learned routing.
- **Cross-instance share.** Workers can `merge` each other's patterns:
drop store JSON to a shared S3 path every N minutes, have peers
pull + merge. Higher `sample_count` wins on conflict.
- **EMA alpha.** The store uses bounded-α EMA (α=0.2 after 10
samples). Vendor migrations register within ~10 fresh observations.
This is intentional, don't change it without measuring the
vendor-flip-detection bench.
## Third-party adapter, operational notes
`ThirdPartyCaptchaSolver` talks the 2captcha-compatible HTTP protocol
(2captcha / CapMonster / CapSolver share it). Production gotchas:
- **Rate limits are vendor-side.** captchaforge does no client-side
rate limiting. If you submit 100 captchas/sec to 2captcha you'll
see `ERROR_NO_SLOT_AVAILABLE`. The retry layer treats this as
terminal and won't burn polling slots, you'll get an immediate
fail-through to the next solver. Add your own rate limiter upstream
if your service generates bursts.
- **Balance monitoring.** `ERROR_ZERO_BALANCE` is terminal. Expose a
metric on `Outcome::Error` events tagged `terminal=true` and alert
before balance hits zero.
- **Service fallback isn't built-in.** If 2captcha goes down,
`ThirdPartyCaptchaSolver` doesn't auto-fail over to CapMonster.
Build that orchestration above the chain, instantiate two
third-party solvers, install both, let the chain race them.
## Browser lifecycle
captchaforge doesn't manage chromium, your service does (via
`chromiumoxide::Browser`). Operational guidance:
- **One browser per worker, many pages per browser.** Cheaper than
spawning a browser per request. The bench harness uses this shape
(`BrowserPool`); copy it.
- **Recycle the browser every N solves OR every M minutes.**
Long-lived chromium accumulates state (cookies, localStorage,
shaders, GPU surfaces) and eventually OOMs. Default in production:
recycle after 100 solves or 1 hour, whichever first.
- **Crash recovery.** If `chromiumoxide::Page::evaluate` returns an
error mentioning a closed CDP connection, the browser died. Catch
and respawn, captchaforge surfaces the error cleanly via the
outer `Result`, but doesn't restart the browser for you.
## Security
- **Forbid `unsafe_code`.** The crate already does (`#![forbid(unsafe_code)]`).
- **API key in env, not files.** See "Tier-A config" above. The
`Config::warn_on_inline_secrets` warning isn't a hard error; treat
it as one.
- **Don't log raw tokens at INFO.** Captcha tokens are bearer tokens
for the next request. Telemetry should record `token_present: true`,
not the token itself. If you absolutely must log the token (debug
scenarios), log to a separate stream that gets shorter retention.
- **Cookies cached alongside tokens.** `TokenCache` stores
`CapturedCookie` data. Same retention story as tokens, at-rest
encryption matters if you persist this anywhere.
## Upgrade procedure
Pre-1.0 releases may break public API. To upgrade safely:
1. Read the `CHANGELOG.md` entry for your target version.
2. Bump in a feature branch; run `cargo test --workspace` AND the
real-WAF bench. If both pass, the vendor probes still work.
3. Watch `Success` rate in production for the first hour after
deploy. If it dips noticeably, the vendor changed their DOM and
the upgrade matters; otherwise you're fine.
## Known limitations
- **No HTTP/2 server-push handling.** If a vendor pushes resources
out-of-band that the captcha widget needs, the page won't see them
unless chromium handles it natively (which it does for the major
vendors).
- **No mobile user-agent profile.** Defaults look like a desktop
Chrome. If you're solving captchas on a mobile-only flow, set the
UA via your `chromiumoxide::Browser` config.
- **No proxy support layer.** captchaforge doesn't manage proxies.
If you need them, configure `chromiumoxide::Browser` with the
proxy at launch, captchaforge inherits whatever the page's
network stack uses.
## When in doubt
Open an issue at <https://github.com/santhreal/captchaforge/issues>
with a reproducer, the relevant `journalctl` / log output, and the
output of `captchaforge --version`. Don't paste API keys.