captchaforge 0.2.39

Captcha detection and solving for Firefox and BiDi-driven browsers. Detection, vendor solver scaffolding, trusted cross-origin click delivery into nested OOPIFs, and stealth personas are implemented and tested; broad live-vendor solve rates are not yet benchmarked.
Documentation
# captchaforge: Production deployment guide

This is the operator's manual for running captchaforge in a long-lived
service. Library API reference is in `README.md`; module-level
architecture is in `ARCHITECTURE.md`. This document answers the
"what do I configure / monitor / on-call when X happens" questions.

## Stealth: required for production

Modern WAFs (Cloudflare, Akamai, PerimeterX/HUMAN, DataDome) detect
headless Chromium *trivially* via `navigator.webdriver === true`,
empty `navigator.plugins`, missing `window.chrome`, and headless WebGL
fingerprints. **No amount of "realistic mouse movement" recovers from
those signals.** Without stealth overrides, captchaforge's behavioural
solver against real Cloudflare interactive challenges will reliably
return `success: false`.

**Apply stealth before the first navigation:**

```rust
let page = browser.new_page("about:blank").await?;
captchaforge::stealth::apply_stealth(&page).await?;
page.goto("https://protected.example.com/login").await?;
// ... captchaforge::auto_solve(&page).await ...
```

`auto_solve()` also calls `apply_stealth` defensively, but ONLY on
the page passed to it, by then the first navigation has already
loaded and the WAF has already fingerprinted you. Always apply
stealth on the page object BEFORE any `goto()`.

What stealth covers (see `src/stealth.rs` for the full surface):
- `navigator.webdriver` returns `undefined`
- `navigator.plugins` reports a plausible PDF plugin set
- `navigator.languages` returns `["en-US", "en"]`
- `window.chrome` exposes the expected runtime/loadTimes/csi/app
- WebGL vendor/renderer report `Intel Inc. / Intel Iris OpenGL Engine`
- `Notification.permission` matches the page's actual permission state

What stealth does NOT cover, these are YOUR responsibility:
- **User-Agent string.** Set a real Chrome UA on
  `chromiumoxide::BrowserConfig`. NEVER ship `HeadlessChrome/...`.
- **TLS fingerprint (JA3/JA4).** Different chromium build or a
  network-layer proxy. Out of scope here.
- **IP reputation.** Use residential proxies if your traffic gets
  blocked by ASN.

## Sanity-check before you ship

Before captchaforge handles a real request:

1. **Run the real-WAF bench against live sitekeys.** Proves the
   production pipeline survives real Cloudflare / hCaptcha / Google
   scripts, not just self-hosted fixtures:

   ```sh
   cd bench/
   cargo run --release -- --suite real-waf --instances 1
   ```

   Pass condition: every entry in the `SUCC%` column is `100.0%` and
   `SOLVER` is one of `BehavioralBypass` / `VisionLLM` / `AudioBypass` /
   `ThirdPartyService`. If any row reports `FAILED` with
   `"token looks fake"`, the chain reported success but didn't actually
   inject a real token, likely a regression in the per-vendor JS
   probe; investigate before shipping.

2. **Run the throughput suite at production concurrency.** One minute,
   the same `max_concurrent` you'll see live. Throughput target is
   workload-specific; a single 4090-class GPU on Ollama qwen3-vl:30b
   sustains ~10 image-grid solves/min. PoW captchas are CPU-bound and
   sustain hundreds/min.

3. **Set the cache TTL to match upstream token lifetime.** Default is
   60s; bump to the vendor's published token-validity window if you
   have real-domain traffic. A high `expired_misses : hits` ratio in
   `TokenCache::stats()` means TTL is shorter than the typical
   re-visit window (bump it).

## Configuration layers

Three layers, lightest → heaviest. Higher layers beat lower ones.

### 1. Compiled defaults

Whatever ships in `SolveConfig::default()` / `ChainConfig::default()` /
`TokenCache::new()`. The library is usable with zero config; defaults
target a single-process headless-browser deployment.

### 2. Tier-A operational config (`.captchaforge.toml`)

The recommended layer for production. Drop a config file and the
discovery search path is walked:

- `$CAPTCHAFORGE_CONFIG` (explicit path)
- `./.captchaforge.toml`
- `./captchaforge.toml`
- `$XDG_CONFIG_HOME/captchaforge/config.toml` (or `~/.config/...`)

See `.captchaforge.toml.example` for the schema. Unknown keys are a
parse error so typos surface instead of being silently ignored.

**Don't put secrets in this file.** The third-party API key is the
only secret captchaforge handles, and `Config::load_from_path` emits
a `tracing::warn!` if it sees `[third_party] api_key` set inline.
Use the `CAPTCHAFORGE_THIRDPARTY_API_KEY` env var or a secrets-store
sidecar instead.

### 3. Env vars (per-process overrides)

| Variable | Default | Purpose |
|---|---|---|
| `CAPTCHAFORGE_CONFIG` | (unset) | Explicit Tier-A config path |
| `CAPTCHAFORGE_VLM_ENDPOINT` | `http://localhost:11434` | Ollama base URL |
| `CAPTCHAFORGE_VLM_MODEL` | `qwen3-vl:30b` | VLM model name |
| `CAPTCHAFORGE_THIRDPARTY_API_KEY` | (unset) | 2captcha-compatible key |
| `CAPTCHAFORGE_THIRDPARTY_ENDPOINT` | (per service) | Override service base URL |

Builder methods on the solver structs always beat env, which always
beats Tier-A, which always beats compiled defaults.

## Telemetry: wire it on day one

The default chain ships with `NoopTelemetry`. Production deployments
**must** install a real `SolverTelemetry` impl before going live 
without one you have no visibility into success rates, p95 latency,
or which solvers are actually winning.

### Recommended sinks (drop-in)

- **Prometheus**: implement `SolverTelemetry::record` to push
  `captchaforge_solve_duration_seconds{outcome,solver,kind}` and
  `captchaforge_solve_total{outcome,solver}` counters. Cache hits
  fire as `solver: "TokenCache"` so dashboards distinguish solved-
  from-scratch vs cache-served.

- **OpenTelemetry**: same shape, attach `SpanKind::Client` for the
  third-party HTTP span, `SpanKind::Internal` for behavioral/VLM/
  audio. Wire to your existing exporter.

- **Sentry-as-error-sink**: implement `record` to dispatch
  `Outcome::Error` events to `sentry::capture_message`. Don't
  forward `Outcome::Failure`: that's expected business-as-usual
  (the chain falls through to the next solver).

- **SQLite-tail**: append-only event log to a local SQLite for offline
  audit. Useful for the "why did this user see a captcha twice" debug
  question 3 weeks later.

### What to alert on

- `Outcome::Error` rate > 1% over 10 min → the chain is hitting
  unexpected exceptions. Likely chromium crashed or VLM endpoint went
  down.
- `Outcome::Timeout` rate > 5% → `per_solver_timeout_ms` is too low
  OR upstream is slow. Check vendor status pages.
- Cache `hit_rate` flips from steady-state to ~0 → a TTL setting
  changed, or domain pattern shifted (e.g. new path includes
  per-request token).
- `Success` count flat-lines but `Failure` count climbs → vendor
  changed their detection logic. Time to look at `PatternStore` to
  see if the EMA flipped winners; if a method has gone from
  `success_rate: 1.0` to `0.1` over the last 10 samples, that's the
  signal.

## Token cache, what to expect

`TokenCache` is per-process, in-memory, `Mutex<HashMap>`-backed. This
is correct for **single-process services** with steady traffic to a
small set of domains.

Worth knowing:

- **Restart loses state.** If you restart workers more than once per
  minute, the cache is effectively never hot. Either bump TTL +
  reduce restarts, or front captchaforge with a sticky load balancer
  so the same worker handles the same domain.
- **No bounded size.** `TokenCache` will grow without limit if your
  domain set grows without limit. For services that touch arbitrary
  user-supplied URLs (e.g. crawlers), wrap `TokenCache` with your own
  LRU eviction or call `clear()` periodically. Keep an eye on
  `stats().puts` minus `stats().invalidations`.
- **Not shared across processes.** Multi-worker deployments with the
  default cache will each maintain their own, duplicate solves on
  cache misses. For a shared cache, write a thin wrapper around the
  `TokenCache` API that proxies to Redis / memcached / your existing
  cache layer; the public methods (`get`, `put_full`, `invalidate`)
  are the contract.

## Pattern store, when to persist

`PatternStore` is per-process by default. For a long-running service:

- **Persist on shutdown.** `chain.with_pattern_store(store)` and call
  `store.save_to_path(...)` from your shutdown handler. On next start,
  `store.load_from_path(...)` rehydrates the learned routing.
- **Cross-instance share.** Workers can `merge` each other's patterns:
  drop store JSON to a shared S3 path every N minutes, have peers
  pull + merge. Higher `sample_count` wins on conflict.
- **EMA alpha.** The store uses bounded-α EMA (α=0.2 after 10
  samples). Vendor migrations register within ~10 fresh observations.
  This is intentional, don't change it without measuring the
  vendor-flip-detection bench.

## Third-party adapter, operational notes

`ThirdPartyCaptchaSolver` talks the 2captcha-compatible HTTP protocol
(2captcha / CapMonster / CapSolver share it). Production gotchas:

- **Rate limits are vendor-side.** captchaforge does no client-side
  rate limiting. If you submit 100 captchas/sec to 2captcha you'll
  see `ERROR_NO_SLOT_AVAILABLE`. The retry layer treats this as
  terminal and won't burn polling slots, you'll get an immediate
  fail-through to the next solver. Add your own rate limiter upstream
  if your service generates bursts.
- **Balance monitoring.** `ERROR_ZERO_BALANCE` is terminal. Expose a
  metric on `Outcome::Error` events tagged `terminal=true` and alert
  before balance hits zero.
- **Service fallback isn't built-in.** If 2captcha goes down,
  `ThirdPartyCaptchaSolver` doesn't auto-fail over to CapMonster.
  Build that orchestration above the chain, instantiate two
  third-party solvers, install both, let the chain race them.

## Browser lifecycle

captchaforge doesn't manage chromium, your service does (via
`chromiumoxide::Browser`). Operational guidance:

- **One browser per worker, many pages per browser.** Cheaper than
  spawning a browser per request. The bench harness uses this shape
  (`BrowserPool`); copy it.
- **Recycle the browser every N solves OR every M minutes.**
  Long-lived chromium accumulates state (cookies, localStorage,
  shaders, GPU surfaces) and eventually OOMs. Default in production:
  recycle after 100 solves or 1 hour, whichever first.
- **Crash recovery.** If `chromiumoxide::Page::evaluate` returns an
  error mentioning a closed CDP connection, the browser died. Catch
  and respawn, captchaforge surfaces the error cleanly via the
  outer `Result`, but doesn't restart the browser for you.

## Security

- **Forbid `unsafe_code`.** The crate already does (`#![forbid(unsafe_code)]`).
- **API key in env, not files.** See "Tier-A config" above. The
  `Config::warn_on_inline_secrets` warning isn't a hard error; treat
  it as one.
- **Don't log raw tokens at INFO.** Captcha tokens are bearer tokens
  for the next request. Telemetry should record `token_present: true`,
  not the token itself. If you absolutely must log the token (debug
  scenarios), log to a separate stream that gets shorter retention.
- **Cookies cached alongside tokens.** `TokenCache` stores
  `CapturedCookie` data. Same retention story as tokens, at-rest
  encryption matters if you persist this anywhere.

## Upgrade procedure

Pre-1.0 releases may break public API. To upgrade safely:

1. Read the `CHANGELOG.md` entry for your target version.
2. Bump in a feature branch; run `cargo test --workspace` AND the
   real-WAF bench. If both pass, the vendor probes still work.
3. Watch `Success` rate in production for the first hour after
   deploy. If it dips noticeably, the vendor changed their DOM and
   the upgrade matters; otherwise you're fine.

## Known limitations

- **No HTTP/2 server-push handling.** If a vendor pushes resources
  out-of-band that the captcha widget needs, the page won't see them
  unless chromium handles it natively (which it does for the major
  vendors).
- **No mobile user-agent profile.** Defaults look like a desktop
  Chrome. If you're solving captchas on a mobile-only flow, set the
  UA via your `chromiumoxide::Browser` config.
- **No proxy support layer.** captchaforge doesn't manage proxies.
  If you need them, configure `chromiumoxide::Browser` with the
  proxy at launch, captchaforge inherits whatever the page's
  network stack uses.

## When in doubt

Open an issue at <https://github.com/santhreal/captchaforge/issues>
with a reproducer, the relevant `journalctl` / log output, and the
output of `captchaforge --version`. Don't paste API keys.