Auth-Cloudflare ☁️
[!NOTE]
Auth Cloudflare Workers AI adds direct, OpenAI-compatible access to Cloudflare-hosted Workers AI text-generation models available to your Cloudflare account. No
custom_providerswiring, no bash URL adaptation, no stale model lists: install the plugin, export two environment variables, andhermes modeloffers the account's Workers AI chat models. One provider. Two env vars. Zero hand-rolled YAML.
Install ⚡
Cloudflare ships as a Hermes model-provider plugin (pure Python) backed by two Rust crates. No Rust toolchain is required for the common path.
As a Hermes plugin (recommended)
Terminal
Terminal
# Workers & Pages → Overview
# Account → Workers AI → Write
The account ID is operational metadata, not a secret. The API token is a secret - scope it to Account → Workers AI → Write (some dashboard versions label the same permission Workers AI → Edit) and nothing else. Do not request DNS, Workers Scripts, R2, D1, KV, Pages, Zero Trust, or account administration permissions.
[!IMPORTANT]
The picker works pure-Python (no compiled dependencies). The
auth-cloudflareexecutable backs the full command surface -hermes cloudflare doctor / catalog refresh / model inspect, catalog caching, and the conformance commands (model verify --suite smoke|tool-loop,model health).download.shinstalls that executable (checksum-verified, atomic, fail-closed) to a managed path; without it the picker still works through the in-process fallback, but the diagnostics and conformance commands are unavailable.
Binary CLI (optional but recommended)
The Rust core ships as a single executable. Install it any of these ways:
Terminal - cargo install (needs a Rust toolchain):
auth-cloudflare is the engine the plugin talks to; auth-hermes-cloudflare
is the optional utility (install, upgrade, doctor, status, link,
uninstall) that wraps the same installer.
Terminal - via the plugin installer (checksum-verified download from the
GitHub release, no Rust toolchain needed):
The binary is discovered in this order (plugin locate_auth_cloudflare_binary):
AUTH_CLOUDFLARE_BIN env → PATH → ~/.hermes/bin/auth-cloudflare → plugin
bin/ → plugin binaries/. cargo install puts it on PATH; download.sh
installs to ~/.hermes/bin/auth-cloudflare by default (override with
BINARY_DIR, or set AUTH_CLOUDFLARE_BIN for an exact path).
Terminal - verify the install:
From source
Terminal
JSON CLI contract 🧾
The auth-cloudflare binary emits machine-readable JSON for every command
that honors --format json: version, doctor, catalog get, catalog list, and policy get. This section pins the stable envelope shape that the
Hermes plugin consumes.
catalog get envelope
| Field | Type | Meaning |
|---|---|---|
schema_version |
int | Catalog schema version (1) |
source |
string | live | cache | fallback - where the records came from |
fetched_at |
string | RFC 3339 timestamp of the snapshot |
cache_status |
string | fresh | stale | none |
default_model |
string | Provider default model id (@cf/...) |
model_count |
int | Number of records in the resolved snapshot |
experimental_included |
bool | true - the fetch runs hide_experimental=false |
deprecated_included |
bool | false - the fetch runs include_deprecated=false |
models |
array of objects | One object per model (fields below) |
Each models[] object carries:
| Field | Type | Meaning |
|---|---|---|
id |
string | Model id (@cf/...) |
display_name |
string | Human-readable model name |
status |
string | Policy status (recommended, available, experimental, …) |
primary_agent_eligible |
bool | Whether the model may serve as the primary agent model |
context_tokens |
int | Context window size in tokens |
pricing_per_million |
object | input, cached_input, output per-million USD (nullable) |
capabilities |
object | chat, tools, reasoning verdicts |
catalog list returns the same envelope with models as an ordered array of
model-id strings (picker order). The other JSON commands are version
(name/version handshake), doctor (redacted status), and policy get
(current model policy).
The Problem 🔥
Hermes discovers provider catalogs at the OpenAI-standard GET …/models.
Cloudflare's Workers AI surface does not implement that endpoint - the
account-aware catalog lives at /ai/models/search and answers in the
OpenRouter format instead.
So every manual integration ends up hand-rolled: an account ID pasted into a base URL by hand, a model list maintained as YAML, embedding and safety-classifier models polluting the picker, and no reliable way to know which models actually support tool calling. The agent spends its budget maintaining configuration instead of using the models.
The plugin registers a real cloudflare provider that owns that routing: the
account ID is injected from the environment, the catalog is fetched from the
/ai/models/search endpoint, non-chat models are filtered out, and a curated
fallback catalog keeps the picker alive when the network is unavailable.
How It Works ⚙️
Pipeline
ENV VARS ──────────────► auth-hermes-cloudflare ────────► Hermes model picker
│
├─ CLOUDFLARE_ACCOUNT_ID → derived base URL (…/ai/v1)
├─ CLOUDFLARE_API_TOKEN → Bearer header only, never logged
├─ catalog discovery → GET …/ai/models/search (OpenRouter)
├─ chat filter → @cf/ chat models only, no guards
└─ offline fallback → curated catalog in the profile
Hermes decides:
• Online → live, account-aware catalog
• Offline → curated fallback, picker still works
Four steps:
- Register -
register_provider()addsauth-cloudflare-workers-aito the Hermes provider registry (aliases includecloudflare,cloudflare-workers-ai,workers-ai,cf-workers-ai, and more). - Resolve - the account ID from the environment derives every endpoint
(
base_url,models_url,verify_url) - one source of truth, shared by the Python plugin and the Rust core. - Discover - the catalog is fetched from the real search endpoint
(
format=openrouter&per_page=1000), not from a nonexistent/models. - Filter - safety classifiers (
llama-guard-3-8b) and non-chat modalities (embedding, image, audio, video) never reach the primary picker.
Architecture 🏗️
Layout
crates/auth-cloudflare/ ← Core: auth, catalog, cache (cdylib + rlib)
auth.rs ← account/token resolution → endpoint construction
catalog.rs ← ModelRecord, ModelRole, CapabilityState
cache.rs ← account-scoped cache paths
crates/auth-hermes-cloudflare/ ← Hermes integration (cdylib + rlib)
lib.rs ← re-exports core types for tool schemas/hooks
plugins/auth-hermes-cloudflare/ ← Hermes plugin (submodule → Auth-Hermes-Cloudflare)
__init__.py ← register_provider, lazy URLs, fetch_models
profiles/dev-cloudflare/ ← working Hermes profile (provider block)
skills/ ← cloudflare-* skills
.playform/ ← plan + development conversation archive
| Route | Purpose |
|---|---|
POST …/accounts/<ACCOUNT_ID>/ai/v1/chat/completions |
OpenAI-compatible inference |
GET …/accounts/<ACCOUNT_ID>/ai/models/search |
Account-aware model catalog (OpenRouter format) |
GET /client/v4/user/tokens/verify |
Token health check |
POST …/accounts/<ACCOUNT_ID>/ai/run/<model> |
Native REST inference (not used by the plugin) |
All endpoint/auth logic lives in the Rust core (crates/auth-cloudflare);
the Python plugin is a thin in-process provider that mirrors it for the
picker and wizard paths.
[!NOTE]
plugins/auth-hermes-cloudflare/is a separate repo (PlayForm/Auth-Hermes-Cloudflare), tracked here as a git submodule.
Provider Surface 🔧
| Aspect | Value |
|---|---|
| Provider name | auth-cloudflare-workers-ai |
| Aliases | cloudflare, cloudflare-ai, auth-cloudflare-workers-ai, cloudflare-workers-ai, workers-ai, cf-workers-ai, cf |
| Display name | Auth Cloudflare Workers AI |
| API mode | chat_completions |
| Auth type | api_key |
| Base URL | derived from the account ID - fixed_base_url, the setup wizard never prompts for an override |
| Health check | disabled (no /models endpoint); token verify is used instead |
| Signup | dash.cloudflare.com/profile/api-tokens |
| Default model | @cf/deepseek-ai/deepseek-v4-flash-0731 |
| Fallback catalog | 22 curated chat models compiled into the profile |
Skills 🧠
The repo ships three Hermes skills alongside the plugin:
| Skill | Purpose |
|---|---|
cloudflare-dev-workflow |
Reverse-PR git workflow - feat-dev/trunk integration, Source remote, no direct pushes |
cloudflare-operations |
Operational patterns - endpoints, catalog refresh, troubleshooting |
cloudflare-release-workflow |
Release process - version sync, Cloudflare/v* tag naming, BINARY_VERSION, download scripts |
Configuration 🎛️
Everything is driven by two environment variables - no recompile, no config
file to keep in sync. The AUTH_CLOUDFLARE_* names are the canonical ones;
the CLOUDFLARE_* names remain as the legacy Hermes-compatible aliases.
| Variable | Role | Secret |
|---|---|---|
CLOUDFLARE_ACCOUNT_ID / AUTH_CLOUDFLARE_ACCOUNT_ID |
account ID (Workers & Pages → Overview) | no |
CLOUDFLARE_API_TOKEN / AUTH_CLOUDFLARE_API_TOKEN |
API token (Account → Workers AI → Write; some dashboards label it Edit) | yes |
The token is only ever sent as a Bearer header - never logged, never echoed,
never rendered by Debug/Display (the Rust core redacts it in both).
config.yaml
model:
default: "@cf/deepseek-ai/deepseek-v4-flash-0731"
provider: cloudflare
providers:
cloudflare:
api_key_env: CLOUDFLARE_API_TOKEN
base_url: https://api.cloudflare.com/client/v4/accounts/${CLOUDFLARE_ACCOUNT_ID}/ai/v1
api_mode: chat_completions
[!TIP]
A ready-made working profile lives at
profiles/dev-cloudflare/(config.yamlwith the provider block above pluscloudflare-dev-workflowas the default skill). Copy it and set your own env vars.
Models 📊
The default is @cf/deepseek-ai/deepseek-v4-flash-0731 - DeepSeek V4 Flash:
1,310,720-token context, function calling, reasoning, multimodal; $0.44/M
input and $1.32/M output at Cloudflare's published rates. It is the
development default; GLM-5.3 Flash stays experimental until conformance
thresholds are met.
Discovery is live and account-aware: the catalog is fetched from
/ai/models/search in the OpenRouter format and filtered to chat-capable
@cf/ models. When the network is unavailable the plugin falls back to a
curated catalog compiled into the profile (22 chat models).
[!WARNING]
The OpenRouter-format catalog response omits tool-calling metadata entirely. The plugin maintains a verified capability table instead of inferring
unsupportedfrom absent metadata - capability unknown ≠ capability unsupported.
Curated fallback catalog (22 chat models)
@cf/deepseek-ai/deepseek-v4-flash-0731 @cf/moonshotai/kimi-k2.7-code
@cf/deepseek-ai/deepseek-v4-pro-0813 @cf/openai/gpt-oss-120b
@cf/openai/gpt-oss-20b @cf/zai-org/glm-5.3
@cf/qwen/qwen3.8-27b @cf/qwen/qwen3-30b-a3b-fp8
@cf/qwen/qwen2.5-coder-32b-instruct @cf/meta/llama-4-scout-17b-16e-instruct
@cf/meta/llama-3.3-70b-instruct-fp8-fast @cf/mistralai/mistral-small-3.1-24b-instruct
@cf/nvidia/nemotron-3-120b-a12b @cf/ibm-granite/granite-4.0-h-micro
@cf/zai-org/glm-4.7-flash @cf/moonshotai/kimi-k2.6
@cf/deepseek-ai/deepseek-r1-distill-qwen-32b @cf/meta/llama-3.1-8b-instruct-fp8
@cf/meta/llama-3.2-1b-instruct @cf/meta/llama-3.2-3b-instruct
@cf/meta/llama-3.2-11b-vision-instruct @cf/qwen/qwq-32b
Scope 🎯
Supported:
- Direct Cloudflare-hosted
@cf/...Workers AI text-generation models. - OpenAI-compatible Chat Completions.
- Hermes-owned tools: terminal, filesystem, browser, Git, and installed skills.
- Account-aware Workers AI catalog discovery.
Not supported in this release:
- Cloudflare AI Gateway third-party models.
- Anthropic Messages, Gemini-native, or provider-specific API protocols.
- Image, video, embedding, speech, reranking, or safety-only models as the primary Hermes agent.
- Cloudflare MCP account operations.
- Automatic model failover.
Troubleshooting ❓
| Symptom | Cause / fix |
|---|---|
could not verify this endpoint via …/ai/v1/models |
Expected - Cloudflare has no OpenAI /models endpoint. The provider disables health probing (supports_health_check=False) and discovers via /ai/models/search instead. |
Missing environment variable CLOUDFLARE_ACCOUNT_ID |
The error names the exact variable and prints the fix: export it from Workers & Pages → Overview → Account ID. |
Missing environment variable CLOUDFLARE_API_TOKEN |
Export a custom token scoped to Account → Workers AI → Write (some dashboards label it Edit). |
| Token verify returns 401 / 403 | Wrong, expired, or under-scoped token - check GET /client/v4/user/tokens/verify and re-create the token with Account → Workers AI → Edit. |
| Catalog fetch fails (API error, missing data array, timeout) | Non-fatal by design - the provider falls back to the curated 22-model catalog and the picker keeps working. |
Development 🛠️
Terminal
- CI
Check.ymlruns fmt, clippy, and tests on every push/PR. - CI
Build.ymlbuilds release executables for four targets (aarch64/x86_64 macOS + Linux) onCloudflare/v*tags and attaches them to the release - the exact assetsdownload.shfetches. - Git flow is a reverse-PR workflow:
feat-dev/trunkbranches,Sourceremote, no direct pushes - see thecloudflare-dev-workflowskill.
Relationship to Hermes Agent 🔗
Hermes intentionally keeps third-party vendor providers out of the core tree - the maintainers' preferred extension path is a standalone model-provider plugin. Cloudflare implements exactly that contract:
ProviderProfilesubclass +register_provider()- the same import side-effect pattern every bundled Hermes provider follows.fixed_base_url=True- the base URL is derived from the account ID, so the setup wizard never asks for a manual Base URL override.- Lazy URL properties - plugin discovery runs before the profile
.envis loaded, so URLs compute fromos.environat access time, never at import.
→ Model-provider plugin developer guide
Contributing 🤝
| Want to… | Start here |
|---|---|
| Report a bug | Open an issue |
| Suggest a feature | Start a discussion |
| Submit a PR | Fork & open a PR |
| Ask a question | Discussions Q&A |
No contribution is too small. First-time contributors are especially welcome.
License 📜
Released under CC0-1.0 - public domain.
Built with ❤️ by PlayForm.