dbx-tools-model-proxy
Rust proxy between OpenAI or Anthropic clients and Databricks Model Serving protocols.
Install and run the version-matched release binary through dbx:
The first invocation downloads the GitHub release asset matching the installed
@dbx-tools/cli version and host platform. Later invocations reuse the
validated executable.
The public crate can also be installed from crates.io:
Release builds also publish dbx-model-proxy as a GitHub release asset for
each configured platform.
The proxy uses aigw-openai and aigw-anthropic as protocol adapters.
OpenAI Chat Completions and Anthropic Messages requests can target either
Databricks Chat Completions or Responses. Native Responses input currently
targets Responses without conversion.
Authentication comes from dbx-tools-core. The proxy resolves the
selected Databricks profile, obtains cached or refreshed credentials, and
retries one upstream 401 after refreshing the rejected token.
dbx-tools-model discovers the workspace's serving endpoints, caches the
catalogue for five minutes, and resolves loose model values before forwarding.
For example, "model": "gpt" selects the highest-ranked deployed GPT model.
Authentication And Rate-Limit Identity
The binary creates one DatabricksClient at startup. That client owns upstream
authentication, while each incoming request supplies the principal used to
partition reactive rate-limit cooldowns. Identity selection never calls a
Databricks API:
- With
dbx model-proxy --profile PROFILE,dbx-tools-coreresolves that profile and uses its normal cached token lifecycle. U2M profiles use the Databricks CLI when available and may invoke login when renewal requires it. PAT and U2M requests key cooldowns by the profile name; the token itself is never part of the key. - M2M profiles key cooldowns by OAuth client ID. Token refreshes can replace the access token without changing the rate-limit identity.
- In a Databricks App using App SP, startup resolves
DATABRICKS_HOST,DATABRICKS_CLIENT_ID, andDATABRICKS_CLIENT_SECRET. The key is therefore the normalized App host, service-principal client ID, and resolved serving endpoint. - Trusted
x-forwarded-userandx-forwarded-emailheaders partition requests by App user. When those are absent, the proxy may decodesub,user_id,oid,client_id,azp,email, orpreferred_usernamefrom an incoming bearer JWT without verifying it. This decode is only a local rate-limit partitioning hint and never authenticates the request. - Forwarded OBO identity does not replace the startup client's upstream
credential. The standalone binary has no request-scoped AppKit context, so
automatic App startup normally resolves App SP. A host that needs true OBO
upstream calls must construct request-scoped clients and pass the request
headers through
DatabricksAuthOptions.
The resulting key is [normalized Databricks host, current principal, resolved model]. Different users, service principals, hosts, or serving endpoints never
share a cooldown. Keys are stored only in process memory, have no default count
limit, and disappear when the proxy exits.
Run
The server listens on 127.0.0.1:4000 by default. --port reads
DATABRICKS_APP_PORT when present.
Request bodies default to 25 MB through --max-request-bytes /
MAX_REQUEST_BYTES. Embedded JPEG, PNG, and WebP inputs are detected from their
bytes across OpenAI Chat, Responses, and Anthropic base64 shapes. Images larger
than 2 MB are re-encoded and resized proportionally to at most a 1,568-pixel
edge and 1.15 megapixels before protocol translation. Images at or below 2 MB
and remote image URLs are unchanged, and the proxy never fetches image URLs
itself. Set --image-resize-threshold-bytes or
IMAGE_RESIZE_THRESHOLD_BYTES to change the size threshold.
LOG_LEVEL accepts debug, info, warn, or error, case-insensitively,
and defaults to info. Pass -v or --verbose to select debug when
LOG_LEVEL is absent. An explicit LOG_LEVEL always wins. Normal completions
log only the resolved model, route or protocol pair, streaming mode, status,
and total duration. Rate limiting, exhausted retries, upstream 5xx responses,
and recoverable transport failures log at warn.
Debug completions add request and response byte counts, the immediate TCP peer, raw and calibrated token estimates, reported usage, attempt count, reservation state, queue wait, and adaptive budget without logging request bodies, credentials, identities, encrypted state, or embedded binary content. The token estimate includes model-visible JSON but excludes encrypted reasoning and compaction state, signatures, and embedded image, file, audio, and screenshot payloads. Streams emit their connection event at debug and one completion event at info or warn.
Pass-through usage comes from complete parsed SSE events while the original
chunks are forwarded unchanged. Observation is capped at 1 MB per event; a
malformed or larger event safely retains the estimate. Reported usage
reconciles process-local reservations across Chat Completions, Responses, Codex,
Anthropic translations, and embeddings. Unused output reservations are
credited immediately, and actual output is recorded when no maximum was
specified. Each workspace/model queue admits requests FIFO and wakes its head
when reconciliation frees capacity, without blocking unrelated models. After
three consistent samples outside a five-percent noise band, a bounded per-model
exponential moving ratio calibrates raw input estimates against actual usage.
Streaming Chat Completions defaults
stream_options.include_usage to true; an explicit caller value is
preserved.
Import postman/model-proxy.postman_collection.json into Postman. The
collection has separate folders for --target chat and --target responses;
restart the proxy with the folder's documented command before running it. Set
chatModel and responsesModel to endpoints available in the selected
workspace.
Supported routes:
GET /v1/modelsPOST /v1/embeddingsPOST /v1/chat/completionsPOST /v1/responsesPOST /v1/messagesGET /healthzGET /metricsGET /metrics/snapshotGET /metrics/eventsGET /metrics/prometheus
GET /v1/models reads the cached live serving-endpoint catalogue. Standard
requests receive an OpenAI object / data envelope. An originator header
whose value starts with codex receives a Codex models envelope. Both use
the identities returned by Databricks directly: OpenAI uses the serving
endpoint name, while Codex maps databricks-<model> to the gateway's
system.ai.<model> identity. No databricks/ or dbx/ namespace is added.
Requests targeting Responses default a missing truncation field to "auto";
an explicit caller value is preserved. Databricks AI Gateway accepts this
default on Open Responses for Claude and Kimi and on Codex Responses for GPT
and Kimi. Claude itself is not enabled on the Codex route.
Unfiltered responses list chat/LLM families first and embedding families
second, alphabetically sorting families within each tier. Each family uses the
same version, variant, and class preference as a search for that family.
Recognized unclassified models remain in the chat/LLM tier. Custom and
unrecognized endpoints sort by name last. Codex priorities follow the resulting
order.
Use ?search=gpt to apply the same fuzzy scoring and ordering as model
resolution. Add ?extended=true to include the score, service names,
capability class, profile, task, state, and other catalogue metadata. Extended
output defaults to false.
POST /v1/embeddings resolves the requested model only among deployed
embedding endpoints, forwards the request to that endpoint's invocations
route, and preserves the OpenAI embedding response.
--target responses forces Chat Completions or Anthropic Messages input
through the Responses request translator. --target chat sends canonical
Chat Completions. --target auto selects Responses for Responses clients,
models listed by the current Databricks Responses documentation, Codex clients,
and requests containing Responses-only hosted tools or conversation fields.
Native Responses requests are forwarded without a capability allow-list, so
Databricks-supported function, custom, apply_patch, shell,
image_generation, mcp, and web_search tools, image inputs, conversation
state, background mode, and future request fields remain intact. Chat and
Anthropic image blocks are translated to Responses input_image content.
Chat-hosted tools are preserved in Responses form instead of being rejected as
malformed function tools. A forced --target chat returns a clear client error
for Responses-only features rather than silently dropping them.
Codex model records obtain image-input, web-search, and patch capability sets from the corresponding Databricks documentation pages. The parsed model lists are cached for one day and matched against endpoint, model-service, and provider identities from the live workspace catalogue. The same parser generates a committed snapshot during repository synthesis, and the binary embeds that snapshot as its offline fallback. A failed page refresh retains the matching capabilities from the embedded snapshot without blocking model listing. This avoids embedding a handwritten model/version matrix while still using the unified local execution tool shape expected by current Codex clients.
Codex Catalogue Discovery
Codex discovery is a separate wire contract over the same route:
- a standard
GET /v1/modelsreceives the OpenAIdataenvelope; - a request carrying
originator: codex_cli_rsreceives the Codexmodelsenvelope used bycodex debug modelsand the/modelpicker.
Codex validates the complete remote catalogue before merging it with its bundled
models. One incompatible record causes it to retain the bundled catalogue.
Reasoning levels must therefore remain { effort, description } objects,
web_search_tool_type must remain a non-null enum value, and
supports_search_tool carries the independent capability flag.
The Codex envelope deliberately excludes embeddings, Claude, Gemini,
unrecognized identities, and endpoint names without the required
databricks- prefix. Those models do not become compatible merely because they
appear in the standard OpenAI envelope; adding a family requires separate
Responses and tool-replay validation.
Compare remote and bundled discovery without changing provider configuration:
|
|
Run the bounded real-client regression with an installed Codex 0.148.0:
RUN_CODEX_DISCOVERY_TESTS=1 RUSTC_WRAPPER= \
The test builds the catalogue through models_payload_with_capabilities, serves
it from a loopback-only fixture, uses synthetic authentication and an isolated
temporary CODEX_HOME, verifies the expected HTTP request, and compares
discovered slugs rather than unstable total counts.
Streaming requests use SSE without buffering the upstream response. Matching protocols pass the upstream byte stream through directly, including Chat Completions to Chat Completions and Responses to Responses for Codex clients. Cross-protocol streams pass through aigateway's stateful Chat Completions or Responses parser. Anthropic output uses aigateway's native SSE encoder, while OpenAI Chat Completions output uses the proxy's canonical event encoder.
Responses input currently targets only Responses, so Responses-to-Chat translation is outside the supported route matrix.
Databricks errors are returned with their original status, body, and content
type. The proxy also forwards Retry-After, request and correlation IDs,
rate-limit headers, quota names, and Databricks limit details.
HTTP 429 responses pause the process-local host/principal/model gate described
above. One request probes after the shared cooldown while other streaming and
non-streaming requests for the same key remain paused. The Retry-After
response header controls the delay when present, followed by the documented
Foundation Model API error.retry_after JSON value. An input-token 429 without
either waits for the local token window, or 60 seconds when process-local
history cannot explain the workspace limit. Other 429s use BackON jittered
exponential delays from one second to one minute. Every retry reacquires token
admission and owns exactly one reservation. Every 429 logs a returned
error.message, including the final attempt. The default five retries mean one
initial request plus up to five retries.
After the final attempt, the original 429 status, body, and rate-limit headers
are returned to the caller. Configure RATE_LIMIT_RETRIES,
RATE_LIMIT_INITIAL_DELAY_MS, and RATE_LIMIT_MAX_DELAY_MS, or the matching
CLI flags. Set retries to 0 to disable both retries and coordinated cooldowns.
Only an initial HTTP 429 is retried; an SSE error after streaming begins cannot
be replayed safely.
The process-local token queue reads Databricks' published Enterprise
pay-per-token ITPM and OTPM limits from the same daily documentation cache and
generated-fallback pattern used for model capabilities and retirement status.
Input and output windows are tracked separately for each resolved model and
workspace. Requests reserve a tokenx-rs input estimate plus any explicit
max_output_tokens, max_completion_tokens, or max_tokens value. Claude
Sonnet 4 reserves its documented 1,000-token default when no output limit is
present. An active queue rejects an input estimate above its complete
per-minute budget with a local structured 429 rather than clamping it.
RATE_LIMIT_MODE / --rate-limit-mode accepts auto, on, or off and
defaults to auto. Auto mode leaves each workspace/model key unthrottled until
its first 429 message containing Exceeded workspace input tokens,
case-insensitively. A matching 429 starts with a 50 percent input-budget
penalty, repeated matches add 25 percentage points up to a 90 percent penalty,
and each clean recovery step removes 10 percentage points. Output admission is
not reduced.
Recovery requires both ten minutes since the latest matching 429 and ten clean upstream 2xx responses. Later steps require another five minutes and ten clean responses. Zero penalty starts one final full-budget probation interval before the key becomes inactive again. Idle time alone never relaxes a key, and a renewed matching 429 immediately tightens or reactivates it. The state is process-local adaptive congestion control and resets when the process exits. It does not claim ownership of the complete workspace quota.
The complete configured or documented input limit remains the per-request
ceiling. The temporary penalty only reduces the rolling budget. A request above
the adaptive budget but within the complete ceiling can run as the sole
reservation in an empty window, so adaptive recovery cannot make a valid
request impossible to admit. on applies complete budgets immediately and
never decays. off never applies them. Provisioned throughput bypasses token
admission in every mode. An unknown model limit retains shared cooldown and
retry behavior without creating an active no-op limiter.
Use INPUT_TOKENS_PER_MINUTE / --input-tokens-per-minute and
OUTPUT_TOKENS_PER_MINUTE / --output-tokens-per-minute to override the
published limits. Set PROVISIONED_THROUGHPUT=true or pass
--provisioned-throughput to disable both TPM windows. QPH remains enforced by
Databricks because process-local tracking cannot coordinate a workspace across
proxy replicas. /healthz exposes aggregate process-local counters for
automatic activation, tightening, relaxation, probation, deactivation,
reactivation, admission waits, oversized rejections, post-admission input 429s,
retry reacquisition, and fallback full-window delays.
Metrics And Dashboard
Metrics use the existing proxy listener and are enabled by default. The normal
release binary defaults to --metrics=ui:
uicollects metrics and serves all metrics routes, including the dashboard;collectcollects metrics and serves JSON, SSE, and Prometheus without UI routes;offremoves metrics collection, middleware, history, and routes;trueselects the fullest mode compiled into the binary;falsealiasesoff.
Use METRICS for the same values. A non-loopback --host keeps collecting but
returns 404 from every metrics route unless --metrics-public or
METRICS_PUBLIC=true explicitly acknowledges remote exposure. Forwarded
headers do not bypass this safeguard. Put an operator-owned authenticated proxy
in front before exposing model names, traffic rates, and limit pressure.
/metrics/snapshot returns the bounded dashboard payload. Pass
?model=<resolved-model> to include the selected model's bounded history;
ordinary snapshots and SSE retain only aggregate model summaries.
/metrics/events streams five-second snapshot events with SSE.
/metrics/prometheus returns Prometheus text from the in-process recorder.
/metrics serves the static dashboard embedded in the normal release binary.
The dashboard provides 1h, 6h, and 24h ranges, model and outcome filters,
request, token, and latency line graphs, model latency and error summaries, the
adaptive rate-limit timeline, and process-wide retention status. GridStack
provides drag-and-drop placement and widget resizing without a runtime CDN or
framework server. The bounded widget geometry is stored in browser
localStorage, restored after reloads and proxy restarts, and synchronized
across open tabs. Metric history remains process-local and is never stored in
the browser. The reset-layout icon restores and persists the canonical widget
geometry.
History remains in process memory:
- five-second buckets retain the latest hour;
- one-minute rollups retain the latest 24 hours;
- at most 32 named model series are retained, with additional names combined
under
other; - aggregate history targets and caps retained data at 16 MiB;
- no request events, bodies, identities, peers, credentials, or model traffic are written to disk.
The dashboard follows the approved
Figma frame
and canonical branding/brand.yaml tokens. Assets are static and require no
runtime CDN or Node server. Validate their canonical tokens and Figma status
markers with:
Cargo Features
The default feature is metrics-ui, which includes metrics. A headless build
keeps collection and machine endpoints without dashboard assets:
A metrics-free build accepts only --metrics=false and excludes the recorder,
histograms, Prometheus exporter, embedded assets, and metrics routes:
Metrics persistence is deliberately not implemented. Process-local memory is the only supported retention layer until operational evidence justifies the separate optional SQLite phase.