dbx-tools-model-proxy 0.9.3

Multi-protocol Databricks model proxy
# dbx-tools-model-proxy

Rust proxy between OpenAI or Anthropic clients and Databricks Model Serving
protocols.

Install and run the version-matched release binary through `dbx`:

```sh
dbx model-proxy --profile PROFILE
```

The first invocation downloads the GitHub release asset matching the installed
`@dbx-tools/cli` version and host platform. Later invocations reuse the
validated executable.

The public crate can also be installed from crates.io:

```sh
cargo install dbx-tools-model-proxy
```

Release builds also publish `dbx-model-proxy` as a GitHub release asset for
each configured platform.

The proxy uses `aigw-openai` and `aigw-anthropic` as protocol adapters.
OpenAI Chat Completions and Anthropic Messages requests can target either
Databricks Chat Completions or Responses. Native Responses input currently
targets Responses without conversion.

Authentication comes from `dbx-tools-core`. The proxy resolves the
selected Databricks profile, obtains cached or refreshed credentials, and
retries one upstream `401` after refreshing the rejected token.
`dbx-tools-model` discovers the workspace's serving endpoints, caches the
catalogue for five minutes, and resolves loose model values before forwarding.
For example, `"model": "gpt"` selects the highest-ranked deployed GPT model.

## Authentication And Rate-Limit Identity

The binary creates one `DatabricksClient` at startup. That client owns upstream
authentication, while each incoming request supplies the principal used to
partition reactive rate-limit cooldowns. Identity selection never calls a
Databricks API:

- With `dbx model-proxy --profile PROFILE`, `dbx-tools-core` resolves that
  profile and uses its normal cached token lifecycle. U2M profiles use the
  Databricks CLI when available and may invoke login when renewal requires it.
  PAT and U2M requests key cooldowns by the profile name; the token itself is
  never part of the key.
- M2M profiles key cooldowns by OAuth client ID. Token refreshes can replace the
  access token without changing the rate-limit identity.
- In a Databricks App using App SP, startup resolves
  `DATABRICKS_HOST`, `DATABRICKS_CLIENT_ID`, and
  `DATABRICKS_CLIENT_SECRET`. The key is therefore the normalized App host,
  service-principal client ID, and resolved serving endpoint.
- Trusted `x-forwarded-user` and `x-forwarded-email` headers partition requests
  by App user. When those are absent, the proxy may decode `sub`, `user_id`,
  `oid`, `client_id`, `azp`, `email`, or `preferred_username` from an incoming
  bearer JWT without verifying it. This decode is only a local rate-limit
  partitioning hint and never authenticates the request.
- Forwarded OBO identity does not replace the startup client's upstream
  credential. The standalone binary has no request-scoped AppKit context, so
  automatic App startup normally resolves App SP. A host that needs true OBO
  upstream calls must construct request-scoped clients and pass the request
  headers through `DatabricksAuthOptions`.

The resulting key is `[normalized Databricks host, current principal, resolved
model]`. Different users, service principals, hosts, or serving endpoints never
share a cooldown. Keys are stored only in process memory, have no default count
limit, and disappear when the proxy exits.

## Run

```sh
cargo run --manifest-path packages/rs/model-proxy/Cargo.toml -- \
  --profile PROFILE --target auto
```

The server listens on `127.0.0.1:4000` by default. `--port` reads
`DATABRICKS_APP_PORT` when present.

Request bodies default to 25 MB through `--max-request-bytes` /
`MAX_REQUEST_BYTES`. Embedded JPEG, PNG, and WebP inputs are detected from their
bytes across OpenAI Chat, Responses, and Anthropic base64 shapes. Images larger
than 2 MB are re-encoded and resized proportionally to at most a 1,568-pixel
edge and 1.15 megapixels before protocol translation. Images at or below 2 MB
and remote image URLs are unchanged, and the proxy never fetches image URLs
itself. Set `--image-resize-threshold-bytes` or
`IMAGE_RESIZE_THRESHOLD_BYTES` to change the size threshold.

`LOG_LEVEL` accepts `debug`, `info`, `warn`, or `error`, case-insensitively,
and defaults to `info`. Request summaries include protocol, selected model,
streaming mode, status, latency, raw request bytes, a fast `tokenx-rs` token
estimate, and the immediate TCP peer IP and port without logging request bodies
or credentials. The peer can be a local or platform proxy rather than the end
user. The estimate includes model-visible JSON but excludes encrypted
reasoning/compaction state, signatures, and embedded image, file, audio, and
screenshot payloads. Buffered responses also report upstream input, output, and
total usage. Streams log connection and completion separately; completion
includes response bytes, total duration, cancellation/failure state, and usage
from both translated and pass-through Chat or Responses events. Pass-through
usage comes from complete parsed SSE events while the original chunks are
forwarded unchanged. Observation is capped at 1 MB per event; a malformed or
larger event safely retains the estimate. Reported usage reconciles process-local
reservations across Chat Completions, Responses, Codex,
Anthropic translations, and embeddings. Unused output reservations are
credited immediately, and actual output is recorded when no maximum was
specified. Each workspace/model queue admits requests FIFO and wakes its head
when reconciliation frees capacity, without blocking unrelated models. After
three consistent samples outside a five-percent noise band, a bounded
per-model exponential moving ratio calibrates raw input estimates against
actual usage. Logs include both the raw estimate and applied factor. Streaming
Chat Completions defaults
`stream_options.include_usage` to `true`; an explicit caller value is
preserved.

Import `postman/model-proxy.postman_collection.json` into Postman. The
collection has separate folders for `--target chat` and `--target responses`;
restart the proxy with the folder's documented command before running it. Set
`chatModel` and `responsesModel` to endpoints available in the selected
workspace.

Supported routes:

- `GET /v1/models`
- `POST /v1/embeddings`
- `POST /v1/chat/completions`
- `POST /v1/responses`
- `POST /v1/messages`
- `GET /healthz`

`GET /v1/models` reads the cached live serving-endpoint catalogue. Standard
requests receive an OpenAI `object` / `data` envelope. An `originator` header
whose value starts with `codex` receives a Codex `models` envelope. Both use
the identities returned by Databricks directly: OpenAI uses the serving
endpoint name, while Codex maps `databricks-<model>` to the gateway's
`system.ai.<model>` identity. No `databricks/` or `dbx/` namespace is added.
Requests targeting Responses default a missing `truncation` field to `"auto"`;
an explicit caller value is preserved. Databricks AI Gateway accepts this
default on Open Responses for Claude and Kimi and on Codex Responses for GPT
and Kimi. Claude itself is not enabled on the Codex route.
Unfiltered responses list chat/LLM families first and embedding families
second, alphabetically sorting families within each tier. Each family uses the
same version, variant, and class preference as a search for that family.
Recognized unclassified models remain in the chat/LLM tier. Custom and
unrecognized endpoints sort by name last. Codex priorities follow the resulting
order.

Use `?search=gpt` to apply the same fuzzy scoring and ordering as model
resolution. Add `?extended=true` to include the score, service names,
capability class, profile, task, state, and other catalogue metadata. Extended
output defaults to `false`.

`POST /v1/embeddings` resolves the requested model only among deployed
embedding endpoints, forwards the request to that endpoint's `invocations`
route, and preserves the OpenAI embedding response.

`--target responses` forces Chat Completions or Anthropic Messages input
through the Responses request translator. `--target chat` sends canonical
Chat Completions. `--target auto` selects Responses for Responses clients,
models listed by the current Databricks Responses documentation, Codex clients,
and requests containing Responses-only hosted tools or conversation fields.

Native Responses requests are forwarded without a capability allow-list, so
Databricks-supported `function`, `custom`, `apply_patch`, `shell`,
`image_generation`, `mcp`, and `web_search` tools, image inputs, conversation
state, background mode, and future request fields remain intact. Chat and
Anthropic image blocks are translated to Responses `input_image` content.
Chat-hosted tools are preserved in Responses form instead of being rejected as
malformed function tools. A forced `--target chat` returns a clear client error
for Responses-only features rather than silently dropping them.

Codex model records obtain image-input, web-search, and patch capability sets
from the corresponding Databricks documentation pages. The parsed model lists
are cached for one day and matched against endpoint, model-service, and provider
identities from the live workspace catalogue. The same parser generates a
committed snapshot during repository synthesis, and the binary embeds that
snapshot as its offline fallback. A failed page refresh retains the matching
capabilities from the embedded snapshot without blocking model listing. This
avoids embedding a handwritten model/version matrix while still using the
unified local execution tool shape expected by current Codex clients.

### Codex Catalogue Discovery

Codex discovery is a separate wire contract over the same route:

- a standard `GET /v1/models` receives the OpenAI `data` envelope;
- a request carrying `originator: codex_cli_rs` receives the Codex `models`
  envelope used by `codex debug models` and the `/model` picker.

Codex validates the complete remote catalogue before merging it with its bundled
models. One incompatible record causes it to retain the bundled catalogue.
Reasoning levels must therefore remain `{ effort, description }` objects,
`web_search_tool_type` must remain a non-null enum value, and
`supports_search_tool` carries the independent capability flag.

The Codex envelope deliberately excludes embeddings, Claude, Gemini,
unrecognized identities, and endpoint names without the required
`databricks-` prefix. Those models do not become compatible merely because they
appear in the standard OpenAI envelope; adding a family requires separate
Responses and tool-replay validation.

Compare remote and bundled discovery without changing provider configuration:

```sh
codex debug models | jq '.models[] | {slug, visibility}'
codex debug models --bundled | jq '.models[] | {slug, visibility}'
```

Run the bounded real-client regression with an installed Codex 0.148.0:

```sh
RUN_CODEX_DISCOVERY_TESTS=1 RUSTC_WRAPPER= \
  cargo test -p dbx-tools-model --test model \
  codex_real_client_discovers_fixture_catalogue --offline -- --nocapture
```

The test builds the catalogue through `models_payload_with_capabilities`, serves
it from a loopback-only fixture, uses synthetic authentication and an isolated
temporary `CODEX_HOME`, verifies the expected HTTP request, and compares
discovered slugs rather than unstable total counts.

Streaming requests use SSE without buffering the upstream response. Matching
protocols pass the upstream byte stream through directly, including Chat
Completions to Chat Completions and Responses to Responses for Codex clients.
Cross-protocol streams pass through aigateway's stateful Chat Completions or
Responses parser. Anthropic output uses aigateway's native SSE encoder, while
OpenAI Chat Completions output uses the proxy's canonical event encoder.

Responses input currently targets only Responses, so Responses-to-Chat
translation is outside the supported route matrix.

Databricks errors are returned with their original status, body, and content
type. The proxy also forwards `Retry-After`, request and correlation IDs,
rate-limit headers, quota names, and Databricks limit details.

HTTP 429 responses pause the process-local host/principal/model gate described
above. One request probes after the shared cooldown while other streaming and
non-streaming requests for the same key remain paused. The `Retry-After`
response header controls the delay when present, followed by the documented
Foundation Model API `error.retry_after` JSON value. An input-token 429 without
either waits for the local token window, or 60 seconds when process-local
history cannot explain the workspace limit. Other 429s use BackON jittered
exponential delays from one second to one minute. Every retry reacquires token
admission and owns exactly one reservation. Every 429 logs a returned
`error.message`, including the final attempt. The default five retries mean one
initial request plus up to five retries.
After the final attempt, the original 429 status, body, and rate-limit headers
are returned to the caller. Configure `RATE_LIMIT_RETRIES`,
`RATE_LIMIT_INITIAL_DELAY_MS`, and `RATE_LIMIT_MAX_DELAY_MS`, or the matching
CLI flags. Set retries to `0` to disable both retries and coordinated cooldowns.
Only an initial HTTP 429 is retried; an SSE error after streaming begins cannot
be replayed safely.

The process-local token queue reads Databricks' published Enterprise
pay-per-token ITPM and OTPM limits from the same daily documentation cache and
generated-fallback pattern used for model capabilities and retirement status.
Input and output windows are tracked separately for each resolved model and
workspace. Requests reserve a `tokenx-rs` input estimate plus any explicit
`max_output_tokens`, `max_completion_tokens`, or `max_tokens` value. Claude
Sonnet 4 reserves its documented 1,000-token default when no output limit is
present. An active queue rejects an input estimate above its complete
per-minute budget with a local structured 429 rather than clamping it.

`RATE_LIMIT_MODE` / `--rate-limit-mode` accepts `auto`, `on`, or `off` and
defaults to `auto`. Auto mode leaves each workspace/model key unthrottled until
its first 429 message containing `Exceeded workspace input tokens`,
case-insensitively. `on` applies budgets immediately; `off` never applies them.

Use `INPUT_TOKENS_PER_MINUTE` / `--input-tokens-per-minute` and
`OUTPUT_TOKENS_PER_MINUTE` / `--output-tokens-per-minute` to override the
published limits. Set `PROVISIONED_THROUGHPUT=true` or pass
`--provisioned-throughput` to disable both TPM windows. QPH remains enforced by
Databricks because process-local tracking cannot coordinate a workspace across
proxy replicas. `/healthz` exposes process-local counters for automatic
activation, admission waits, oversized rejections, post-admission input 429s,
retry reacquisition, and fallback full-window delays.