Skip to main content

Crate edge_completions

Crate edge_completions 

Source
Expand description

§edge-completions

CI crates.io docs.rs license

An ergonomic, typed Rust SDK and optional command-line client for OpenAI-compatible chat completions through Cloudflare. The first supported convenience model is moonshotai/kimi-k3.

This is an independent open-source project. It is not affiliated with, endorsed by, or sponsored by Cloudflare, Inc. Cloudflare is a trademark of Cloudflare, Inc.

§Why this crate

  • A native ChatCompletions trait with Send futures and no boxing on the generic path.
  • An explicit DynChatCompletions adapter when runtime type erasure is needed.
  • A sealed typestate request builder that rejects illegal construction sequences at compile time.
  • Validated newtypes for account IDs, API tokens, models, URLs, timeouts, and limits.
  • Typed tool schemas, arguments, and results without a public untyped JSON escape hatch.
  • thiserror errors for configuration, transport, provider, response, and tool failures.
  • Drop-based request cancellation, typed whole-request timeouts, and no hidden retry or background-task policy.
  • Redacted token debug output and bounded response-body decoding.
  • An opt-in edge-completions command that prints typed assistant text and never prints provider envelopes.
  • No autonomous tool loop: your application retains authorization and execution control.

§Install

Add the library:

cargo add edge-completions

Install the optional command-line client:

cargo install edge-completions --features cli

The cli feature is intentionally disabled for library consumers, so SDK-only builds do not compile command-line dependencies. The minimum supported Rust version is 1.86.

Reqwest’s asynchronous transport is Tokio-backed, so applications using the SDK need a Tokio runtime. Typed tool definitions use Schemars and Serde:

cargo add tokio --features macros,rt-multi-thread
cargo add schemars
cargo add serde --features derive

§Configuration

Create a scoped Cloudflare API token and set:

export CLOUDFLARE_ACCOUNT_ID="your_cloudflare_account_id"
export CLOUDFLARE_API_TOKEN="your_least_privilege_api_token"

Client::from_env() also accepts the original KIMI3_ON_CLOUDFLARE_API_KEY variable as a compatibility fallback. New applications should use CLOUDFLARE_API_TOKEN.

You can find the account ID in the Cloudflare dashboard after selecting your account, or follow Cloudflare’s account and zone ID guide. If Wrangler is installed, wrangler whoami also reports the active account ID. Never commit either value.

Cloudflare routes third-party models such as Kimi K3 through the account’s default AI Gateway when no gateway header is present. Set GatewayId only when the request must use a named gateway. Follow Cloudflare’s AI Gateway setup guide when selecting least-privilege token permissions.

§Command line

Validate configuration without making an API request:

source ~/.zshrc
edge-completions check

Expected output:

configuration is valid

Send a prompt and print only the assistant’s text:

edge-completions chat \
  --temperature 0.2 \
  --max-tokens 200 \
  "Explain typed API boundaries in one sentence."

Prompts can also come from standard input:

printf '%s\n' 'Explain capability traits concisely.' | edge-completions chat

Use edge-completions help chat for all chat options. The command returns typed, sanitized errors on standard error. It does not expose raw request or response envelopes and does not execute model-proposed tools. See the complete CLI reference, including exit codes and the base-URL security contract.

§Library quickstart

use edge_completions::{AssistantOutput, ChatMessage, ChatRequest, Client, Error};

async fn answer() -> Result<(), Error> {
    let client = Client::from_env()?;
    let request = ChatRequest::kimi_k3_builder()
        .message(ChatMessage::user(
            "Explain typed API boundaries in one sentence.",
        ))
        .build();

    let completion = client.chat(&request).await?;
    let text = match completion.first_choice()?.message().output() {
        AssistantOutput::Text(text) => text,
        AssistantOutput::TextAndToolCalls { text, .. } => text,
        AssistantOutput::ToolCalls(_) | AssistantOutput::Empty => {
            return Err(Error::MissingContent);
        }
    };

    println!("{text}");
    Ok(())
}

Run the complete example:

source ~/.zshrc
cargo run --example simple_chat

Expected result: the example prints the model identifier followed by validated assistant text. It never prints the provider response envelope.

§Typed tool calls

A tool contract ties together its generated JSON Schema, validated input type, and serializable output type:

use edge_completions::ToolDefinition;
use schemars::JsonSchema;
use serde::{Deserialize, Serialize};

struct GetWeather;

#[derive(Deserialize, JsonSchema)]
struct WeatherArguments {
    city: String,
}

#[derive(Serialize)]
struct WeatherReport {
    temperature_celsius: i16,
}

impl ToolDefinition for GetWeather {
    type Arguments = WeatherArguments;
    type Output = WeatherReport;

    const NAME: &'static str = "get_weather";
    const DESCRIPTION: &'static str = "Get the weather for a city";
}

Decode a provider proposal only through the matching contract:

let call = completion.first_choice()?.message().first_tool_call()?;
let validated_call = call.validate::<GetWeather>()?;

// The application authorizes and executes the action here.
let report = get_weather(validated_call.arguments());
let result_message = validated_call.result(&report)?;

Run the complete two-turn example:

source ~/.zshrc
cargo run --example weather_tool

The example returns deterministic sample weather; it does not call a weather service or claim to provide a live forecast.

Output follows this shape. The final sentence depends on the model:

Tool requested: get_weather
Validated city: San Francisco
Final answer: ...

§Compile-time composition

The typestate builder represents valid request states as types. Adding a message or tool consumes one state and returns the next, while trait bounds control which operations exist:

let request = ChatRequest::kimi_k3_builder()
    .message(ChatMessage::user("Weather in Paris?"))
    .tool(FunctionTool::for_tool::<GetWeather>()?)
    .tool_choice(ToolChoice::Required)
    .build();

Calling build before message, or tool_choice before tool, does not compile. AssistantOutput models response alternatives as an exhaustive sum type, so callers must handle text, tool calls, text and tool calls, and empty output.

GuaranteeEnforced byFailure point
A request has at least one messageChatRequestBuilder typestateCompilation
Tool choice follows at least one toolWithTools trait boundCompilation
A tool result matches its tool contractValidatedToolCall<T>Compilation
Every supported assistant outcome is handledExhaustive AssistantOutput matchCompilation
Provider data matches the declared contractTyped deserialization and validationRuntime

The same design has a small categorical interpretation: request states are objects, legal transitions are composable morphisms, and assistant output is a coproduct with a product branch. External network and model data still require typed runtime validation.

See the type-system guide for the complete state graph, compile-fail examples, and the boundary between static guarantees and runtime checks.

§Native async boundary

Use generic dispatch for the normal application boundary:

use edge_completions::{
    ChatCompletion, ChatCompletions, ChatRequest, Error,
};

async fn answer<C>(
    ai: &C,
    request: &ChatRequest,
) -> Result<ChatCompletion, Error>
where
    C: ChatCompletions,
{
    ai.complete(request).await
}

ChatCompletions::complete returns impl Future + Send. The static path keeps the concrete future and performs no heap allocation for trait dispatch.

Use explicit type erasure only when the concrete implementation is selected at runtime:

use edge_completions::{
    ChatCompletion, ChatRequest, DynChatCompletions, Error,
};

async fn answer_dynamic(
    ai: &dyn DynChatCompletions,
    request: &ChatRequest,
) -> Result<ChatCompletion, Error> {
    ai.complete_boxed(request).await
}

DynChatCompletions returns BoxChatFuture, making its one allocation per call visible in the API. See the async execution guide for cancellation, deadlines, concurrency, runtime ownership, and retry policy.

Dropping either future cancels the in-flight exchange. The SDK spawns no task, retains no partial response, and performs no retry. RequestTimeout covers the connection and complete response body; expiry returns Error::Timeout.

§Migrating from 0.3

Version 0.4 replaces the async-trait capability method with a native Rust future. This is an intentional pre-1.0 compatibility change.

0.3 API0.4 replacement
&dyn ChatCompletionsGeneric C: ChatCompletions
Arc<dyn ChatCompletions>Arc<dyn DynChatCompletions>
ai.complete(request) through dynai.complete_boxed(request)
#[async_trait] impl ChatCompletionsNative implementation with async fn complete

Client::chat, request and response types, typestate construction, and typed tool contracts are unchanged.

§Migrating from 0.2

Version 0.3 introduced these additive type-system APIs, and version 0.4 retains them. Existing ChatRequest::new, ChatRequest::kimi_k3, with_tool, with_tools, and ToolCall::arguments_for calls remain available. New code should prefer these replacements:

0.2 APIPreferred 0.3 APIBenefit
ChatRequest::kimi_k3(messages)ChatRequest::kimi_k3_builder().message(...).build()Non-empty messages are proven at compile time
request.with_tools(tools, choice).tool(...).tool_choice(choice)Tool choice cannot precede a tool
call.arguments_for::<T>()call.validate::<T>()The validation proof remains available for result encoding
message.content() plus tool_calls()message.output()All supported output combinations are handled together

No automatic migration is required. Adopt the new API when compile-time guarantees are useful at the call site.

Keep the environment-based credential lookup while customizing the transport:

use edge_completions::{Client, GatewayId};

fn client() -> Result<Client, edge_completions::Error> {
    Client::builder_from_env()?
        .gateway_id(GatewayId::new("production-gateway")?)
        .build()
}

§Features

FeatureDefaultPurpose
cliNoBuilds the installable edge-completions command.

§Error and security model

  • Every library failure is typed with thiserror; production code contains no unwrap, expect, or panic path.
  • Configured request deadlines return Error::Timeout; caller cancellation drops the future and therefore returns no SDK result.
  • Local request cardinality and tool-choice sequencing are enforced by typestate; provider responses and model-produced tool calls are checked at runtime.
  • ValidatedToolCall<T> is a proof-carrying value that binds decoded arguments and encoded output to the same ToolDefinition at compile time.
  • API tokens use redacted Debug and are never included in errors.
  • Success and failure bodies are size-bounded. Unknown error bodies are not exposed.
  • Only HTTPS endpoints are accepted, except loopback HTTP for local tests.
  • Tool names and arguments are model-controlled and remain untrusted until validate::<T>() produces a ValidatedToolCall<T> witness. The compatibility helper arguments_for::<T>() performs the same validation and returns the arguments.
  • Tool execution is intentionally outside this crate.
  • The HTTP adapter explicitly disables automatic retries. Applications own any retry classification and bounded concurrency policy.

A 402 Payment Required with Cloudflare code 2021 means the account or AI Gateway lacks usable Workers AI balance/provider billing. Add the required balance or configure BYOK before retrying the live examples.

§Validation

cargo fmt --all -- --check
cargo clippy --locked --all-targets --all-features -- -D warnings
cargo test --locked --all-targets --all-features
cargo test --locked --doc --all-features
RUSTDOCFLAGS="-D warnings -D missing_docs" cargo doc --locked --no-deps --all-features
cargo deny check
cargo audit
cargo publish --locked --dry-run

Tests cover exact request serialization, native and dynamic trait use, future Send guarantees, caller cancellation, typed timeouts, concurrent calls, dropped connections, typed tool-call round trips, HTTP authentication, response validation, provider errors, response limits, TLS policy, and secret redaction against local servers. Tests do not make live provider calls.

§Scope and limitations

  • Supported now: non-streaming chat completions, typed function tools, Kimi K3 convenience construction, optional AI Gateway ID, timeout and body-size policy.
  • Not supported yet: streaming, multimodal content, embeddings, Responses API, provider-specific raw extension maps, or autonomous tool execution.
  • Provider contract drift remains possible. Contract changes should land with a versioned test before expanding the public API.

See the architecture, async execution guide, type-system guide, contributing guide, CLI reference, security policy, and release process.

§License

Licensed under either of Apache License, Version 2.0 or MIT license at your option.

Modules§

async_model
Native async dispatch, explicit type erasure, cancellation, and deadlines.
request_state
Compile-time states used by ChatRequestBuilder.
type_system
Compile-time request states, exhaustive output handling, and typed tool proofs.

Structs§

AccountId
A validated Cloudflare account identifier.
ApiBaseUrl
A validated API base URL.
ApiToken
A validated API token whose debug output is always redacted.
AssistantMessage
A typed assistant message returned by the provider.
ChatChoice
One provider-generated completion choice.
ChatCompletion
A typed chat-completion response.
ChatMessage
A typed chat message accepted by the provider request contract.
ChatRequest
A validated, serializable chat-completion request.
ChatRequestBuilder
A typestate builder that makes invalid request-construction sequences fail to compile.
Client
HTTP implementation of the typed chat-completion capability.
ClientBuilder
Builder for transport policy and optional Cloudflare AI Gateway routing.
FunctionTool
A validated OpenAI-compatible function-tool definition.
GatewayId
A validated value for Cloudflare’s optional cf-aig-gateway-id header.
MaxTokens
A validated, non-zero completion-token limit.
ModelId
A validated provider model identifier, such as moonshotai/kimi-k3.
ProviderErrorCode
A numeric error code returned by Cloudflare’s API envelope.
RequestTimeout
A validated request timeout.
ResponseSizeLimit
A validated maximum response-body size.
Temperature
Sampling temperature constrained to the provider’s documented [0, 2] range.
ToolCall
A function-tool call proposed by the model.
Usage
Token accounting returned by the provider.
ValidatedToolCall
Proof that a model-produced tool call matched and decoded through T.

Enums§

AssistantOutput
Exhaustive alternatives carried by a typed assistant message.
Error
Errors returned while configuring or calling the provider endpoint.
FinishReason
Why the provider stopped generating a completion.
InvalidConfiguration
Failures raised while validating local SDK configuration.
ProviderFailure
A sanitized failure decoded from a non-success provider response.
ToolChoice
Provider setting for selecting function tools.
ToolError
Failures while defining, decoding, or encoding a typed tool contract.

Constants§

ACCOUNT_ID_ENV
Environment variable read by Client::from_env for the Cloudflare account ID.
API_TOKEN_ENV
Environment variable read by Client::from_env for the Cloudflare API token.

Traits§

ChatCompletions
Provider-independent capability boundary for chat-completion adapters.
DynChatCompletions
Object-safe adapter for runtime-selected chat-completion implementations.
ToolDefinition
A typed tool contract used for schema generation, argument decoding, and result encoding.

Type Aliases§

BoxChatFuture
Boxed future returned by DynChatCompletions.