Skip to main content

Module zai

Module zai 

Source
Expand description

Z.ai / 智谱 GLM (BigModel) proprietary extensions to the OpenAI-compatible API.

Everything in this module is gated on the zai cargo feature.

GLM is served behind an OpenAI-compatible endpoint at https://open.bigmodel.cn/api/paas/v4 (the international brand is Z.ai, https://api.z.ai/api/paas/v4), authenticated with the usual Authorization: Bearer <key> header. Most of the request body is standard OpenAI; this module collects the parts that are not.

GLM’s divergences fall into two groups, handled differently:

  • Generic controls that GLM merely spells its own way live directly on RequestBody as zai-gated fields — do_sample and tool_stream. The keys GLM shares with another provider are unified rather than duplicated: thinking and user_id are gated on any(deepseek, zai), request_id on any(vllm, zai), so enabling several provider features at once never emits a key twice.
  • Platform-ecosystem extensions — features tied to Zhipu’s own platform rather than to text generation — are grouped here:
    • PlatformParams, #[serde(flatten)]ed into the request body through RequestBody::zai_platform (currently the watermark_enabled compliance flag).
    • the retrieval and web_search tool types, added as zai-gated variants of RequestTool;
    • the web_search results GLM returns at the top level of a chat completion.
use openai_interface::chat::create::request::{Message, RequestBody};
use openai_interface::zai::{PlatformParams, SearchEngine, WebSearchTool};

let request = RequestBody {
    messages: vec![Message::user("最近有什么关于 Rust 的新闻?")],
    model: "glm-4.6".to_string(),
    do_sample: Some(false),
    zai_platform: Some(PlatformParams {
        watermark_enabled: Some(false),
    }),
    tools: Some(vec![
        openai_interface::chat::create::request::RequestTool::WebSearch {
            web_search: WebSearchTool {
                search_engine: Some(SearchEngine::SearchProJina),
                enable: Some(true),
                ..Default::default()
            },
        },
    ]),
    ..Default::default()
};

let json = serde_json::to_value(&request).unwrap();
assert_eq!(json["do_sample"], serde_json::json!(false));
assert_eq!(json["watermark_enabled"], serde_json::json!(false));
assert_eq!(json["tools"][0]["type"], serde_json::json!("web_search"));

§Fields that need no gate

  • reasoning_effort is already an ungated RequestBody field, and its ReasoningEffort enum already covers every GLM value (none, minimal, low, medium, high, xhigh, max). GLM only honours it when thinking is enabled, and the value-to-depth mapping is model-specific.
  • temperature, top_p and max_tokens are standard keys; GLM simply accepts narrower ranges (temperature and top_p in [0, 1], two decimals; max_tokens up to 131072). No extra field models a range.
  • reasoning_content on the response message and streamed delta is covered by the always-available reasoning feature.
  • GLM’s extra finish_reason values (sensitive, model_context_window_exceeded, network_error) are preserved by FinishReason::Unknown rather than modelled as variants, so an unexpected terminator never fails deserialization.

§Not implemented

GLM also exposes an mcp tool type, image/video generation, embeddings and a separate GLM Coding Plan endpoint. None are documented precisely enough to model here; this feature covers the fields GLM adds to the OpenAI chat completions endpoint.

See the GLM chat-completions reference and its Z.ai equivalent.

Structs§

PlatformParams
GLM platform-ecosystem request parameters, flattened into the request body.
RetrievalTool
GLM’s retrieval tool: grounds the answer in one of Zhipu’s knowledge bases.
WebSearchResult
One web-search result GLM returns in the top-level web_search array of a chat completion.
WebSearchTool
GLM’s web_search tool: lets the model call Zhipu’s web search.

Enums§

ContentSize
How much of each searched page to keep.
ResultSequence
Where the search results are placed relative to the answer.
SearchEngine
The web-search backend GLM should use.
SearchRecencyFilter
How recent a web-search result may be.