Skip to content

Model detail

glm-5.3-flashx

Provided by GLM
Pay-as-you-go

GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s. Built on the same hybrid sparse and linear attention architecture (320B total parameters, 18B active), it is suited for efficient coding, visual understanding, and long-horizon agent tasks with a 1M-token context window.

Model specs

Context length
1M
Max output
128K
I/O modalities
Text / Image / File / Video
Released
2026-09

02

API endpoints

One MixRoute gateway, OpenAI-compatible

  • OpenAI-compatible /v1/chat/completions POST

03

Pricing

Flat rate, unit: /1M Tokens

Input

$0.3700 /1M Tokens

Completion

$1.2500 /1M Tokens

Cache read

$0.0750 /1M Tokens

05

Selection summary

Quickly judge whether it fits your workload

What is this model good for?

High-speed variant of Z.ai's GLM-5.3-Flash

GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s.

What should you check before using it?

1,000,000 context, 128,000 output tokens

glm-5.3-flashx supports reasoning, tool calling, structured output, streaming, prompt caching, vision input and the OpenAI SDK.

Why use it through MixRoute?

glm-5.3-flashx takes text and images on one route

Text and image inputs share the same route and the same key. MixRoute serves glm-5.3-flashx on the OpenAI-compatible endpoint /v1/chat/completions; point your existing client at https://api.mixroute.ai/v1 with the model ID glm-5.3-flashx and re-run your integration tests.

07

Capabilities & compatibility

Unconfirmed capabilities stay "Unknown," no guesses

Available:

Tool calling

glm-5.3-flashx supports function and tool calls.

EX Send a tools definition; the response returns tool_calls to execute and feed results back.
Available:

Structured output

glm-5.3-flashx can produce JSON-formatted output on demand.

EX response_format: { type: "json_schema" } forces valid JSON output.
Available:

Streaming

glm-5.3-flashx supports streamed responses.

EX stream: true streams tokens incrementally, so chat UIs render as they arrive.
Available:

Prompt caching

glm-5.3-flashx supports prompt caching under supported routing conditions.

EX Keep the system prompt and tool definitions at the start of the request; repeated prefixes are automatically billed at the cache-read rate.
Available:

Vision input

glm-5.3-flashx can process image inputs alongside text.

EX Pass a base64-encoded image in a image_url content part alongside text messages.
Available:

OpenAI SDK

Call glm-5.3-flashx through MixRoute's OpenAI-compatible endpoint.

EX Change base_url to https://api.mixroute.ai/v1; everything else follows the OpenAI style.

08

Token cost estimator

Live estimate from this page's pricing, not an actual bill

Share of the same prefix read repeatedly, up to 100%

09

FAQ

Reasoning tokens are billed as output on glm-5.3-flashx, at $1.25 per 1,000,000 token. They are generated before the visible answer and count toward both latency and spend, so a higher reasoning setting costs more even when the final answer is short.

One endpoint, a testable decision

Test this model and alternate routes with the same request format

Start from a real workload, then let quality, total cost, and failure conditions decide whether to send production traffic.