Model detail
glm-5.3-flashx
GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s. Built on the same hybrid sparse and linear attention architecture (320B total parameters, 18B active), it is suited for efficient coding, visual understanding, and long-horizon agent tasks with a 1M-token context window.
Model specs
- Context length
- 1M
- Max output
- 128K
- I/O modalities
- Text / Image / File / Video
- Released
- 2026-09
02
API endpoints
One MixRoute gateway, OpenAI-compatible
-
OpenAI-compatible
/v1/chat/completionsPOST
03
Pricing
Flat rate, unit: /1M Tokens
Input
$0.3700 /1M Tokens
Completion
$1.2500 /1M Tokens
Cache read
$0.0750 /1M Tokens
05
Selection summary
Quickly judge whether it fits your workload
What is this model good for?
High-speed variant of Z.ai's GLM-5.3-Flash
GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s.
What should you check before using it?
1,000,000 context, 128,000 output tokens
glm-5.3-flashx supports reasoning, tool calling, structured output, streaming, prompt caching, vision input and the OpenAI SDK.
Why use it through MixRoute?
glm-5.3-flashx takes text and images on one route
Text and image inputs share the same route and the same key. MixRoute serves glm-5.3-flashx on the OpenAI-compatible endpoint /v1/chat/completions; point your existing client at https://api.mixroute.ai/v1 with the model ID glm-5.3-flashx and re-run your integration tests.
07
Capabilities & compatibility
Unconfirmed capabilities stay "Unknown," no guesses
Tool calling
glm-5.3-flashx supports function and tool calls.
tools definition; the response returns tool_calls to execute and feed results back.
Structured output
glm-5.3-flashx can produce JSON-formatted output on demand.
response_format: { type: "json_schema" } forces valid JSON output.
Streaming
glm-5.3-flashx supports streamed responses.
stream: true streams tokens incrementally, so chat UIs render as they arrive.
Prompt caching
glm-5.3-flashx supports prompt caching under supported routing conditions.
Vision input
glm-5.3-flashx can process image inputs alongside text.
image_url content part alongside text messages.
OpenAI SDK
Call glm-5.3-flashx through MixRoute's OpenAI-compatible endpoint.
base_url to https://api.mixroute.ai/v1; everything else follows the OpenAI style.
08
Token cost estimator
Live estimate from this page's pricing, not an actual bill
Share of the same prefix read repeatedly, up to 100%
09
FAQ
Reasoning tokens are billed as output on glm-5.3-flashx, at $1.25 per 1,000,000 token. They are generated before the visible answer and count toward both latency and spend, so a higher reasoning setting costs more even when the final answer is short.
One endpoint, a testable decision
Test this model and alternate routes with the same request format
Start from a real workload, then let quality, total cost, and failure conditions decide whether to send production traffic.