Skip to content

Model detail

qwen3.6-max-preview

Provided by Alibaba
Pay-as-you-go Dynamic pricing 2 tiers

Qwen3.6-Max-Preview is a proprietary frontier model from Alibaba Cloud built on a sparse mixture-of-experts architecture with approximately 1 trillion total parameters. It is optimized for agentic coding, tool use, and long-context reasoning, supporting a 262K token context window. The model includes an integrated thinking mode that preserves reasoning traces across multi-turn conversations and supports structured output and function calling. Access is available exclusively through the Alibaba Cloud Model Studio and Qwen Studio APIs; no open weights are provided.

Model specs

Context length
Max output
I/O modalities
Released
2026-04

02

API endpoints

One MixRoute gateway, OpenAI-compatible

  • OpenAI-compatible /v1/chat/completions POST

03

Tiered pricing

Unit: /1M Tokens

Tier Input /1M Tokens Output /1M Tokens Cache read /1M Tokens Cache write /1M Tokens
short Length ≤ 128K $1.3000 $7.8000 $0.1300 $1.6250
long Length > 128K $2.0000 $12.0000 $0.2000 $2.5000

08

Token cost estimator

Live estimate from this page's pricing, not an actual bill

Share of the same prefix read repeatedly, up to 100%

One endpoint, a testable decision

Test this model and alternate routes with the same request format

Start from a real workload, then let quality, total cost, and failure conditions decide whether to send production traffic.