Skip to content

Model detail

qwen3.6-plus

Provided by Alibaba
Pay-as-you-go Dynamic pricing 2 tiers

Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and high-performance inference. Compared to the 3.5 series, it delivers major gains in agentic coding, front-end development, and overall reasoning, with a significantly improved “vibe coding” experience. The model excels at complex tasks such as 3D scenes, games, and repository-level problem solving, achieving a 78.8 score on SWE-bench Verified. It represents a substantial leap in both pure-text and multimodal capabilities, performing at the level of leading state-of-the-art models.

Model specs

Context length
Max output
I/O modalities
Released
2026-04

02

API endpoints

One MixRoute gateway, OpenAI-compatible

  • OpenAI-compatible /v1/chat/completions POST

03

Tiered pricing

Unit: /1M Tokens

Tier Input /1M Tokens Output /1M Tokens Cache read /1M Tokens Cache write /1M Tokens
short Length ≤ 256K $0.5000 $3.0000 $0.0500 $0.6250
long Length > 256K $2.0000 $6.0000 $0.2000 $2.5000

08

Token cost estimator

Live estimate from this page's pricing, not an actual bill

Share of the same prefix read repeatedly, up to 100%

One endpoint, a testable decision

Test this model and alternate routes with the same request format

Start from a real workload, then let quality, total cost, and failure conditions decide whether to send production traffic.