Chinese LLMs Compared: Qwen, DeepSeek, Kimi, GLM, MiniMax and ERNIE

“Every list I read says a different Chinese model wins. Which one do I actually pick?” In September 2026 the choice usually comes down to six families: Qwen, DeepSeek, Kimi, GLM, MiniMax, and ERNIE. No family wins every job. The top row moves with your workload, your hardware, your context needs, and whether you need downloadable weights. DeepSeek’s V4 line, GLM-5.3 and Kimi K3 carry open weights under their own licenses; Qwen3.8-27B and ERNIE 4.5 are Apache 2.0. A “Chinese model” label means the lab is Chinese, not that the model writes Chinese well.
Which Chinese-developed LLM families lead in 2026?
No single family leads, and that is the part worth keeping. Turing Post’s 2026 family review is an outside comparison, not MixRoute’s ranking, and it assigns current strengths this way: DeepSeek V4 for long-context reasoning and agents, Qwen3.8 for coding and multimodal work, Kimi K3 for long-horizon multi-agent tasks, and GLM-5.3 for coding and tool use. The review also states the condition that matters more than any row in it: the best choice depends on workload, hardware, context requirements, and whether you need open weights. So the useful question is not which family sits on top. It is which of those four constraints your project actually hits first.

Six families behind one “Chinese model” label
A ranking filtered by “Chinese” tells you where the lab sits, not how well the model writes the language. That gap explains a lot of the apparent disagreement between leaderboards. These are the families behind the filter:
- Qwen from Alibaba, which pairs a large hosted MoE flagship with open-weight releases at several sizes.
- DeepSeek, whose current V4 line carries MIT-licensed weights for long-context reasoning, coding, and agent work.
- Kimi from Moonshot AI, whose K3 model is open weight under the custom Kimi K3 License and built for long-horizon agent tasks.
- GLM from Z.ai, the Tsinghua-linked lab whose GLM-5.3 is a post-training upgrade of the GLM-5.2 base for software engineering, terminal work, and long-running agents.
- MiniMax, whose M3 model competes as a lower-cost multimodal worker.
- ERNIE from Baidu, which is no longer only a hosted family: Baidu released ERNIE 4.5 weights and development toolkits under Apache 2.0 while keeping a separate hosted ERNIE API.
A leader that changes with the filter
Watch the ordering move when the filter moves. The BenchLM snapshot dated September 18, 2026 lists Kimi K3 at 74.43 and Qwen3.8 Max at 73.17 under BenchAlign v5.2. Both are composite scores, not task pass rates. A composite number does not tell you whether the checkpoint you can download is the one that was scored, so the release and the deployment option behind a row decide whether it belongs on your shortlist.

How do Qwen, DeepSeek, Kimi K3, GLM, MiniMax and ERNIE compare?
Each family ships several releases, and the release you can actually call or download is the row that counts. The table below lists representative 2026 releases with the figures each lab publishes for that release. The weight column separates hosted products from downloadable checkpoints, because a family name does not tell you which release you can run or which license covers it.

| Family (lab) | Representative releases available in 2026 | Reported parameters (release-specific) | Context window | License and weight status | Hosting and API access | Workload the lab targets | Chinese-language evaluation to run |
|---|---|---|---|---|---|---|---|
| Qwen (Alibaba) | Qwen3.8 Max (hosted); Qwen3.8-27B (open); separate downloadable Max-scale text-only checkpoint | Max: 2.4T / 95B active; Qwen3.8-27B: 27B dense | Hosted Max route listed at 1M; some Qwen3.8 releases list 262K native extending up to 1M | Hosted Max is a hosted product, not open weight; Qwen3.8-27B is Apache 2.0; the downloadable Max-scale checkpoint has its own terms | Hosted Max accepts text, images, and video and returns text; smaller dense releases fit more practical deployments | Long-running coding, complex agents, frontend and visual debugging | Test regional terminology and instruction following in both Traditional and Simplified Chinese |
| DeepSeek | DeepSeek V4 Pro 0813; DeepSeek V4-Flash sibling | Public materials report V4 Pro family at 1.6T / 49B active; the DeepSeek V4 Pro model card lists V4-Flash at 284B / 13B active | 1M | MIT-licensed weights for the current V4 line | Hosted API or private deployment; thinking and non-thinking modes, tool calls, and JSON output on the hosted route | Deep reasoning, long-context coding, complex agents, backend and terminal work | Test factual accuracy, long-context recall, and the editing needed for customer replies |
| Kimi (Moonshot AI) | Kimi K3 | 2.8T / 104B active | 1M | Open weights under the custom Kimi K3 License, not MIT; full weights and technical report released | Hosted API, or self-hosting on substantial multi-GPU infrastructure; consult the current official deployment guide, not a fixed hardware figure | Multi-agent research, long-horizon coding, tool use, native multimodal input | Test whether research summaries preserve citations and separate evidence from inference |
| GLM (Z.ai) | GLM-5.3; earlier GLM-5.2 remains available | GLM-5.3 is an MoE reported at roughly 750B total / 40B active; it shares its base with GLM-5.2 | 1M | Official GLM-5.3 weights are downloadable under the GLM-5.3 License; GLM-5.2 is the earlier MIT-licensed release. Read the actual license rather than inferring terms from a version number | Hosted API plus official local deployment documentation for the open checkpoint | Software engineering, terminal work, long-horizon agents, defensive code review, tool use | Test bilingual technical explanations and whether Chinese comments match the code |
| MiniMax | MiniMax M3 | Approximately 428B / 23B active | 1M | Open weights under the MiniMax Community License, not MIT; review terms before redistribution or model-as-a-service use | Hosted API with text, image, and document input; self-hosting is possible under license terms | Lower-cost multimodal analysis and high-volume subtasks | Test tone consistency and correction effort across a batch of short customer tasks |
| ERNIE (Baidu) | ERNIE 4.5 family (open; released in 2025); Baidu hosted ERNIE API | Parameters vary by named variant: MoE variants use 47B or 3B active parameters, the largest model has 424B total, and a 0.3B dense model is also in the family | Not stated in the materials reviewed here | ERNIE 4.5 weights and development toolkits are released under Apache 2.0; the hosted ERNIE API is a separate Baidu commercial route with its own terms | ERNIE 4.5 supports local or cloud deployment from weights; hosted ERNIE is available through Baidu’s API | Multimodal reasoning, instruction following, enterprise and Chinese-market products per Baidu’s ERNIE 4.5 announcement | Test your domain vocabulary, regional wording, and the requested output format |
What the weight and license column does not tell you
Open weight means the parameters are downloadable under a model-specific license. That is not the same as an OSI-approved license, unrestricted commercial use, or inexpensive deployment, so read each release’s terms on their own.
- Hosted Qwen3.8 Max accepts images and video. The separate downloadable Max-scale Qwen checkpoint is described as text-only with mandatory thinking. Features on the hosted product do not carry onto the downloadable artifact.
- GLM-5.2 is MIT-licensed and GLM-5.3 shares its base, but that does not make GLM-5.3 MIT. The GLM-5.3 License is its own document.
- ERNIE 4.5 shipped under Apache 2.0. That release says nothing about every hosted or future ERNIE product.
What do the benchmark scores behind these comparisons actually measure?
A score answers the question its method was built for, and nothing wider. Chinese-model comparisons mix three kinds of numbers: a public evidence-aware snapshot, a broad composite index, and vendor-reported results. Each method sets a ceiling on the claim a score can support.
BenchLM’s Supported and Estimated labels
BenchLM marks rows Supported or Estimated so you can see how much evidence sits under each one. Its overall score combines capabilities, while task-specific columns answer narrower questions. Read the method, the source dates, and the uncertainty range next to the score. A small point difference on its own is not a reason to move your production model.
The Artificial Analysis Intelligence Index is broad, not coding-only
You will also see the Artificial Analysis Intelligence Index quoted in Chinese-model posts, and it is easy to read as a coding ranking. It is not one. The index folds coding, terminal, reasoning, knowledge, and mathematics evaluations into a single composite, which makes it useful for building a shortlist and weak as the last word on a deployment. Coding decisions need coding-specific evidence gathered under your own agent setup.
Vendor-reported results come from the lab that published them
Z.ai publishes the most specific current GLM coding results, and those figures are vendor-reported. The company reports that GLM-5.3 improved from 4.6 to 28.3 on Terminal-Bench 3.0 and from 46.2 to 66.9 on DeepSWE v1.1. On Z.ai’s private Code Bench, GLM-5.3 reached 34.5% at max effort while using about 75,000 output tokens per task, and GLM-5.2 scored 23.4% while using roughly 96,000 tokens. Those numbers describe terminal work and software-engineering tasks under Z.ai’s own evaluation conditions, and they are not independent measurements. Keep that scope attached when you repeat them.
Which Chinese LLM should you use for Chinese-language writing and conversation?
None of the rows above measure Chinese output. A leaderboard row tells you where a model was built and how it scored on broad categories, not how a Simplified Chinese support reply or a Traditional Chinese product page reads. If your product ships copy that customers read in Chinese, the result you can trust comes from a prompt set that mirrors that work, scored blind, with a native speaker deciding the close calls. Published scores can inform the shortlist. They cannot settle it.

Start with families that advertise Chinese or multilingual strength
Three families say something about Chinese or multilingual work in their own materials. Alibaba’s Qwen3-generation materials advertise official support for more than 100 languages and dialects, and the Qwen3.8 line carries that positioning into coding and agent work. ERNIE is Baidu’s Chinese-market line, and ERNIE 4.5 now gives you an Apache 2.0 checkpoint to test instead of only a hosted API. GLM began as bilingual English-Chinese ChatGLM models before scaling into larger MoE architectures. Qwen, ERNIE, and GLM are reasonable places to start a Chinese-language trial, and none of them is a proven winner for your register and your audience.
Run a blind pass with your own production prompts
Build a small prompt set that mirrors what your product does: conversational replies, customer support responses, product copy, technical documentation, and any long-form Chinese writing your team produces. If your audience spans both writing systems, cover both, because Traditional and Simplified Chinese carry separate register and terminology conventions. Pull domain vocabulary from your own work, apply the same system message and tone constraints you use in production, and score the outputs with the model names hidden so they cannot bias the grade. Where tone or cultural fit decides the outcome, a native speaker reviews the drafts. No published ranking score settles Chinese-language quality, so the blind pass runs before you commit to a model.
Three prompts you can reuse in a Chinese-language trial
Treat these as starting inputs and swap the fictional details for your own. Keep the model names hidden from the reviewer and mark each output for factual accuracy, regional wording, format compliance, and the edits needed before use.
| Task | Prompt | What to check |
|---|---|---|
| Traditional Chinese support | 請用台灣繁體中文回覆:「我已付款,為什麼餘額還沒增加?」已知資訊只有付款時間與交易編號;請先問一個必要問題,不承諾到帳時間。 | Asks for the missing payment status without inventing a cause or deadline |
| Simplified Chinese extraction | 从「订单A17已付款;B08尚未付款;C22状态未知」提取订单编号和付款状态,输出JSON数组,不添加原文没有的信息。 | Preserves all three states and produces valid JSON |
| Technical explanation | 用繁體中文向新手說明:HTTP 429 可能是流量限制,也可能是額度不足。請分別給出下一步,不把兩者都寫成等待重試。 | Maintains the distinction and gives an action for each case |
A run only compares to the next one when the model ID, prompt, settings, latency, reported usage, and reviewer edits travel with it. These prompts define a trial. They do not report model results.
If DeepSeek is still on your shortlist, see how DeepSeek and ChatGPT differ on everyday language and reasoning tasks, then put both through the same blind pass before you commit.
Which Chinese LLM should you pick for coding and long-running agents?
The coding shortlist is a set of trial candidates grouped by workload and by license or hosting constraint, not a ranking. Compare accepted results under the same agent setup, because a coding leaderboard does not tell you which model finishes patches inside your tool harness.
The coding shortlist, by workload and constraint
- Text-only repository and terminal loops: GLM-5.3 is the candidate to test first for long-running software engineering and terminal work. Z.ai’s vendor-reported results support that role, and official GLM-5.3 weights are now downloadable under the GLM-5.3 License. If your team also weighs the earlier MIT-licensed GLM-5.2 release, compare GLM-5.2 with Claude Opus 4.8 and Fable 5 before setting a default.
- MIT-licensed self-hosting for backend work: DeepSeek V4 Pro 0813 is the permissive-weight candidate, with a 1M context and thinking and non-thinking modes on the API. Check the exact output setting for the route you use; a large maximum output recommended for local high-effort reasoning is not automatically the hosted API cap.
- Frontend, screenshot, and video work: Qwen3.8 Max accepts text, images, and video and returns text, which fits rebuilding an interface from a screenshot or debugging a rendered frontend against a reference.
- Difficult multimodal coding with open weights: Kimi K3 is the high-capability candidate when visual evidence or very hard reasoning justifies its cost and infrastructure weight. Confirm the Kimi K3 License and the current official deployment guide before choosing it, because a 2.8T-parameter model needs substantial multi-GPU infrastructure.
- High-volume, lower-cost subtasks: MiniMax M3 is positioned as a budget multimodal worker for triage and first-pass review. Verify current route rates and license terms instead of assuming the whole family is always the cheapest option.
The model and the agent are a package
A coding benchmark winner does not automatically become the best coding agent. The agent decides which repository files enter the context, how tool output is presented, which edit formats are allowed, when tests run, what happens after a failed command, and how much context survives between turns. Run the same model inside two different agent systems and the completion rate can move a long way, so pick the model and the agent together. If your team already runs an agent harness, that harness is the thing to hold constant while you swap models. A one-million-token context window is also not an instruction to paste an entire repository into every request. Generated files, old logs, and unrelated documentation pull attention away from the execution path that matters.
Should you self-host an open-weight Chinese LLM or call an API?
Settle the hosting and licensing question before you compare capability numbers, because a model you cannot run or cannot reach does not solve your problem. The open-weight Chinese families do not share one license or one hardware profile, and the API routes that expose them are not identical to the upstream model cards.
Where the open-weight licenses split
- DeepSeek: the current V4 line is MIT-licensed, which gives you a permissive self-hosting option when MIT is the license you want. Apache-2.0 releases are permissive in the same way.
- Qwen: Alibaba licenses per release. Qwen3.8-27B is Apache 2.0, while hosted Qwen3.8 Max and the downloadable Max-scale Qwen checkpoint stay separate artifacts with separate terms.
- GLM: Z.ai offers official GLM-5.3 weights under the GLM-5.3 License. GLM-5.3 shares its base with GLM-5.2, so the older version number does not make the older model smaller or cheaper.
- Kimi: K3 is downloadable under the custom Kimi K3 License.
- MiniMax: M3 sits under the MiniMax Community License, which needs a read before redistribution or a model-as-a-service offering.
- ERNIE: Baidu released ERNIE 4.5 weights and toolkits under Apache 2.0, while the hosted ERNIE API remains a separate commercial route.
What self-hosting actually costs
Downloadable weights are not the same as cheap deployment. Hardware requirements vary enormously across this set: some Qwen open releases and other smaller models can run on a single workstation, while trillion-parameter MoE checkpoints such as Kimi K3 at 2.8T require substantial multi-GPU infrastructure for practical inference. If you want a Max-scale Qwen on your own hardware, use the separate downloadable Max-scale checkpoint, because the hosted route is not a self-hosting entry. The two also differ in modality: hosted Qwen3.8 Max accepts text, images, and video and returns text, while the downloadable Max-scale Qwen checkpoint is text-only with mandatory thinking.
An open checkpoint also says nothing about the cost of running it well, and an open-weight row on a leaderboard does not mean the model is easy to host. For a team without that infrastructure, a hosted API can be the cheaper operational choice even when the weights are free to download. Compare current per-token rates for the exact release and route you plan to use, and remember that reasoning tokens, cache state, retries, and output length all affect the final bill.
API routes can be narrower than the model
A route a host exposes is not the model’s full specification. One portal route may serve GLM-5.3 as text-to-text while another host adds multimodal input or a different context ceiling. The same gap runs inside a single family: the hosted Qwen3.8 Max route accepts images and video, while the downloadable Max-scale Qwen checkpoint is text-only with mandatory thinking. Read the route documentation before you promise a feature to your team, and treat the model card as the starting point rather than the contract.
Once your model list and deployment path are set, the remaining decision is where to buy API access. MixRoute’s model catalog documents exposed model IDs, API modes, and context specifications for the models it lists, and the catalog is a reference, not a quality ranking or a proof of Chinese-language suitability. Verify the exact model ID, modality, context limit, and API mode for the release you choose. Decide whether a MixRoute prepaid credit tier fits your model budget before you commit to a provider.
FAQ
Does winning a “best Chinese LLM” ranking mean the model is best at Chinese-language work?
No. In the ranking cited above, “Chinese model” means the originating lab is Chinese, and the overall score says nothing about Chinese-language quality, so a high composite and a good Traditional Chinese product page are two different measurements. Test your own register: support replies, Traditional and Simplified wording, and the domain terms your customers use. Score the outputs blind with a native speaker on the panel, because the failure you are looking for is tone and terminology, and that shows up in the edits rather than in a benchmark table.
Which Chinese-developed model is the overall leader in 2026?
Not in any way that settles your decision. Rankings differ by benchmark, by release, and by deployment option, and the September 18, 2026 BenchLM snapshot places Kimi K3 above Qwen3.8 Max on its overall composite. Change the filter and the first row moves with it. A point or two on a composite is smaller than the spread your own acceptance criteria will produce, so shortlist two or three releases, run the same tasks and the same reviewer against all of them, and let the accepted results decide.
Which Chinese-developed LLM should I pick for coding?
Shortlist for the repository task and the deployment constraint, then compare accepted results under the same agent setup, because vendor scores alone will not pick your winner. A reasonable first test is GLM-5.3 for text-only long-horizon coding, DeepSeek V4 Pro 0813 for MIT-licensed backend work, Qwen3.8 Max for frontend and visual tasks, and Kimi K3 where difficult multimodal coding justifies the infrastructure cost. GLM-5.2 remains available as an earlier MIT-licensed alternative. Count the whole bill, since retries, reasoning tokens, and tool output all land on it.
Can I download and run DeepSeek, Qwen, Kimi or GLM locally?
Yes, for specific releases, and the license differs for each one. DeepSeek’s V4 line is MIT-licensed, Qwen3.8-27B and ERNIE 4.5 are Apache 2.0, GLM-5.3 uses the GLM-5.3 License, and Kimi K3 uses the custom Kimi K3 License along with substantial multi-GPU hosting needs. Check the inference engine and hardware requirements for the exact checkpoint before you provision anything. If your team is new to serving weights, a smaller release from the same family on one workstation teaches you the serving stack before you commit GPU capacity to a trillion-parameter checkpoint.
Are “open weight” and “open source” the same thing?
No. Open weight means the parameters are downloadable under that model’s own license, and the terms can still restrict commercial use, redistribution, or a model-as-a-service offering, so whether the license is OSI-approved is a separate question from whether the download exists. The clauses worth reading first are the ones covering commercial use, redistribution, and service offerings. Downloadable parameters do not guarantee inexpensive deployment either, so price the hardware and the serving stack for the exact release before you plan around free weights.
How much context can the current Chinese-developed families handle?
Read the route, because advertised context, input limit, and maximum output are three separate numbers. The current flagships list one-million-token context windows, some Qwen3.8 releases list 262K native context extending toward 1M, and the output ceiling differs by release and by route. Long context also adds to the bill, since you pay for the tokens you send on each request, and a window packed with generated files and old logs gives the model more to sift through than the task needs.
Can Chinese-developed families be called through one standard API instead of five separate SDKs?
Yes. A gateway can expose several models through documented API routes, so one OpenAI-compatible client can reach more than one provider while you keep a single code path. Check each model ID, modality, and API mode before you switch traffic, because every model still carries its own input types, context ceiling, and reasoning settings, and a route that accepts images for one model can be text-only for another. Point a non-critical service at the new endpoint first and confirm the model ID you send is the release you benchmarked.