Skip to content
Gateway Architecture

Which LLM Should You Use for Coding?

POSTED ON UPDATED ON 15 min read MixRoute

Which LLM Should You Use for Coding?
Which LLM Should You Use for Coding?

A leaderboard-topping coding model does not automatically transfer to your repository. What decides the choice is three things: the task (completion, debugging, repository edits, or a long-running agent), the deployment (a hosted API or your own GPUs), and the budget. In the September 2026 snapshots, the reported proprietary leaders are Claude Fable 5.1 at a Coding Index of 81.6, GPT-5.6 Sol at 78.3 under xhigh effort, and Claude Opus 5 at 78.0, while Kimi K3 (max) leads the open-weight entries at 76.2. Those numbers say nothing about your harness, so use them to build a shortlist and then run one real task from your own backlog before you commit.

Read a benchmark score the way you read a resume: useful for a first cut, but the decision happens when the candidate works on your own code. Use the dated sources below to narrow the field, then run a small private evaluation on a task from your repository before you point real traffic at one model.

Which coding model fits your deployment constraint?

Deployment decides which candidates you can use at all, so settle it before you compare scores. If running GPUs is not the work you want to take on, buy hosted access and let the provider run the hardware. If code has to stay inside your network, or you want control over how models are served, self-hosting becomes part of the job.

Self-hosting trades a per-token invoice for utilization, hardware, and electricity: what each generated token really costs depends on how hard you run the box. Weight size on its own does not turn into a predictable bill, so judge a candidate by the exact checkpoint, the quantization, the context setting, and the serving runtime.

Choose the deployment first; Provider operates hardware; verify pricing and data terms.; Check quantization, runtime and context memory.; Plan infrastructure and the exact model license.
Selection guidance from the article; no measured winner is implied.

When infrastructure is not your task

Hosted access removes hardware planning and hands you three constraints instead: per-token pricing, provider rate limits, and whatever data policy the provider publishes. Which model you want depends on the loop you are in. Interactive completion and debugging reward a fast model, while repository edits and agentic work want stronger reasoning and a longer output budget.

MixRoute does not change which model is best. What it changes is the surface you compare on: one OpenAI-compatible endpoint reaches models from several providers, so you test against documented model IDs instead of a single vendor’s console.

For work that stays inside the editor, the loop matters as much as the checkpoint. Read the practical guide to keeping a vibe coding loop productive before you spend on hardware or API credits, and log reasoning effort and tool settings as evaluation variables next to model choice. Both move results, and neither should go unrecorded.

Self-host with a checkpoint, not a family name

The starting point is a specific artifact you can download. Qwen3.8-27B is a good example of what to verify: its official model card, as of September 7, 2026, lists an Apache-2.0 release, 27 billion parameters in a causal language model with a vision encoder for native image and video understanding, a 262,144-token native context extensible to roughly 1 million tokens with RoPE scaling, thinking on by default, and a reasoning_effort field that accepts xhigh, medium, or low. A hosted Qwen Cloud variant with a 1-million-token default and built-in tools is a separate offering the card lists as coming soon, so confirm which checkpoint actually serves your request before you plan around it.

Memory tracks precision and context, not parameter count. Pinggy’s deployment notes, published August 2026, estimate the Qwen3.8-27B weights at about 14 to 17 GB in 4-bit, roughly 28 GB at FP8, and about 56 GB at BF16. Those are published estimates, not a MixRoute measurement, and they do not guarantee any particular GPU. The KV cache adds more on top at long context, so check memory at the context length you actually intend to run rather than assuming a card will fit.

How do proprietary coding models compare on benchmarks?

Reported scores are shortlist evidence, and the leader changes with the benchmark. In the WhatLLM coding dataset, snapshot dated September 7, 2026, the reported Coding Index leaders are Claude Fable 5.1 at 81.6, GPT-5.6 Sol at 78.3 under the xhigh effort setting, and Claude Opus 5 at 78.0. Those underlying scores come from Artificial Analysis data whose source measurement date is not supplied, so read them as a dated snapshot rather than a live verdict.

Keep benchmark conditions attached; Keep the benchmark version and date.; Record reasoning effort, harness and fallback.; A reported index starts a shortlist.
Selection guidance from the article; no measured winner is implied.

What the reported leaders show

Every number carries an effort setting and a fallback configuration, and both change what it means. Claude Fable 5.1 is reported with adaptive thinking, max effort, and a default fallback; GPT-5.6 Sol is reported at xhigh effort. A fallback can mask model behavior during evaluation, so a fallback-enabled score is not the same measurement as a run that never fell back.

Terminal work gets its own version of the test. The WhatLLM page, snapshot dated September 17, 2026, names GPT-5.6 Sol at max effort as its terminal-work candidate, with a reported Terminal-Bench Hard score of 66.0 percent, the highest among models with a result for that exact benchmark. That is a different effort setting from the xhigh Sol entry above, which the ranking table lists at 61 percent on Terminal-Bench Hard, and Claude Fable 5.1 has no Terminal-Bench result in the same table. Coverage is incomplete, so the result does not establish a winner for other tasks or for other Terminal-Bench versions.

For repository-scale edits, the number to look at is SWE-bench Pro. Claude Opus 4.8 appears with a reported vendor benchmark score of 69.2 in Pinggy’s comparison notes and in Onyx’s side-by-side table, last updated July 20, 2026. Seeing the same value in several tables guards against transcription errors; it is not independent replication. The figure is vendor-reported, so it cannot establish a universal ranking.

Reported model and effort Metric Reported value Source snapshot
Claude Fable 5.1 (default fallback) Coding Index 81.6 WhatLLM, read 2026-09-07, Artificial Analysis reported data
GPT-5.6 Sol (xhigh) Coding Index 78.3 WhatLLM, read 2026-09-07
Claude Opus 5 (max effort) Coding Index 78.0 WhatLLM, read 2026-09-07
GPT-5.6 Sol (max) Terminal-Bench Hard 66.0% WhatLLM terminal candidate, read 2026-09-17
Claude Opus 4.8 SWE-bench Pro 69.2 Vendor-reported figure listed in Pinggy and Onyx tables

Before choosing between the top two reported leaders, read the dedicated GPT-5.6 Sol versus Claude Fable 5 comparison, which covers the differences a single index number hides, including reasoning effort behavior and fallback handling.

What are the best open-weight models for self-hosting?

Kimi K3, Qwen3.8-27B, and GLM-5.3 are three separate downloadable checkpoints, and each sits in a different hardware and licensing position. Which one you can run comes down to whether you have a cluster, a single GPU, or something in between. A hosted product name such as Qwen3.8-Max does not tell you what the downloaded weights can do.

Kimi K3: highest reported open-weight index, cluster required

At 76.2 on the Coding Index in the WhatLLM dataset snapshot dated September 7, 2026, Kimi K3 (max) is the highest open-weight entry in that dataset. That is not a universal open-weight win: the index does not cover every task, and the deployment cost is steep. The official Kimi K3 model card, September 2026, reports a 2.8-trillion-parameter model with 104 billion active parameters, 896 experts with 16 active per token, a 1-million-token context, native vision through MoonViT-V2, MXFP4 weights from quantization-aware training, and a bespoke Kimi K3 License. Pinggy’s August 2026 notes estimate about 1.4 TB of resident weights and report Moonshot’s guidance of 64 or more accelerators, so if you are sizing a single box, this is not the checkpoint you are sizing for. Weight size alone does not guarantee that any hardware fits, and the custom license needs a separate agreement for some resale use cases.

Qwen3.8-27B: the single-GPU class

For one GPU, Qwen3.8-27B is the practical local option. The official card reports 73.0 on Terminal-Bench 2.1, 61.7 on SWE-bench Pro, 90.3 on LiveCodeBench v6, and 42.2 on DeepSWE 1.1, all Alibaba-reported under the settings documented in the card’s footnotes. Watch the harness footnote: the card names the Claude Code harness for the SWE-bench Pro and DeepSWE rows, and that footnote does not carry over to LiveCodeBench v6. Apache-2.0 licensing plus a Pinggy-estimated 14 to 17 GB footprint at 4-bit puts it first on the list when the deployment is a single 24 GB card or a 32 GB Mac.

GLM-5.3: separate model, separate license

This one shares the GLM-5.2 base model, and its official model card, September 2026, attributes every gain to post-training. The card reports a 50 percent improvement over GLM-5.2 on Z.ai’s in-house code bench, open-weights state-of-the-art claims on Terminal-Bench 3.0 and Agents’ Last Exam, and a Terminal-Bench 2.1 score of 88.2 in its vendor comparison table. Those are vendor-reported results under named harnesses and reasoning settings, so they belong on a shortlist, not in a universal claim.

Evaluation context in that card changes by task. The footnotes used a 1-million-token window for specified evaluations such as NL2Repo and Agents’ Last Exam, 400K for DeepSWE, and 300K for HLE with tools. Those are evaluation settings, not a universal hosted context limit, so check the serving configuration you actually intend to run.

GLM-5.3 downloads under its own GLM-5.3 license, which is not MIT, and it is not the same artifact as GLM-5.3-Flash. Do not infer the license, context length, or hardware footprint of one from the other. Check GLM-5.3’s official deployment guide for the precision, context, and serving setup it requires. To see how the previous-generation GLM-5.2 lines up against leading closed models, read the GLM-5.2 versus Opus 4.8 versus Fable 5 comparison before you pick a family.

Checkpoint Reported size Context License Deployment notes
Qwen3.8-27B 27B dense with vision encoder 262,144 native, extensible to about 1M via RoPE Apache-2.0 Pinggy estimate: about 14 to 17 GB weights at 4-bit, before KV cache
Kimi K3 2.8T total, 104B active 1M tokens Kimi K3 License (custom) About 1.4 TB resident at MXFP4; Moonshot guidance of 64+ accelerators, per Pinggy
GLM-5.3 Not stated in the captured official card; same base model as GLM-5.2 Evaluation context varies: 1M for NL2Repo and Agents’ Last Exam, 400K for DeepSWE, 300K for HLE with tools GLM-5.3 License (custom) Official card lists SGLang, vLLM, TokenSpeed, Transformers, KTransformers, Unsloth, and Ascend paths

How much does coding API usage cost per million tokens?

Every provider bills per 1 million tokens, prices input separately from output, and charges a higher rate for output. Coding widens that gap, because an agent produces output tokens across many turns and retries. Treat input tokens and output tokens as two different quantities and multiply each by its own rate.

Cost arithmetic with an illustrative rate

Assume 2 USD per million input tokens and 10 USD per million output tokens for a hypothetical calculation. A workload that consumes 80 million input tokens and 20 million output tokens costs 160 USD for input plus 200 USD for output, or 360 USD total. Those are illustrative rates, not provider quotes. Real rates vary by model, provider, cache-hit status, and date, and the token total has to include retries and every agent turn rather than just the first prompt and the final answer. Whether a rejected or failed request is billed also varies by provider, so check the billing policy before you budget.

When you are ready to turn that estimate into actual spend, compare current MixRoute rates for the models on your shortlist before you fund an account. Cached-input and output prices differ even within one model family, so read the rate table for the exact checkpoint and context mode you plan to call.

Where reported prices come from

Published price tables go stale. Pinggy’s August 2026 notes, for example, report the Kimi K3 API at 3.00 USD per million input tokens, 0.30 USD per million cache-hit input tokens, and 15.00 USD per million output tokens. That is a dated observation from a third-party page, not a current provider quote. Read the provider’s own pricing documentation for the model version and context you intend to run, and treat a leaderboard’s price column as a starting point only.

How to run a reproducible coding evaluation?

A benchmark score tells you what a model did on someone else’s tasks under someone else’s harness. A reproducible evaluation replaces that with a result you can defend: the same repository, the same starting commit, the same task, the same harness, the same model settings, the same grader, and a written record of tokens and cost. Without those pins, a comparison drifts into measuring the harness or the prompt instead of the model.

Separate model tests from bundle tests; Same starting commit, prompt and grader.; Use compatible shared settings for model comparisons.; Count tokens, cost, retries and accepted changes.
Selection guidance from the article; no measured winner is implied.

Pin the variables that change the outcome

Start with one issue or failing test from your own backlog and record its starting commit. For an isolated model comparison, use one compatible harness, one settings block that every candidate supports, and the same task, prompt, and grader for each model. Published benchmark harnesses are provenance, not requirements: Anthropic and GLM rows in vendor cards run under Claude Code, OpenAI numbers run under Codex, and Kimi K3 runs under the Kimi Code harness in its official card. Those names describe how each published number was produced. If you want to compare native bundles, such as a vendor’s agent paired with its own model, label that a bundle comparison and do not attribute the difference to the model alone.

Log these for every run, because each one moves both the score and the token count:

  • the reasoning_effort value
  • temperature and top-p
  • maximum output tokens
  • context window
  • timeout

Record tokens, cost, and review, not just pass or fail

Capture the reported usage fields from each API response: input tokens and output tokens per call, summed across every agent turn and retry. Multiply those totals by the rate you are actually quoted to get cost per accepted result. Decide what counts as accepted before you start, and have a human review the diff. A model that passes the test but rewrites unrelated code is not a clean pass, and no automated score captures that.

Before switching between providers, or between a hosted API and a local runtime, check the exposed model IDs and the supported API modes in the current model documentation for each side. MixRoute, for example, lists IDs such as claude-fable-5-1 and gpt-5.6-sol in its model reference, so a comparison should call the exact ID that will serve the request rather than a shorthand name that could resolve to a different checkpoint. Run the comparison yourself, because public rankings do not measure your agent harness, your review standards, or the cost of a failed change.

FAQ

What is the best proprietary LLM for coding?

No single model wins, and the ranking depends on which benchmark and which effort setting you read. In the WhatLLM snapshot dated September 7, 2026, Claude Fable 5.1 reports a Coding Index of 81.6 with a default fallback, GPT-5.6 Sol reports 78.3 at xhigh effort, and Claude Opus 5 reports 78.0. Then pick by your own loop rather than the top row: interactive completion rewards a fast model, agentic repository work rewards reasoning depth and a long output budget, and the token bill tracks the number of turns, not the gap between two index scores.

Which open-weight model has the highest coding index?

Kimi K3 (max) leads the open-weight entries in the WhatLLM dataset at 76.2, but leading one dataset does not make it the right download. It needs cluster-class hardware: an estimated 1.4 TB of resident weights at MXFP4, with Moonshot guidance of 64 or more accelerators. From a single GPU, Qwen3.8-27B is the realistic trial candidate at an estimated 14 to 17 GB in 4-bit, assuming the context you need fits. GLM-5.3 is a separate checkpoint with its own license and its own serving requirements.

What hardware runs Qwen3.8-27B for coding?

A single 24 GB GPU or a 32 GB Mac is the practical starting point. Pinggy’s notes dated August 2026 estimate about 14 to 17 GB for the weights at 4-bit before the KV cache, with an 18 GB Ollama download, and those are estimates rather than measurements. Check memory at the context length you intend to run, because the KV cache grows with context and the card that fits a short prompt may not fit a long one. If infrastructure is not the work you want to take on, hosted access skips the sizing question.

How much does Kimi K3 cost via API?

Check the provider’s own pricing page for the version you will call, because cached and uncached input bill at different rates and rates move. Pinggy’s August 2026 coverage reported the Kimi K3 API at 3.00 USD per million input tokens, 0.30 USD per million cache-hit input tokens, and 15.00 USD per million output tokens. Budget by turns rather than by prompt, since a coding agent’s total is the sum of every turn and retry, and a 15.00 USD output rate makes loops that keep regenerating the same file expensive.

What benchmark shows terminal coding ability?

Terminal-Bench, and only within the same version. In the WhatLLM coding dataset, snapshot dated September 17, 2026, GPT-5.6 Sol at max effort reports a Terminal-Bench Hard score of 66.0 percent, the highest among models with a result for that exact benchmark, while Claude Fable 5.1 has no Terminal-Bench result there. The xhigh Sol entry sits at 61 percent on Hard, so the effort setting is part of the number. Terminal-Bench 2.1 and Terminal-Bench 3.0 are different tests with different difficulty.

Can I self-host GLM-5.3 on one node?

Not from the card alone. GLM-5.3’s official deployment guide is what you size from, and the card version checked in September 2026 does not state a one-node memory footprint. It does list SGLang, vLLM, TokenSpeed, Transformers, KTransformers, Unsloth, and Ascend NPU paths. Its evaluations ran at context windows from 300K to 1M tokens, so settle on the configuration you mean to serve first. GLM-5.3 uses its own custom license, which is not MIT, and GLM-5.3-Flash is a separate checkpoint with separate requirements.


Scan to share
Scan to share
Gateway Architecture Chinese LLMs Compared: Qwen, DeepSeek, Kimi, GLM, MiniMax and ERNIE MixRoute 19 min read Gateway Architecture Codex vs Claude Code: Which Coding Workflow Fits You? MixRoute 20 min read Gateway Architecture Free AI APIs Compared: Recurring Free Tiers, Trial Credits, and Local Models MixRoute 17 min read