Buyer's guide

Best AI coding models (2026)

Sixteen coding models that exist on their labs’ own pages, closed and self-hostable, with licences, vendor prices and benchmark results labelled by who produced them.

Last verified 6 October 2026

In short
On the only independent coding leaderboard we could read in full, GPT-6 Astra, Gemini 3.8 Flash and Claude Opus 5 tie at 74%. The best open-weight results, GLM-5.3 and Kimi K3, sit at 69%. Most newer models have no independent measurement yet, and lab-reported numbers cannot be compared with each other.
Read this first

How we rank

We did not build a composite score. A weighted blend of benchmark numbers from different labs, harnesses and effort levels would look precise and mean nothing. Instead:

  • Existence first. A model is listed only if it is on its lab’s own documentation, pricing page or model card. Names that appear in the press but not on the lab’s pages are excluded, and listed at the end.
  • Independent evidence orders the list. Models that appear on the DeepSWE v1.1 leaderboard (run by Datacurve, one harness for every model, updated 22 September 2026) are ordered by that result. Ties within the error bars are listed alphabetically, not ranked against each other. Models not on the board follow, alphabetically, and say so.
  • Self-reported numbers are shown but labelled. Each says who ran it, the date and the caveat the source gives. Terminal-Bench, FrontierCode and similar figures from different cards are not comparable. We never average, round or convert them.
  • Prices are vendor list prices read on 6 October 2026, with the source linked. “Not verified” means the vendor page was not read or does not publish one. Open-weight prices are the hosted API price, not the cost of self-hosting.
  • Not captured. We could not read the rows of the Terminal-Bench 4.0 leaderboard, or independent leaderboards for SWE-bench Verified, SWE-bench Pro, LiveCodeBench and Aider polyglot, so no independent figure from those appears. Where a lab reports one itself, it is labelled as such.
Short answers

Best for…

PickModelWhy
Best overallGPT-6 AstraTop of the independent DeepSWE board at 74% ±3%, tied within error with Gemini 3.8 Flash and Claude Opus 5. It is the most expensive per token, so overall here means best evidence, not cheapest.
Best open-weightGLM-5.3Joint-highest open-weight result on the independent board (69% ±3%, tied with Kimi K3 at 69% ±5%), under a licence whose revenue condition applies only above US$10 billion of model-as-a-service revenue.
Best for agentsClaude Opus 5.5Anthropic’s recommended default for agentic work with a 66.4% self-reported Terminal-Bench 4.0 score. The evidence is vendor-reported plus a measured predecessor, so treat this pick as weaker than the two above.
Best valueGemini 3.8 Flash (closed) and GLM-5.3-Flash (open)Gemini 3.8 Flash ties the top on DeepSWE at $2.36 average task cost and $0.75 / $3.75 per 1M tokens until 31 December 2026. GLM-5.3-Flash is 63% ±4% at $0.24 average task cost on the same board.
Best for private or air-gappedGLM-5.3 for capability, Qwen3.8-27B for modest hardwareBoth are open weights with licences that do not restrict internal use. GLM-5.3 is measured independently; Qwen3.8-27B is a self-reported 27B dense model that is far easier to host, and Devstral Small 2 is lighter still.
Picks are judgements from the evidence below, not a score. Where the evidence is thin, the pick says so.
Closed weights

Closed coding models

API-only models. Strong on capability and convenience; code leaves your infrastructure, so check data-processing terms. None of the vendor pages we read states EU regions for the API.

#ModelIndependent DeepSWE v1.1Price per 1M (in / out)ContextLicence
1Gemini 3.8 Flash74% ±1% (high effort)$0.75 / $3.751,048,576 input / 65,536 output tokensProprietary
2GPT-6 Astra74% ±3% (xhigh effort)$10 / $501,050,000 tokens (max output 128,000)Proprietary
3Claude Fable 5.1Not independently measured$10 / $501M tokens (max output 128K)Proprietary
4Claude Opus 5.5Not independently measured$4 / $201M tokens (max output 128K)Proprietary
5Claude Sonnet 5.5Not independently measured$2 / $101M tokens (max output 128K)Proprietary
6GPT-6.1 SolNot independently measured$2 / $101,050,000 tokens (max output 128K)Proprietary
7Grok 4.7Not independently measured$2 / $6500K tokensProprietary
#1 · Google · closed weights

Gemini 3.8 Flash

Best for: Best value among closed models on current evidence: tied with the top on DeepSWE at roughly half the average task cost of GPT-6 Astra.

Released: 2 September 2026 (model card). Licence: proprietary, API access only

Size: not verified. Context window: 1,048,576 input / 65,536 output tokens.

Price per 1M tokens (input / output): $0.75 / $3.75. Standard tier through 31 December 2026. From 1 January 2027: $1.50 input, $7.50 output. (vendor pricing, read 2026-10-06)

Independent result (DeepSWE v1.1): 74% ±1% (high effort).

Other benchmarks, as published by the party named:
  • DeepSWE v1.1: 73.7%. Self-reported by Google self-reported (model card), 2026-09-02. No footnote on harness or effort for coding rows. (source)
  • Terminal-Bench 4.0: 19.1%. Self-reported by Google self-reported (model card), 2026-09-02. No footnote. The same card lists Claude Opus 5 at 51.8%, so this model is far behind the leaders on that benchmark. (source)

Hosting: API only. EU regions: not verified on the pages we read.

Watch-outs:

  • The launch price doubles on 1 January 2027.
  • Weak on Terminal-Bench 4.0 per Google’s own card: strong on repository-fix tasks does not mean strong on long terminal sessions.
#2 · OpenAI · closed weights

GPT-6 Astra

Best for: The strongest independently measured result among closed models, with a 1M-token context and a $4.43 average task cost on that board.

Released: September 2026 (OpenAI page not readable by us; date from a secondary source). Licence: proprietary, API access only

Size: not verified. Context window: 1,050,000 tokens (max output 128,000).

Price per 1M tokens (input / output): $10 / $50. Prompts over 272K input tokens are billed at 2x input and 1.5x output. Batch and Flex 50%. (vendor pricing, read 2026-10-06)

Independent result (DeepSWE v1.1): 74% ±3% (xhigh effort).

Other benchmarks, as published by the party named:
  • Terminal-Bench 4.0: 57.9%. Self-reported by OpenAI-reported, as quoted on Anthropic’s Opus 5.5 page, 2026-09-22. High effort per Anthropic’s footnote. Secondary reports say 57.7%; unresolved because the OpenAI announcement page was not readable. (source)

Hosting: API only. EU regions and data-residency terms: not verified on the pages we read.

Watch-outs:

  • $10 / $50 per 1M tokens, the highest list price in this guide.
  • Statistically tied with Gemini 3.8 Flash and Claude Opus 5 on DeepSWE.
#3 · Anthropic · closed weights

Claude Fable 5.1

Best for: The most demanding reasoning and long-horizon work where Anthropic points to its top tier; for coding Anthropic recommends Opus 5.5 as the default.

Released: 1 September 2026. Licence: proprietary, API access only

Size: not verified. Context window: 1M tokens (max output 128K).

Price per 1M tokens (input / output): $10 / $50. Cache read $0.25. (vendor pricing, read 2026-10-06)

Independent result (DeepSWE v1.1): Not independently measured.

Other benchmarks, as published by the party named:
  • Terminal-Bench 4.0: 55.8%. Self-reported by Anthropic self-reported (comparison column on the Opus 5.5 page), 2026-09-22. Reported at each model’s highest effort; Fable 5.1’s effort not stated separately. (source)

Hosting: Anthropic API, Amazon Bedrock, Google Cloud and Microsoft Foundry. EU regions: not verified.

Watch-outs:

  • Same list price as GPT-6 Astra and double Opus 5.5’s.
  • The independent DeepSWE board shows its predecessor Claude Fable 5 at 70% ±3%, which is not this model.
#4 · Anthropic · closed weights

Claude Opus 5.5

Best for: Long-horizon agentic coding: Anthropic’s recommended default, and the predecessor Opus 5 is on the independent DeepSWE board at 74%.

Released: 22 September 2026. Licence: proprietary, API access only

Size: not verified. Context window: 1M tokens (max output 128K).

Price per 1M tokens (input / output): $4 / $20. Cache read $0.20. Fast mode $8 / $40. (vendor pricing, read 2026-10-06)

Independent result (DeepSWE v1.1): Not independently measured.

Other benchmarks, as published by the party named:
  • Terminal-Bench 4.0: 66.4%. Self-reported by Anthropic self-reported, 2026-09-22. xhigh effort, ±2.6 points standard error. Anthropic cites the public leaderboard for the earlier Claude Opus 5 at 51.8% (Claude Code harness) and says it reproduced 52.3%. (source)
  • FrontierCode v1.1 (Main): 54.4%. Self-reported by Anthropic self-reported, 2026-09-22. Max effort. At default (medium) effort the page text says 54.6%. (source)

Hosting: Anthropic API, Amazon Bedrock, Google Cloud and Microsoft Foundry. EU region availability: not verified on the model page.

Watch-outs:

  • Not on the independent DeepSWE board yet; the 74% belongs to the predecessor, Claude Opus 5, and is not transferable.
  • Its own self-reported numbers use different benchmarks and effort levels from other labs.
#5 · Anthropic · closed weights

Claude Sonnet 5.5

Best for: Daily coding-agent work at a fifth of Fable 5.1’s price.

Released: 28 September 2026. Licence: proprietary, API access only

Size: not verified. Context window: 1M tokens (max output 128K).

Price per 1M tokens (input / output): $2 / $10. List price on the models overview page. (vendor pricing, read 2026-10-06)

Independent result (DeepSWE v1.1): Not independently measured.

Other benchmarks, as published by the party named:
  • Terminal-Bench 4.0: 70.6%. Self-reported by Anthropic self-reported, 2026-09-28. Effort for this cell not stated. It is higher than the 66.4% Anthropic reports for Opus 5.5 at xhigh, so compare only with care. (source)
  • CursorBench 4.0: 55.5%. Self-reported by Anthropic self-reported, 2026-09-28. Effort not stated for the cell. (source)

Hosting: Anthropic API, Amazon Bedrock, Google Cloud and Microsoft Foundry. EU regions: not verified.

Watch-outs:

  • No independent measurement yet. The earlier Claude Sonnet 5 sits at 54% ±4% on DeepSWE, with a high average task cost.
  • Anthropic notes that it scores lower at max than at xhigh effort on FrontierCode because of out-of-scope edits.
#6 · OpenAI · closed weights

GPT-6.1 Sol

Best for: OpenAI says it nearly matches GPT-6 Astra at a lower cost, and its Codex docs recommend it for complex coding.

Released: 29 September 2026 (secondary source; OpenAI page not readable). Licence: proprietary, API access only

Size: not verified. Context window: 1,050,000 tokens (max output 128K).

Price per 1M tokens (input / output): $2 / $10. Cached input $0.10. (vendor pricing, read 2026-10-06)

Independent result (DeepSWE v1.1): Not independently measured.

Other benchmarks, as published by the party named: not verified

Hosting: API and Codex. EU regions: not verified.

Watch-outs:

  • No coding benchmark for it in a primary source we could read, and no independent measurement. The DeepSWE board lists gpt-5.6-sol, a different model.
#7 · xAI · closed weights

Grok 4.7

Best for: A lower-priced closed option that xAI’s models page recommends for code.

Released: not verified. Licence: proprietary, API access only

Size: not verified. Context window: 500K tokens.

Price per 1M tokens (input / output): $2 / $6. Cached input $0.50. Prompts over 200K are billed at $4 / $12 for the whole request. (vendor pricing, read 2026-10-06)

Independent result (DeepSWE v1.1): Not independently measured.

Other benchmarks, as published by the party named: not verified

Hosting: API only. EU regions: not verified.

Watch-outs:

  • No coding benchmark on a primary page and no independent measurement. The DeepSWE board lists grok-4.6, a different model, at 67% ±2%.
Open weights

Open-weight coding models you can self-host

Weights you can run on your own or EU-hosted infrastructure, so source code never leaves it. Licences differ: “open weights” does not mean unrestricted. See our guide to self-hosted models for enterprises for sizing, serving stacks and licence traps, and our sovereign cloud guide for where to run them.

#ModelIndependent DeepSWE v1.1Price per 1M (in / out)ContextLicence
1GLM-5.369% ±3% (max effort)not verifiednot verifiedGLM-5.3 License (MIT-style, custom)
2Kimi K369% ±5% (max effort)$3 / $151,048,576 tokensKimi K3 License (MIT-style, custom)
3DeepSeek V4 Pro (0813)63% ±6% (max effort)$0.66 / $1.981M tokens (API max output 384K)MIT
4GLM-5.3-Flash63% ±4% (max effort)not verifiednot verifiedMIT
5DeepSeek V4.1 FlashNot independently measured$0.15 / $0.61M tokens (API max output 384K)MIT
6Devstral Small 2 (24B)Not independently measurednot verified256K tokensApache 2.0
7Kimi K2.7 CodeNot independently measured$0.95 / $4262,144 tokensModified MIT
8Mistral Medium 3.5 (128B)Not independently measurednot verified256K tokensModified MIT (revenue cap)
9Qwen3.8-27BNot independently measurednot verified262,144 native, extensible to 1,000,000 with YaRNApache 2.0
#1 · Z.ai (Zhipu) · open weights

GLM-5.3

Best for: The strongest independently measured open-weight model, tied with Kimi K3, under a permissive licence for almost every enterprise.

Released: Hugging Face repository created 25 August 2026. Licence: GLM-5.3 License (custom, MIT-style). Companies running a “Model as a Service” business with more than US$10 billion revenue in 12 months must pass Z.AI’s security review before commercial use. (licence)

Size: About 753B parameters in the published FP8 weights; active parameters not stated on the card. Context window: not verified.

Price per 1M tokens (input / output): not verified

Independent result (DeepSWE v1.1): 69% ±3% (max effort).

Other benchmarks, as published by the party named:
  • Terminal-Bench 2.1: 88.2. Self-reported by Z.ai self-reported (model card), 2026-08. Not comparable with other labs’ Terminal-Bench figures: harness and version differ. (source)

Hosting: Weights on Hugging Face (FP8; a BF16 repo also exists). The card lists SGLang, vLLM, TokenSpeed, Transformers, KTransformers. Hardware guidance: not verified. Self-host or use a European GPU host.

Watch-outs:

  • Very large: published FP8 weights are about 753B parameters, so expect a multi-GPU node.
  • API price and context window: not verified. The card notes an emergent cyber-exploitation capability from post-training and has no safety section.
#2 · Moonshot AI · open weights

Kimi K3

Best for: Frontier-scale open weights with a 1M-token context and native low-precision release, for teams with the hardware.

Released: Hugging Face repository created 13 June 2026. Licence: Kimi K3 License (custom, MIT-style). A “Model as a Service” business above US$20 million revenue needs a separate agreement; products above 100M monthly users or US$20M monthly revenue must display “Kimi K3”. Neither applies to internal use. (licence)

Size: 2.8T total, 104B active. Context window: 1,048,576 tokens.

Price per 1M tokens (input / output): $3 / $15. Cached input $0.30. Hosted API price from Moonshot, not the cost of self-hosting. (vendor pricing, read 2026-10-06)

Independent result (DeepSWE v1.1): 69% ±5% (max effort).

Other benchmarks, as published by the party named:
  • DeepSWE v1.1: 67.5. Self-reported by Moonshot self-reported, 2026-09. Max effort, Kimi Code harness. Moonshot’s card says the official leaderboard shows 67.3 on the mini-SWE-agent harness; the page we read on 6 October shows 69% ±5%. (source)
  • Terminal-Bench 2.1: 88.3. Self-reported by Moonshot self-reported, 2026-09. Kimi Code harness; other models’ scores on the card are best across harnesses. (source)

Hosting: Native MXFP4 weights with MXFP8 activations; the card recommends vLLM, SGLang and TokenSpeed. Minimum hardware: not verified. Weights are 2.8T parameters.

Watch-outs:

  • Needs a large multi-GPU cluster to self-host.
  • The licence has conditions that matter if you resell model access.
#3 · DeepSeek · open weights

DeepSeek V4 Pro (0813)

Best for: MIT-licensed weights and a very low hosted price, with a 1M-token context.

Released: 13 August 2026 (DeepSeek API change log). Licence: MIT (licence)

Size: About 1.65T parameters (Hugging Face metadata); active parameters not verified. Context window: 1M tokens (API max output 384K).

Price per 1M tokens (input / output): $0.66 / $1.98. Off-peak rate; peak hours (Monday to Friday 01:00 to 04:00 and 06:00 to 10:00 UTC) cost 2x. Cached input $0.022. (vendor pricing, read 2026-10-06)

Independent result (DeepSWE v1.1): 63% ±6% (max effort). The board names its row deepseek-v4-pro; we did not confirm that it is the 0813 build.

Other benchmarks, as published by the party named:
  • Terminal Bench 2.1: 87.9. Self-reported by DeepSeek self-reported, 2026-08-13. Minimal mode of the DeepSeek harness, max reasoning effort. (source)
  • DeepSWE: 62.7. Self-reported by DeepSeek self-reported, 2026-08-13. Same harness. The card lists Kimi K3 at 67.5 and Claude Fable 5 at 70.0 under the same method. (source)

Hosting: The model card gives a vLLM example on a single 4-GPU GB300 node and points to vLLM and SGLang recipes for other hardware.

Watch-outs:

  • Trails Kimi K3 and GLM-5.3 on the independent board.
  • Check the hosted API’s data-processing terms before sending code; self-hosting avoids the question.
#4 · Z.ai (Zhipu) · open weights

GLM-5.3-Flash

Best for: Best value among open weights on current evidence: 63% ±4% on DeepSWE at an average task cost of $0.24 on that board.

Released: Hugging Face repository created 25 August 2026. Licence: MIT (licence)

Size: 320B total, 18B active. Context window: not verified.

Price per 1M tokens (input / output): not verified

Independent result (DeepSWE v1.1): 63% ±4% (max effort).

Other benchmarks, as published by the party named: not verified

Hosting: Weights on Hugging Face; the card lists SGLang, vLLM, TokenSpeed, Transformers, KTransformers. Hardware guidance: not verified.

Watch-outs:

  • Z.ai’s claim of approaching Claude Opus 4.8 is a lab claim. On the independent board this model scores 63% ±4% and Opus 4.8 scores 59% ±2%.
  • API price and context window: not verified.
#5 · DeepSeek · open weights

DeepSeek V4.1 Flash

Best for: The cheapest strong open-weight option to self-host: small active-parameter count and MIT weights.

Released: 10 September 2026 (DeepSeek API change log). Licence: MIT (licence)

Size: 552B backbone parameters; 8B active per token at prefill and 16B at decode (model card). Context window: 1M tokens (API max output 384K).

Price per 1M tokens (input / output): $0.15 / $0.6. Off-peak; peak hours are 2x ($0.30 / $1.20). Cached input $0.003. (vendor pricing, read 2026-10-06)

Independent result (DeepSWE v1.1): Not independently measured.

Other benchmarks, as published by the party named:
  • DeepSWE v1.1: 74.2. Self-reported by DeepSeek self-reported, 2026-09-10. mini-SWE harness, 8 samples per task. DeepSeek’s own table gives 72.6 on its minimal harness, 69.8 with the Claude Code harness and 65.6 with Codex, which shows how much the harness matters. (source)
  • Terminal-Bench 4.0: 31.2. Self-reported by DeepSeek self-reported, 2026-09-10. Same table lists Claude Opus 5 at 51.8, so it is well behind the leaders there. (source)

Hosting: Weights on Hugging Face with an inference folder; the card points to DeepSeek’s own recipe. Minimum hardware: not verified.

Watch-outs:

  • Not on the independent board (its row is the earlier deepseek-v4-flash at 53% ±4%, a different model), so the 74.2 is a self-report only.
#6 · Mistral AI · open weights

Devstral Small 2 (24B)

Best for: Air-gapped or single-GPU coding assistance: the card says it runs on a single RTX 4090 or a Mac with 32GB RAM.

Released: Hugging Face repository created 28 November 2025. Licence: Apache 2.0 (licence)

Size: 24B dense. Context window: 256K tokens.

Price per 1M tokens (input / output): not verified

Independent result (DeepSWE v1.1): Not independently measured.

Other benchmarks, as published by the party named:
  • SWE-bench Verified: 68.0%. Self-reported by Mistral self-reported (model card), 2025-11. Harness not stated on the card text we read. (source)
  • Terminal Bench 2: 22.5%. Self-reported by Mistral self-reported (model card), 2025-11. Older benchmark version than the 2026 figures above. (source)

Hosting: vLLM, SGLang, llama.cpp, LM Studio and Ollama are named on the card.

Watch-outs:

  • An older model than the others here, and well behind on harder agentic benchmarks.
#7 · Moonshot AI · open weights

Kimi K2.7 Code

Best for: A coding-specific open model with a low hosted price, for teams already standardised on Moonshot.

Released: Hugging Face repository created 11 June 2026. Licence: Modified MIT (licence)

Size: About 1T total; active not verified. Context window: 262,144 tokens.

Price per 1M tokens (input / output): $0.95 / $4. Cached input $0.19. A high-speed variant is $1.90 / $8.00. (vendor pricing, read 2026-10-06)

Independent result (DeepSWE v1.1): Not independently measured.

Other benchmarks, as published by the party named: not verified

Hosting: Weights on Hugging Face. Serving stack and hardware: not verified.

Watch-outs:

  • We could not read benchmark results for it, and it is not on the independent board.
#8 · Mistral AI · open weights

Mistral Medium 3.5 (128B)

Best for: Smaller companies and individuals wanting one instruct, reasoning and coding model with broad multilingual coverage (24 languages listed on the card).

Released: Hugging Face repository created 31 March 2026. Licence: Modified MIT. Not usable under the licence if your company’s global consolidated monthly revenue exceeded US$20 million in the preceding month; larger companies need a commercial licence from Mistral. (licence)

Size: 128B dense. Context window: 256K tokens.

Price per 1M tokens (input / output): not verified

Independent result (DeepSWE v1.1): Not independently measured.

Other benchmarks, as published by the party named:
  • SWE-bench Verified: 77.6%. Self-reported by Mistral self-reported (model card), 2026-03. Harness not stated on the card text we read. (source)

Hosting: The card shows vLLM with tensor parallelism 8 and SGLang images for Hopper and Blackwell; llama.cpp via community GGUF. Exact VRAM: not verified.

Watch-outs:

  • The revenue cap in the licence rules it out for most enterprises without a separate deal.
#9 · Alibaba Qwen · open weights

Qwen3.8-27B

Best for: Hardware-light private deployment: dense 27B under Apache 2.0, with an official FP8 repository.

Released: Hugging Face repository created 5 August 2026. Licence: Apache 2.0 (licence)

Size: 27B dense. Context window: 262,144 native, extensible to 1,000,000 with YaRN.

Price per 1M tokens (input / output): not verified

Independent result (DeepSWE v1.1): Not independently measured.

Other benchmarks, as published by the party named:
  • SWE-bench Pro: 61.7. Self-reported by Alibaba Qwen self-reported (model card), 2026-08. Card compares it with Qwen3.6-27B at 53.5. Harness not verified. (source)
  • Terminal Bench 2.1: 73.0. Self-reported by Alibaba Qwen self-reported (model card), 2026-08. Not comparable with other labs’ figures. (source)

Hosting: Official FP8 repo; the card names Transformers, vLLM, SGLang and TokenSpeed. Hardware guidance: not verified.

Watch-outs:

  • Not on the independent board. Larger Qwen3.8 models carry custom licences with revenue conditions.
Independent evidence

The independent leaderboard in full

DeepSWE v1.1 is run by Datacurve on 113 tasks written from scratch, using one harness (mini-swe-agent) for every model. The table is the full board as read on 6 October 2026, including models that are not in this guide, copied unedited. Scores are Pass@1 with the reported spread; models within each other’s spread are not meaningfully ordered. Reasoning effort differs per row. Source.

RankModel (as named on the board)Pass@1Average costEffort
1gpt-6-astra74% ±3%$4.43xhigh
2gemini-3.8-flash74% ±1%$2.36high
3claude-opus-574% ±4%$11.84max
4gpt-5.6-sol73% ±3%$6.46max
5claude-fable-570% ±3%$13.41xhigh
6glm-5.369% ±3%$3.99max
7kimi-k369% ±5%$4.65max
8grok-4.667% ±2%$3.45medium
9gpt-5.6-luna67% ±4%$0.61max
10gpt-5.567% ±6%$7.23xhigh
11gemini-3.7-flash65% ±3%$2.03medium
12glm-5.3-flash63% ±4%$0.24max
13deepseek-v4-pro63% ±6%$1.67max
14claude-opus-4.859% ±2%$13.22max
15qwen3.8-max57% ±3%$3.73xhigh
16muse-spark-1.255% ±2%$3.70xhigh
17claude-sonnet-554% ±4%$26.40max
18deepseek-v4-flash53% ±4%$0.46max
19gemini-3.6-flash47% ±4%$2.21high
20glm-5.244% ±2%$3.92max
21gemini-3.5-flash36% ±4%$3.45high
Rows such as claude-opus-5, claude-fable-5, gpt-5.6-sol and grok-4.6 are earlier models than the ones named in this guide. We do not transfer their scores to their successors.
Checklist

How to evaluate on your own codebase

  • Pull 20 to 30 real tasks from your own repositories: bug fixes with a failing test, small features, one refactor, and the three tasks that went badly last quarter.
  • Freeze the harness. Use the same agent loop, tools, context limit and timeout for every model, and record the reasoning effort. Differences in harness moved DeepSeek’s own DeepSWE result by nearly nine points.
  • Run each task more than once. Report pass rate with its spread, not a single run. The independent board shows spreads of up to six points.
  • Score with your tests, not with a model’s opinion. Then read five passing diffs: passing tests can hide an unwanted change.
  • Measure cost per solved task, not price per token. Average task cost on the independent board ranges from $0.24 to $26.40.
  • Test privacy and licence early: where code is sent, retention terms, and whether an open-weight licence limits your revenue or use.
  • For open weights, test on the hardware you will actually run, at the quantisation you will actually run. Published scores are usually at full or lab-chosen precision.
  • Re-run when a model version changes. Names such as “flash”, “pro” and “sol” are reused across releases.

Swfte can help you run this: see open-source model testing, route the candidates through Connect to compare them with one interface, and use the LLM API price index for current prices.

Left out

Models we considered and left out

  • Gemini 3.2 Pro, DeepSeek V4.5, Llama 5. Never shipped. Not on the labs’ model lists. Previously published on this site in error and removed.
  • Claude Opus 4.8. Exists on Anthropic’s docs but is superseded by Opus 5.5. It appears only as a reference row on the DeepSWE board (59% ±2%).
  • GPT-5.5, GPT-5.3 Codex, GPT-6 Sol. GPT-5.5 retires from ChatGPT and Codex on 14 October 2026. No OpenAI page listed GPT-5.3 Codex, and GPT-6 Sol is named only as “when available”.
  • Qwen3.8-2.4T-A95B, Qwen3.8-Flash-Next. Exist on Hugging Face under custom Qwen licences with revenue conditions. Left out because we could not confirm the mapping to the benchmark rows or read the benchmark tables.
  • Qwen 3.7. No open weights found; only a hosted “plus” name appears.
  • MiniMax M3. Exists, but its community licence is non-commercial by default.
  • Llama 4. Model cards are gated, so we could not verify the licence terms.
  • gpt-oss, Gemma 4, Z.ai and MiniMax coding variants, Codestral. Not checked for coding evidence in this pass. Gemma 4 and gpt-oss are covered in the self-hosted models guide.

Common questions

What is the best AI model for coding in 2026?
On the independent DeepSWE v1.1 leaderboard, GPT-6 Astra, Gemini 3.8 Flash and Claude Opus 5 are tied within error at 74%. The newest Claude and GPT releases are not yet on it. There is no single winner: it depends on your language, your agent harness and what a solved task costs you.
What is the best open-weight coding model I can self-host?
GLM-5.3 and Kimi K3 are the joint-highest open-weight results on the independent board (69%). Both are very large. For modest hardware, Qwen3.8-27B (Apache 2.0) and Devstral Small 2 (Apache 2.0) are far easier to host, but only have lab-reported results.
Can I trust the benchmark numbers on model cards?
Treat them as vendor claims. Each lab picks its own harness, effort level and benchmark version, so numbers from different cards are not comparable. This page labels every figure as independent or self-reported, and never averages or converts them.
Is a self-hosted coding model private enough for source code?
Self-hosting keeps prompts and code on infrastructure you control, which closes the main exposure of sending code to a third-party API. It does not remove other risks: licence terms, supply-chain checks on the weights, and access control on the serving stack still apply.
Why are some models missing, such as Gemini 3.2 Pro, DeepSeek V4.5 or Llama 5?
They do not exist: we checked the labs’ own pages and found no such releases. This site once published articles about them and those were removed. Every model here was confirmed on the lab’s own documentation or model card.
Are there open-weight coding models usable commercially?
Yes, with checks. DeepSeek V4.1 Flash and V4 Pro are MIT, Qwen3.8-27B and Devstral Small 2 are Apache 2.0, GLM-5.3 and Kimi K3 have MIT-style licences with revenue conditions aimed at companies reselling model access, and Mistral Medium 3.5 is barred above US$20 million monthly revenue. Read the licence for the exact version you deploy.
Where can I host these models in the EU?
Open-weight models can run on any EU infrastructure you choose; see our guide to sovereign cloud providers. For closed models, none of the vendor pages we read stated EU regions for the API, so confirm data-residency terms with the vendor before sending code.

Deploy a model with Swfte Connect

One gateway, every provider, per-token cost visibility. Swap models without touching your code.