Selection guide
Best LLM for coding: choose by situation, not by a leaderboard
Choose a coding model by your situation (private repo, regulated, cost-bound, offline) using vendor-stated facts and one independent benchmark, read on 2026-10-07.
The best LLM for coding depends on where your code may go, how long the tasks run and what you can spend, so this page is a selection guide and not a ranking. Start from your situation in the first table, shortlist from the hosted and open-weight tables, then test on your own repository. For a ranked list with cited results, see our separate best AI coding models page.
Last verified 2026-10-07. Sources are listed at the end of the page.
Which LLM should I use for coding in my situation?
Start with the constraint that cannot move. The model comes after it.
| Situation | Start with | What to check |
|---|---|---|
| Private repository | Open weights you host, or a hosted model whose data terms you have read. | Self-hosting keeps code inside your network. For a hosted API, read the provider’s data-processing and retention terms before sending code. Region statements were not on every model page read for this guide. |
| Regulated environment | Self-hosted open weights, or a hosted model on a cloud you already govern. | Ask for the licence text, the data-handling terms and where inference runs. Several hosted models are offered through more than one cloud, so your existing contract may already cover one. |
| Cost-bound team | A smaller open-weight model you host, or a hosted model from the lower-priced end of a vendor’s lineup. | Compare cost per finished task, not per token. A cheaper model that needs three attempts can cost more. Read each vendor’s current pricing page. |
| Offline or air-gapped | An open-weight model small enough for your hardware. | Devstral Small 2’s card says it can run on a single RTX 4090 or a Mac with 32 GB RAM. Confirm the licence permits your use and mirror the weights internally. |
| Very long, multi-step tasks | A hosted model with a long context window and a documented agentic focus. | Vendor pages below describe agentic coding focus and 1M-token context. The agent harness around the model matters as much as the model. |
Which model suits editor completion, agents, review and repository questions?
| Coding task | What decides the choice | How to test it |
|---|---|---|
| Editor completion | Latency and cost per call, since completions run on every pause. Check that the model card documents fill-in-the-middle; the cards read for this page were not checked for it, so no completion model is named. | Measure time to first token and acceptance rate on your own files for a week. |
| Agentic coding (the model edits, runs tests, iterates) | The model, its effort setting and the harness together. Vendors describe long-horizon agentic coding as the focus of Claude Opus 5.5 and Gemini 3.8 Flash, and GPT-6.1 Sol for complex coding. | Run the same 20 tasks from your tracker through each model in the same harness, with tests as the judge. |
| Code review | Few false alarms and the ability to read the whole change with its context. | Replay past pull requests where you already know the defects, and count what is caught and what is wrongly flagged. |
| Long-context repository questions | The documented context window, then how well the model uses it. Windows below run from 256K to about 1M tokens. | Ask questions whose answer sits deep in a large file set and check the answers by hand. |
| Local or private coding | Memory budget, licence and runtime support. See the open-weight table and the best local LLM guide. | Run the model on the machine your developers use and watch memory under a long context. |
Which hosted models do the vendors position for coding?
Facts are from each vendor’s own page on 2026-10-07. Prices are not shown because they change; follow the pricing links in the sources.
| Model | Model ID as the vendor lists it | Context window | What the vendor page says |
|---|---|---|---|
| Claude Opus 5.5 (Anthropic) | claude-opus-5-5 | 1M tokens | Built for long-running agentic coding and knowledge work. Available on the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS. |
| Claude Sonnet 5.5 (Anthropic) | claude-sonnet-5-5 | 1M tokens | Anthropic describes it as the best combination of speed and intelligence in its lineup. The same platforms are listed. |
| GPT-6.1 Sol (OpenAI) | gpt-6.1-sol | 1,050,000 tokens | Near-Astra performance at a lower cost for complex coding, computer use and professional work. OpenAI’s Codex model page recommends it for repeated, long-running work. |
| Gemini 3.8 Flash (Google) | gemini-3.8-flash | Up to 1M input tokens | Described as engineered for long-horizon software engineering, autonomous agents and complex enterprise workflows. |
GPT-6 Astra is the model OpenAI calls its most capable on the same Codex page and also appears on the independent board below. Vendor descriptions are positioning, not measurements.
Which open-weight models can I host for coding?
Licence names and runtimes are as the model cards state them. Read the licence file of the exact repository.
| Model | Licence as named | Context window | Runtimes the card names and notes |
|---|---|---|---|
| Devstral Small 2 (24B) | Apache 2.0 | 256K | vLLM (recommended), llama.cpp, Ollama, LM Studio, SGLang, Transformers. Runs on one RTX 4090 or a 32 GB Mac per the card. |
| Qwen3.8-27B | Apache 2.0 | 262,144 native, extensible to 1,000,000 | vLLM, SGLang and TokenSpeed; the card also points to Ollama, llama.cpp and LM Studio quantisations. A vision-language model with thinking on by default. |
| DeepSeek V4.1 Flash | MIT License | Up to 1M tokens | Transformers, vLLM, SGLang and Docker Model Runner. 552B backbone parameters, so plan for a multi-GPU node even though 8B to 16B are active per token. |
| GLM-5.3 | GLM-5.3 License | Not plainly stated in the card text read | vLLM, SGLang, TokenSpeed, Transformers, KTransformers and Unsloth. The licence requires a security review by Z.AI for a Model as a Service business above 10 billion US dollars in revenue. |
| Kimi K3 | Kimi K3 License | 1,048,576 tokens | vLLM, SGLang and TokenSpeed. 2.8T total and 104B active parameters, so it needs a large cluster. The card text read did not detail licence conditions, so read the licence file. |
See best self-hosted models for enterprises for hardware notes and licence traps across more models.
What does an independent coding benchmark show?
DeepSWE v1.1, run by Datacurve, has 113 tasks over 91 repositories in five languages. Every model runs in the mini-swe-agent harness, and the leaderboard was last updated on 2026-09-22. We read it on 2026-10-07. Rows below are listed alphabetically, not by score.
| Board row | Pass@1 | Average cost as the board lists it | Effort |
|---|---|---|---|
| claude-opus-5 | 74% ±4% | $11.84 | max |
| deepseek-v4-pro | 63% ±6% | $1.67 | max |
| gemini-3.8-flash | 74% ±1% | $2.36 | high |
| glm-5.3 | 69% ±3% | $3.99 | max |
| glm-5.3-flash | 63% ±4% | $0.24 | max |
| gpt-6-astra | 74% ±3% | $4.43 | xhigh |
| kimi-k3 | 69% ±5% | $4.65 | max |
Source: deepswe.datacurve.ai, v1.1, updated 2026-09-22. Board rows name specific versions. claude-opus-5 is the earlier Opus, not Opus 5.5, and the board has no row for GPT-6.1 Sol, Claude Sonnet 5.5, Devstral Small 2, Qwen3.8-27B or DeepSeek V4.1 Flash. Cost and effort differ per row.
How should I read coding benchmark scores?
Three of the rows above sit at 74% and their intervals overlap, so the board itself gives no order among them. The benchmark’s own page warns that leading public coding benchmarks are starting to saturate at the frontier, with adjacent configurations overlapping on confidence intervals.
Scores that a vendor reports for its own model are not comparable across vendors. The test version, the harness, the effort level and the number of attempts differ, and some cards compare their model with others under the vendor’s own method. This page quotes no vendor-reported score for that reason.
A benchmark also measures the benchmark’s tasks. Your repository has its own languages, test suite and conventions, so use a score to build a shortlist and your own tasks to decide.
How do I test coding models on my own repository?
1. Pick 20 real tasks
Take closed tickets that had tests or a clear review verdict, across easy, medium and hard. Remove any that need secrets or production access.
2. Fix the harness
Use the same agent, the same tool set and the same prompt template for every model. Change only the model and, separately, the effort setting.
3. Let tests judge
Score each run by whether the original tests pass and whether the diff stays inside the files it should touch. Review diffs that pass for needless edits.
4. Record cost and time
Log tokens, wall-clock time and the number of retries per task, then compute cost per merged task.
5. Repeat three times
Agent runs vary. Report the spread, and prefer a model whose worst run you can live with.
Where Swfte fits
This page is a selection guide. Swfte is not a model lab and builds no coding model. You can choose and use any model above without Swfte. For a ranked view with cited results, see best AI coding models.
Two Swfte products matter once you have a shortlist. Connect is one OpenAI-compatible API with routing, fallback chains, budgets and usage caps, so a team can send each task type to the model that passed its tests and switch without changing tools. Nexus wraps Claude Code and Codex with a policy gate (allow, deny or ask), an audit trail and completion gates. See Claude Code security for the controls, and the comparisons Cursor vs Claude Code, Codex vs Claude Code and OpenCode vs Claude Code.
Swfte provides the technical controls, governance mechanisms and evidence you need to deploy AI within your applicable regulatory, security and policy requirements. The exact posture depends on your use case, jurisdiction, deployment and configuration.
Sources and last verified
Every dated or technical fact on this page was read from the pages below on 2026-10-07. Anything that could not be confirmed is left out or marked as not verified.
- Anthropic: models overview. Model IDs, context windows and positioning for Claude Opus 5.5 and Sonnet 5.5.
- Anthropic: Claude Opus 5.5 model page. Agentic coding positioning and available platforms.
- OpenAI: Codex models. Which models are recommended for which Codex work.
- OpenAI: GPT-6.1 Sol model page. Model ID, context window and description.
- Google AI: Gemini API models. Gemini 3.8 Flash ID and positioning.
- Google DeepMind: Gemini 3.8 Flash model card. Context window.
- DeepSWE v1.1 leaderboard (Datacurve). Benchmark design, harness, update date and the rows quoted.
- Hugging Face: Devstral-Small-2-24B-Instruct-2512. Licence, context window, hardware statement and runtimes.
- Hugging Face: Qwen/Qwen3.8-27B. Licence, context length and runtimes.
- Hugging Face: DeepSeek-V4.1-Flash. Licence, parameters, context and runtimes.
- Hugging Face: zai-org/GLM-5.3. Licence name and runtimes.
- GLM-5.3 licence text. Model as a Service revenue and security-review clause.
- Hugging Face: moonshotai/Kimi-K3. Licence name, parameters, context and runtimes.
Frequently asked questions
What is the best LLM for coding?
It depends on your situation, so this page names no single winner. Private or offline work points to open weights you host. Long agentic tasks point to hosted models with long context and an agentic focus. Test two or three on 20 real tasks from your repository, and judge them by passing tests.
Can I run a coding LLM locally?
Yes. Devstral Small 2 (Apache 2.0) states that it can run on a single RTX 4090 or a Mac with 32 GB RAM, and names Ollama, llama.cpp and LM Studio among its runtimes. Larger models such as DeepSeek V4.1 Flash and Kimi K3 need multi-GPU hardware. Check memory under your real context length.
Are open-weight models good enough for coding?
Two open-weight models are listed close behind on the independent DeepSWE v1.1 board: glm-5.3 and kimi-k3 at 69%, against 74% for three hosted models, with overlapping intervals in places. That board has no row for the smaller models, so for them the only evidence is your own test. Run your own tasks.
Why do coding benchmark scores differ between sources?
Scores differ because the benchmark version, the harness, the reasoning effort and the number of attempts differ. A vendor’s own report and an independent leaderboard can disagree for the same model. This page quotes only the independent DeepSWE board, names its harness, and does not compare vendor-reported scores.
How is this different from the best AI coding models page?
The best AI coding models page ranks closed and open-weight models using cited results and prices. This page is a selection guide: it starts from your situation, such as a private repository or offline work, lists vendor-stated facts and tells you how to test. Use both, and trust your own repository most.
Does Swfte offer a coding model?
No. Swfte is not a model lab. Connect routes requests between models you choose behind one API, and Nexus adds a policy gate, audit trail and completion gates to Claude Code and Codex. Neither replaces the model, and you can use the models on this page without Swfte.