Buyer's guide

Best self-hosted AI models for enterprises (2026)

Fourteen open-weight models an enterprise can run on its own or EU-hosted infrastructure, with licences, sizes, published hardware guidance and serving stacks read from each model card.

Last verified 6 October 2026

In short
For a self-hosted enterprise model the licence usually decides before quality does. Eleven of the fourteen models here are free of conditions for internal use, two have conditions that matter only if you resell model access, and one bars any company above US$20 million in monthly revenue. Hardware needs range from a single GPU to a full multi-GPU node.
Read this first

How we order the list

We found no comparable, independent evaluation covering all of these models, so there is no quality ranking. The order is: Swfte’s own model first, placed by the publisher and disclosed above; then the open-weight models by licence freedom for internal enterprise use; then alphabetically within each group. Every model has a model card or licence file that we read on 6 October 2026. Strengths are the labs’ own claims and are labelled as such. Hardware figures are either published by the lab, or are arithmetic we label as an estimate. Anything we could not confirm says “not verified”.

Decision guide

How to choose by workload

WorkloadCandidatesWhy
General enterprise assistant and RAGMistral Large 3, Mistral Small 4, Granite 4.2 30BLab-claimed document and RAG strengths, 256K or 128K context, Apache 2.0.
Coding and software agentsGLM-5.3, DeepSeek V4.1 Flash, Qwen3.8-27B, Devstral Small 2Coding results published by their labs; see our coding models guide for the independent leaderboard.
EU languagesCommand A+, Mistral Large 3, Mistral Small 4, Granite 4.2 30B, Gemma 4These cards name EU languages explicitly. Test your own languages: language lists are not quality measurements.
Agents and tool useMistral Large 3, Command A+, Granite 4.2 30B, Kimi K3Native function calling or tool calling is a stated design goal on their cards.
One GPU or the edgeNemotron 3.5 Lightning, Devstral Small 2, Gemma 4, Qwen3.8-27BPublished single-GPU guidance (Nemotron, Devstral) or small, official quantised builds.
Safety tooling and guard modelsGemma 4, Granite 4.2 with Granite Guardian, gpt-oss with gpt-oss-safeguardTheir labs publish safety material or companion guard models. Safety behaviour still has to be tested on your data.
Shortlists drawn from published strengths and language lists, not a measured ranking. Test on your own tasks.

For coding specifically, use our coding models guide, which ranks by an independent leaderboard. To test a shortlist before you commit, see open-source model testing; to run it, see deploying models.

The models

Fourteen open-weight models

#ModelLicenceSizeContext
1Swfte Safety model (publisher placement)<licence and availability — founder to fill><parameters — founder to fill><context — founder to fill>
2Command A+ (05-2026)Apache 2.0218B total, 25B active128K input
3DeepSeek V4.1 FlashMIT552B backbone plus 196B sparsely accessed memory (763.2B total in t…1M tokens
4Devstral Small 2 (24B)Apache 2.024B dense256K tokens
5Gemma 4 31B (instruction-tuned)Apache 2.030.7B dense. The family also has a 26B MoE (25.2B total, 3.8B activ…256K tokens (E2B and E4B: 128K)
6GLM-5.3-FlashMIT320B total, 18B active; multimodalnot verified
7gpt-oss-120bApache 2.0116.8B total (Hugging Face metadata); active parameters and context…not verified
8Granite 4.2 30BApache 2.030B dense reasoning model (3B and 8B siblings)128K native, extendable to 512K
9Mistral Large 3 (675B Instruct 2512)Apache 2.0Card says 675B total and 41B active in one place and 673B and 39B i…256K
10Mistral Small 4 (119B A6B)Apache 2.0119B total, 6.5B active per token256K
11NVIDIA Nemotron 3.5 Lightning 30B-A3BOpenMDW License Agreement 1.130B total, 3B active (Mamba-2, MoE and attention hybrid)Up to 1M (256K on a single H100)
12Qwen3.8-27BApache 2.027B dense (hybrid attention) with a vision encoder262,144 native, extensible to 1,000,000
13GLM-5.3GLM-5.3 LicenseAbout 753B in the published FP8 weights; active parameters not statednot verified
14Kimi K3Kimi K3 License2.8T total, 104B active1,048,576 tokens
15Mistral Medium 3.5 (128B)Modified MIT128B dense256K
Swfte’s row has no facts because none are published yet. The other rows are ordered by licence freedom, not quality.
#2 · Cohere · No conditions beyond the standard licence for internal use

Command A+ (05-2026)

Licence: Apache 2.0. Commercial terms: None beyond Apache 2.0. Earlier Command A repositories are CC-BY-NC-4.0 (non-commercial) and are not the same.

Size: 218B total, 25B active. Context window: 128K input. Released: Hugging Face repository created 11 May 2026.

Hardware and quantisation: Model card minimum GPUs: BF16 4x B200 or 8x H100; FP8 2x B200 or 4x H100; W4A4 1x B200 or 2x H100. Quantisation: Official BF16, FP8 and W4A4 repositories.

Serving stacks named on its card: Transformers (from source) and vLLM 0.21.0 or later with the cohere-melody parser.

Strengths: Lab claim: optimised for agentic, multilingual, reasoning-heavy enterprise tasks. Languages: Card lists about 50 languages including most EU languages (bg, ca, cs, da, de, el, es, et, fi, fr, ga, hr, hu, it, lt, lv, mt, nl, pl, pt, ro, sk, sl, sv).

Safety material: not verified

Watch-outs:

  • Check you are using the Apache 2.0 repository, not an earlier non-commercial Command A one.

Sources: Hugging Face: command-a-plus-05-2026-bf16 (read 2026-10-06).

#3 · DeepSeek · No conditions beyond the standard licence for internal use

DeepSeek V4.1 Flash

Licence: MIT. Commercial terms: None beyond MIT.

Size: 552B backbone plus 196B sparsely accessed memory (763.2B total in the published tensors); 8B active per token at prefill and 16B at decode. Context window: 1M tokens. Released: Hugging Face repository created 10 September 2026.

Hardware and quantisation: not verified Quantisation: Published as 8-bit tensors with an FP4 main KV cache described on the card.

Serving stacks named on its card: The card points to DeepSeek’s own inference folder and recipe; it does not name vLLM or SGLang.

Strengths: Lab claim: agentic and coding performance with controllable reasoning effort; the card reports a KV cache of 890 bytes per token. Languages: No EU-language statement on the card.

Safety material: not verified

Watch-outs:

  • Very large to host despite the small active size: all experts must be resident in memory.

Sources: Hugging Face: DeepSeek-V4.1-Flash (read 2026-10-06).

#4 · Mistral AI · No conditions beyond the standard licence for internal use

Devstral Small 2 (24B)

Licence: Apache 2.0 (the larger sibling, Devstral 2 at 123B, is not Apache). Commercial terms: Card: commercial and non-commercial use.

Size: 24B dense. Context window: 256K tokens. Released: Hugging Face repository created 28 November 2025.

Hardware and quantisation: Card: light enough to run on a single RTX 4090 or a Mac with 32GB RAM. Quantisation: Main repository is FP8; community GGUFs exist.

Serving stacks named on its card: vLLM, SGLang, llama.cpp, LM Studio and Ollama are named on the card.

Strengths: Lab claim: agentic coding; 68.0% on SWE-bench Verified (self-reported). Languages: not verified

Safety material: not verified

Watch-outs:

  • An older model, aimed at coding rather than general chat.

Sources: Hugging Face: Devstral-Small-2-24B-Instruct-2512 (read 2026-10-06).

#5 · Google DeepMind · No conditions beyond the standard licence for internal use

Gemma 4 31B (instruction-tuned)

Licence: Apache 2.0 (the card links Google’s Gemma 4 licence page, which renders as standard Apache 2.0 text). Whether Google’s separate prohibited-use policy applies was not confirmed.. Commercial terms: Apache 2.0; no monthly-user cap seen.

Size: 30.7B dense. The family also has a 26B MoE (25.2B total, 3.8B active), 12B, E4B and E2B.. Context window: 256K tokens (E2B and E4B: 128K). Released: Hugging Face repository created 11 March 2026.

Hardware and quantisation: Google docs, weights only including 20% overhead and excluding KV cache: 31B = 69.9 GB BF16, 34.9 GB SFP8, 17.5 GB Q4_0; 26B A4B = 57.7, 28.8, 14.4 GB. Quantisation: Official quantisation-aware training builds: GGUF Q4_0 and W4A16.

Serving stacks named on its card: Card names Transformers and llama.cpp. vLLM and Ollama support was not confirmed from the card.

Strengths: Lab claim: reasoning with configurable thinking, coding, native function calling. Languages: Card: out-of-the-box support for 35 or more languages, pre-trained on over 140.

Safety material: The card has an Ethics and Safety section: the same safety evaluations as Gemini, “major improvements” over Gemma 3, tested without safety filters. No guard model named in the part we read.

Watch-outs:

  • Check Google’s prohibited-use policy against your use case.

Sources: Hugging Face: gemma-4-31B-it, Google: Gemma 4 model sizes and memory (read 2026-10-06).

#6 · Z.ai (Zhipu) · No conditions beyond the standard licence for internal use

GLM-5.3-Flash

Licence: MIT. Commercial terms: None beyond MIT.

Size: 320B total, 18B active; multimodal. Context window: not verified. Released: Hugging Face repository created 25 August 2026.

Hardware and quantisation: not verified Quantisation: A BF16 repository also exists.

Serving stacks named on its card: SGLang, vLLM, TokenSpeed, Transformers, KTransformers, Unsloth (card).

Strengths: Lab claim: outperforms GLM-5.2 at one-tenth of the price and approaches Claude Opus 4.8 on coding and agentic benchmarks. Languages: Tagged English and Chinese only.

Safety material: not verified

Watch-outs:

  • Only English and Chinese declared, so weak evidence for EU-language use.

Sources: Hugging Face: GLM-5.3-Flash (read 2026-10-06).

#7 · OpenAI · No conditions beyond the standard licence for internal use

gpt-oss-120b

Licence: Apache 2.0 (Hugging Face licence tag; licence text not read by us). Commercial terms: Not verified beyond the Apache 2.0 tag.

Size: 116.8B total (Hugging Face metadata); active parameters and context window not verified. Context window: not verified. Released: Hugging Face repository created 4 August 2025.

Hardware and quantisation: not verified Quantisation: Hugging Face tags: MXFP4 (native) and 8-bit.

Serving stacks named on its card: vLLM (Hugging Face library tag); others not verified.

Strengths: Not verified: the model card was not read. Languages: not verified

Safety material: Companion guard models gpt-oss-safeguard-120b and gpt-oss-safeguard-20b are published under Apache 2.0.

Watch-outs:

  • The oldest model in this list (August 2025).
  • We read only registry metadata, not the card, so strengths and hardware are not verified.

Sources: Hugging Face: gpt-oss-120b (read 2026-10-06).

#8 · IBM · No conditions beyond the standard licence for internal use

Granite 4.2 30B

Licence: Apache 2.0. Commercial terms: Card: unrestricted commercial and academic use.

Size: 30B dense reasoning model (3B and 8B siblings). Context window: 128K native, extendable to 512K. Released: 25 August 2026 (model card).

Hardware and quantisation: not verified Quantisation: Official FP8, NVFP4 and MXFP4 repositories; GGUF and MLX builds.

Serving stacks named on its card: vLLM 0.20 or later (needs a custom reasoning parser) and SGLang; GGUF and MLX published.

Strengths: Lab claim: reasoning, code generation, tool calling and agentic workflows, with flexible thinking modes. Languages: Tested: English, German, Spanish, French, Japanese, Portuguese, Arabic, Czech, Italian, Korean, Dutch, Chinese.

Safety material: The card recommends pairing it with Granite Guardian (granite-guardian-4.1-8b) and warns of possible biased or unsafe output.

Watch-outs:

  • Needs a recent vLLM with the custom reasoning parser.

Sources: Hugging Face: granite-4.2-30b (read 2026-10-06).

#9 · Mistral AI · No conditions beyond the standard licence for internal use

Mistral Large 3 (675B Instruct 2512)

Licence: Apache 2.0. Commercial terms: Card: for both commercial and non-commercial purposes.

Size: Card says 675B total and 41B active in one place and 673B and 39B in another; the inconsistency is unresolved.. Context window: 256K. Released: Hugging Face repository created 28 November 2025.

Hardware and quantisation: Card: FP8 on a single node of B200s or H200s; NVFP4 on a single node of H100s or A100s. Quantisation: FP8 (main), NVFP4, BF16 and a speculative-decoding head repository.

Serving stacks named on its card: vLLM (Hugging Face library tag); other engines not checked on this card.

Strengths: Lab claim: best-in-class agentic capabilities with native function calling and JSON output; long-document understanding and RAG. Languages: Card languages: en, fr, es, de, it, pt, nl, zh, ja, ko, ar.

Safety material: not verified

Watch-outs:

  • Needs a full multi-GPU node.
  • Whether Mistral AI is an EU company was not verified from Mistral’s own pages, so this guide does not assert it.

Sources: Hugging Face: Mistral-Large-3-675B-Instruct-2512 (read 2026-10-06).

#10 · Mistral AI · No conditions beyond the standard licence for internal use

Mistral Small 4 (119B A6B)

Licence: Apache 2.0. Commercial terms: Card: open-source licence for commercial and non-commercial use.

Size: 119B total, 6.5B active per token. Context window: 256K. Released: Hugging Face repository created 23 January 2026 (the repository name suggests March 2026; unresolved).

Hardware and quantisation: Card example serves with tensor parallelism 2; no explicit VRAM statement read. Quantisation: Official NVFP4 and speculative-decoding head repositories; community GGUFs.

Serving stacks named on its card: vLLM (recommended), llama.cpp (community GGUFs), LM Studio, SGLang, Transformers.

Strengths: Lab claim: unifies instruct, reasoning and coding in one model; 40% lower latency and 3x requests per second against Mistral Small 3. Languages: Card lists 24 languages including en, fr, de, es, pt, it, pl, ro, sv.

Safety material: not verified

Watch-outs:

  • The release date is inconsistent between the repository metadata and its name.

Sources: Hugging Face: Mistral-Small-4-119B-2603 (read 2026-10-06).

#11 · NVIDIA · No conditions beyond the standard licence for internal use

NVIDIA Nemotron 3.5 Lightning 30B-A3B

Licence: OpenMDW License Agreement 1.1. Commercial terms: Per the licence text: free to deal in the model materials without restriction; keep the licence and notices on redistribution; rights end if you sue claiming the materials infringe. The card says it is ready for commercial use.

Size: 30B total, 3B active (Mamba-2, MoE and attention hybrid). Context window: Up to 1M (256K on a single H100). Released: 11 August 2026 (model card).

Hardware and quantisation: Card: single-GPU deployment on one H100 80GB or A100 80GB; validated on GB200, B200, H100, H200 and A100. Quantisation: Official NVFP4 repository; speculative-decoding variants.

Serving stacks named on its card: vLLM and SGLang recipes on the card.

Strengths: Lab claim: best for customisation (SFT, RL, distillation). Languages: English, Spanish, French, German, Italian, Japanese (and code).

Safety material: not verified

Watch-outs:

  • A custom licence, though a permissive one: have counsel read it.
  • Sibling Nemotron models use a different NVIDIA licence whose terms we did not read.

Sources: Hugging Face: NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 (read 2026-10-06).

#12 · Alibaba Qwen · No conditions beyond the standard licence for internal use

Qwen3.8-27B

Licence: Apache 2.0. Commercial terms: None beyond Apache 2.0.

Size: 27B dense (hybrid attention) with a vision encoder. Context window: 262,144 native, extensible to 1,000,000. Released: Hugging Face repository created 5 August 2026.

Hardware and quantisation: not verified Quantisation: Official FP8 repository (also Apache 2.0).

Serving stacks named on its card: Transformers, vLLM, SGLang and TokenSpeed; the card recommends dedicated engines for production.

Strengths: Lab claim: coding, professional work and long-horizon agentic tasks; 61.7 on SWE-bench Pro and 73.0 on Terminal Bench 2.1 (self-reported). Languages: No language list found on the card.

Safety material: not verified

Watch-outs:

  • Larger Qwen3.8 models carry custom licences with revenue conditions: read the licence of the exact model.

Sources: Hugging Face: Qwen3.8-27B (read 2026-10-06).

#13 · Z.ai (Zhipu) · Permissive, with conditions that bite only in specific cases

GLM-5.3

Licence: GLM-5.3 License (custom, MIT-style). Commercial terms: Free to use, modify, deploy and sell. The one condition: a licensee or affiliate running a “Model as a Service” business with more than US$10 billion revenue in 12 months must pass Z.AI’s security review before commercial use. End-user products embedding the model are excluded from that definition.

Size: About 753B in the published FP8 weights; active parameters not stated. Context window: not verified. Released: Hugging Face repository created 25 August 2026.

Hardware and quantisation: not verified Quantisation: FP8 main repository; a BF16 repository also exists.

Serving stacks named on its card: SGLang, vLLM, TokenSpeed, Transformers, KTransformers, Unsloth and Ascend NPU engines (card).

Strengths: Lab claim: the most capable open-weights model for coding; 88.2 on Terminal-Bench 2.1 (self-reported). Languages: Tagged English and Chinese only.

Safety material: The card notes an emergent cyber-exploitation capability from post-training and has no safety section or guard model.

Watch-outs:

  • Needs a multi-GPU node.
  • The card’s own note on cyber capability means you should add controls around any agent that has network or shell access.

Sources: Hugging Face: GLM-5.3, GLM-5.3 licence text (read 2026-10-06).

#14 · Moonshot AI · Permissive, with conditions that bite only in specific cases

Kimi K3

Licence: Kimi K3 License (custom, MIT-style). Commercial terms: A licensee or affiliate running a “Model as a Service” business above US$20 million revenue in 12 months needs a separate agreement with Moonshot before commercial use. Products above 100M monthly users or US$20M monthly revenue must display “Kimi K3”. Neither applies to internal use with no third-party access.

Size: 2.8T total, 104B active. Context window: 1,048,576 tokens. Released: Hugging Face repository created 13 June 2026.

Hardware and quantisation: not verified Quantisation: Quantisation-aware training: MXFP4 weights with MXFP8 activations.

Serving stacks named on its card: vLLM, SGLang and TokenSpeed (card).

Strengths: Lab claim: long-horizon coding, agentic work and knowledge work with a 1M context. Languages: not verified

Safety material: not verified

Watch-outs:

  • Needs a large multi-GPU cluster.
  • The licence matters if you resell model access to third parties.

Sources: Hugging Face: Kimi-K3 (read 2026-10-06).

#15 · Mistral AI · Restricted: most enterprises cannot use it without a separate deal

Mistral Medium 3.5 (128B)

Licence: Modified MIT. Commercial terms: You are not authorised to exercise any rights if your company’s global consolidated monthly revenue exceeded US$20 million in the preceding month; this applies to derivatives. Larger companies need a commercial licence from Mistral, at Mistral’s sole discretion, or must use Mistral’s hosted service.

Size: 128B dense. Context window: 256K. Released: Hugging Face repository created 31 March 2026.

Hardware and quantisation: Card example uses tensor parallelism 8 and lists SGLang images for Hopper and Blackwell GPUs; no explicit VRAM figure read. Quantisation: A speculative-decoding head repository; community GGUFs.

Serving stacks named on its card: vLLM, SGLang and llama.cpp (community GGUFs).

Strengths: Lab claim: unified instruct, reasoning and coding; 77.6% on SWE-bench Verified (self-reported). Languages: Card lists 24 languages including en, fr, de, es, pt, it, pl, ro, sv, uk, ru, tr.

Safety material: not verified

Watch-outs:

  • Not usable self-hosted by most enterprises without a separate agreement.

Sources: Hugging Face: Mistral-Medium-3.5-128B, Mistral Medium 3.5 licence text (read 2026-10-06).

Hardware sizing

What hardware each model needs

Where a lab publishes guidance we quote it. Where it does not, we show an estimate: parameter count times bytes per parameter, which covers weights only. KV cache grows with context length and concurrency, and the runtime adds overhead (Google’s figures include 20% overhead and still exclude KV cache). Treat estimates as a floor, not a plan, and confirm by loading the model.

ModelTotal parametersHardware: lab-published or our estimate
Command A+ (05-2026)218BModel card minimum GPUs: BF16 4x B200 or 8x H100; FP8 2x B200 or 4x H100; W4A4 1x B200 or 2x H100.
DeepSeek V4.1 Flash763.2BEstimate, weights only: about 763.2 GB (763.2B parameters at 8-bit tensors as published (F8 and I8)). Our arithmetic, not a lab figure.
Devstral Small 2 (24B)24BCard: light enough to run on a single RTX 4090 or a Mac with 32GB RAM.
Gemma 4 31B (instruction-tuned)30.7BGoogle docs, weights only including 20% overhead and excluding KV cache: 31B = 69.9 GB BF16, 34.9 GB SFP8, 17.5 GB Q4_0; 26B A4B = 57.7, 28.8, 14.4 GB.
GLM-5.3-Flash320BEstimate, weights only: about 320 GB (320B parameters at 8-bit; the precision of the main repository was not checked). Our arithmetic, not a lab figure.
gpt-oss-120b116.8BEstimate, weights only: about 58.4 GB (116.8B parameters at native MXFP4 (about 4 bits per parameter; scales add a little)). Our arithmetic, not a lab figure.
Granite 4.2 30B30BEstimate, weights only: about 30 GB (30B parameters at FP8 (official FP8 repository)). Our arithmetic, not a lab figure.
Mistral Large 3 (675B Instruct 2512)675BCard: FP8 on a single node of B200s or H200s; NVFP4 on a single node of H100s or A100s.
Mistral Small 4 (119B A6B)119BCard example serves with tensor parallelism 2; no explicit VRAM statement read.
NVIDIA Nemotron 3.5 Lightning 30B-A3B30BCard: single-GPU deployment on one H100 80GB or A100 80GB; validated on GB200, B200, H100, H200 and A100.
Qwen3.8-27B27BEstimate, weights only: about 27 GB (27B parameters at FP8 (official FP8 repository)). Our arithmetic, not a lab figure.
GLM-5.3753.3BEstimate, weights only: about 753.3 GB (753.3B parameters at FP8 as published). Our arithmetic, not a lab figure.
Kimi K32800BEstimate, weights only: about 1400 GB (2800B parameters at native MXFP4 weights (about 4 bits per parameter; scales add a little)). Our arithmetic, not a lab figure.
Mistral Medium 3.5 (128B)128BCard example uses tensor parallelism 8 and lists SGLang images for Hopper and Blackwell GPUs; no explicit VRAM figure read.
Estimates assume the precision named in the row, which is the one the lab publishes where we could tell.
Licence traps

Where open weights stop being open

  • Revenue caps. Mistral Medium 3.5 bars companies above US$20 million global monthly revenue. It is the clearest example of an “open” model most enterprises cannot use.
  • Model-as-a-service clauses. Kimi K3 (above US$20 million) and GLM-5.3 (above US$10 billion) require a separate deal or a security review if you run a model-as-a-service business. Internal use is outside them. Read the definition before relying on that.
  • Same family, different licence. Cohere’s earlier Command A repositories are non-commercial while Command A+ is Apache 2.0. Mistral’s Devstral Small 2 is Apache 2.0 while its larger sibling is not. Qwen’s larger models carry custom licences while Qwen3.8-27B is Apache 2.0.
  • Non-commercial by default. MiniMax M3 allows only non-commercial use unless you meet its attribution and notice terms, with prior authorisation above US$20 million revenue. We left it out for that reason.
  • Gated cards. Llama 4’s model cards require a login, so we could not verify its licence terms here and did not include it. Read Meta’s licence directly if you consider it.
  • Derivatives inherit terms. Fine-tunes and distillations usually carry the base licence. Mistral Medium 3.5’s revenue cap, for example, applies to derivatives.

This is a reading of the licence texts we could fetch, not legal advice.

Serving

Serving stacks

The engines below are the ones named on the model cards we read. We did not check each project’s own documentation, so treat the table as “what the labs say works”. Our blog covers the layers of a self-hosted stack.

EngineWhat the cards say
vLLMNamed on the cards of Qwen3.8-27B, GLM-5.3 and GLM-5.3-Flash, Kimi K3, Mistral Small 4, Medium 3.5 and Devstral Small 2, Granite 4.2, Command A+, Nemotron 3.5 and (as a library tag) Mistral Large 3 and gpt-oss.
SGLangNamed on the cards of Qwen3.8-27B, GLM-5.3, Kimi K3, Mistral Small 4 and Medium 3.5, Granite 4.2 and Nemotron 3.5.
llama.cppNamed for Gemma 4 and, through community GGUF builds, for the Mistral models. Suited to single-machine and CPU-assisted setups.
OllamaLinked from the Devstral Small 2 card. Support for other models was not verified.
TokenSpeedA newer engine named beside vLLM and SGLang on the Qwen3.8, GLM-5.3 and Kimi K3 cards.
Text Generation Inference (TGI)Not named on any card we read. Its maintenance status was not verified, so we do not recommend it or rule it out.
Checklist

Evaluation checklist

  • Read the licence of the exact repository you will deploy, not the family page, and have counsel confirm internal use, resale and derivative terms.
  • Run your own evaluation set in the languages you need. A language list on a model card is not a quality measurement.
  • Test at the quantisation you will actually serve. Published scores are usually at the lab’s chosen precision.
  • Size for KV cache and concurrency, not just weights: long contexts and many users can cost as much memory as the weights.
  • Test refusal, escalation and prompt-injection behaviour with your policies, before and after any fine-tune.
  • Pin the revision and verify checksums of the weights you download, and keep an internal mirror.
  • Plan the exit: keep a second model that passes the same tests so a licence or support change does not strand you.
Left out

Models we considered and left out

  • Llama 5, DeepSeek V4.5, Gemini 3.2 Pro. Never shipped (Gemini 3.2 Pro would be closed weights in any case). Never listed.
  • Llama 4 Scout and Maverick. Gated model cards: licence terms not verified.
  • MiniMax M3. Non-commercial by default; see licence traps.
  • Qwen3.8-2.4T-A95B and Qwen3.8-Flash-Next. Exist, under custom Qwen licences with model-as-a-service revenue conditions; left out for length, and Flash-Next is described by its lab as an experimental architecture preview.
  • Qwen 3.7. No public weights found.
  • Cohere Command A (2025) and variants. CC-BY-NC-4.0, non-commercial; superseded here by Command A+.
  • Phi and OLMo. No newer flagship found in the organisation listings and cards were not read.

Common questions

What is the best self-hosted AI model for an enterprise?
There is no single best: it depends on workload, languages, hardware and licence. This guide lists 14 open-weight models, orders them by licence freedom rather than quality, and maps them to workloads. For coding specifically, see our coding models guide, which uses an independent leaderboard.
Is “open-weight” the same as “open source”?
No. Open weights means you can download the model. The licence still decides what you may do with it. Apache 2.0 and MIT licences are broad. Several models here have custom licences with revenue or model-as-a-service conditions, and Mistral Medium 3.5 bars companies above US$20 million in monthly revenue.
How much GPU memory do I need?
Weights alone take roughly the parameter count times the bytes per parameter: 27B parameters at 8-bit is about 27 GB. KV cache, concurrency and runtime overhead come on top. Google’s published figures include 20% overhead and still exclude KV cache. The sizing table on this page separates lab-published figures from our arithmetic.
Can I run these models in the EU?
Yes. Open weights run on any infrastructure you control, including European providers. See our guide to sovereign cloud providers for what each provider publishes. Model origin and hosting location are separate questions: a model from a non-EU lab can run entirely on EU infrastructure.
Which models are safe to use out of the box?
None can be called safe without testing on your data. Some labs publish safety material or guard models (Gemma 4, Granite Guardian, gpt-oss-safeguard); GLM-5.3’s card flags an emergent cyber capability. Swfte’s own safety-first model is described by design intent only and has no published evaluation results yet.
Why are Llama 4 and MiniMax M3 not on the list?
Llama 4’s model cards are gated, so we could not verify its terms. MiniMax M3 is non-commercial by default. Both are mentioned in the licence traps. Models that do not exist, such as Llama 5 and DeepSeek V4.5, are never listed.
Where does Swfte fit?
Swfte publishes this guide and lists its own safety model first as the publisher, disclosed at the top. Swfte also provides the platform layer that deploys and governs models like these on infrastructure you choose. The Swfte Safety model is shown by design intent only: its evaluation facts are not yet published.

Deploy a model with Swfte Connect

One gateway, every provider, per-token cost visibility. Swap models without touching your code.