Kimi K3

Moonshot AIflagshipOpen Source

Moonshot AI's 16 Jul 2026 flagship and the clearest evidence that the open-weight tier has caught the frontier: a 2.8-trillion-parameter MoE (16 of 896 experts routed per token, Stable LatentMoE, MXFP4 weights) with a 1,048,576-token context and native text, vision, and video input. Artificial Analysis Intelligence Index 57, ranking #3 overall behind only Claude Fable 5 and GPT-5.6 Sol, and ahead of every other proprietary model on the board. It beats GLM-5.2 across Moonshot's harness — DeepSWE 67.5 vs 46.2, FrontierSWE 81.2 vs 67.3, SWE Marathon 42.0 vs 13.0, Terminal-Bench 2.1 88.3 vs 82.7, GPQA-Diamond 93.5 vs 91.2 — and tops the open board on Humanity's Last Exam (56%) and BrowseComp (91.2). At $3/$15 it undercuts both Claude Opus 5 ($5/$25) and GPT-5.6 Sol ($5/$30) on input and output while scoring higher on independent benchmarks. Two caveats: K3 is a heavy token consumer, and it is a genuine cost outlier among open models at ~$0.94 per task versus DeepSeek V4 Pro's $0.04. Weights shipped on Hugging Face at moonshotai/Kimi-K3 (2.8T total, 104B active) under Moonshot's own Kimi K3 License rather than a standard open licence, so read the terms before building on it. Hallucination rate rose to 51% from K2.6's 39% — cross-check factual output.

Context Window

1M

tokens

Max Output

33K

tokens

Input Price

$3

per 1M tokens

Output Price

$15

per 1M tokens

Speed

55

tokens/sec

Released

Jul 2026

2026-07-16

Blended Cost

$9.00

per 1M tokens

Value Score

10.8

quality per $

Capabilities

ChatVisionFunction CallingCode GenerationReasoning

Benchmarks

Quality Index
97
MMLU Pro
89.5
HumanEval (Coding)
94
MATH
92
Arena ELO
1500

Price vs Quality

6080100$1.00$10.00$100.00Blended cost per 1M tokens (log scale) →GPT-4o — $6.25/1M, quality 85GPT-4o Mini — $0.375/1M, quality 72o3 Mini — $2.75/1M, quality 88o3 — $5.00/1M, quality 94GPT-4.1 — $5.00/1M, quality 89Claude Opus 4 — $45.00/1M, quality 91Claude Sonnet 4 — $9.00/1M, quality 88Claude 3.5 Haiku — $2.40/1M, quality 75Gemini 2.5 Pro — $5.63/1M, quality 92Gemini 2.0 Flash — $0.250/1M, quality 74Llama 4 Maverick — $0.400/1M, quality 80Llama 4 Scout — $0.275/1M, quality 71Mistral Large 2 — $4.00/1M, quality 79Codestral — $0.600/1M, quality 76DeepSeek V3 — $0.685/1M, quality 86DeepSeek R1 — $1.37/1M, quality 91Grok 3 — $9.00/1M, quality 87Grok 3 Mini — $0.400/1M, quality 78Command R+ — $6.25/1M, quality 68Amazon Nova Pro — $2.00/1M, quality 70Qwen 2.5 72B — $0.600/1M, quality 80Qwen 2.5 Coder 32B — $0.300/1M, quality 74Sonar Pro — $9.00/1M, quality 78GPT-5.5 — $17.50/1M, quality 97GPT-5.5 Pro — $105.00/1M, quality 96Claude Opus 4.7 — $15.00/1M, quality 96Gemini 3.1 Pro — $7.00/1M, quality 96DeepSeek V4 Pro — $1.32/1M, quality 89DeepSeek V4 Flash — $0.210/1M, quality 84Qwen 3.6 Plus — $3.50/1M, quality 86Grok 4.3 — $1.88/1M, quality 93Kimi K2.6 — $2.11/1M, quality 92Claude Sonnet 4.6 — $9.00/1M, quality 90GLM-5.1 — $2.03/1M, quality 88Claude Opus 4.8 — $15.00/1M, quality 98Qwen 3.7 Max — $5.00/1M, quality 94Claude Fable 5 — $30.00/1M, quality 100MiniMax M3 — $1.50/1M, quality 89GPT-5.6 Sol — $17.50/1M, quality 98GPT-5.6 Terra — $7.00/1M, quality 93GPT-5.6 Luna — $0.700/1M, quality 83Gemini 3.6 Flash — $2.25/1M, quality 90Gemini 3.5 Flash-Lite — $1.40/1M, quality 79Claude Sonnet 5 — $9.00/1M, quality 93GLM-5.2 — $2.90/1M, quality 91Claude Opus 5 — $15.00/1M, quality 99Claude Mythos 5 — $30.00/1M, quality 100Grok 4.5 — $4.00/1M, quality 94Qwen3.8 Max — $4.00/1M, quality 96Kimi K3 — $9.00/1M, quality 97Kimi K3
Cheaper is left, more capable is up. Nothing in the directory is both cheaper and at least as capable. Self-host models are left off the price axis — their cost depends on the hardware you run them on, not on tokens.

Context Window in Context

  • GPT-5.6 Sol1.1M
  • GPT-5.6 Terra1.1M
  • GPT-5.6 Luna1.1M
  • Kimi K31.0M
  • GPT-4.11M
  • Gemini 2.5 Pro1M
  • Gemini 2.0 Flash1M
  • Llama 4 Maverick1M
Bar length is logarithmic — each full step is 10× the tokens. Shown against the models with the nearest context windows, not the directory extremes.

Open Source: Licensed under Modified MIT

Not on our gateway yet

Kimi K3 is open-weight, so you can run it yourself. Swfte Connect turns your deployment into a managed endpoint that speaks the same API as every other model in this directory.

The request
curl https://api.swfte.com/agents/v2/gateway/chat/completions \
  -H "Authorization: Bearer $SWFTE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "moonshot:kimi-k3",
    "messages": [{"role": "user", "content": "Hello"}]
  }'
Deploy this model with Swfte Connect

About Kimi K3

Kimi K3 is a flagship AI model by Moonshot AI, released on July 16, 2026. It supports a context window of 1049K tokens and can generate up to 33K output tokens.

At $3 per million input tokens and $15 per million output tokens, its blended cost of $9.00/1M tokens places it in the premium pricing tier. Its value score of 10.8 reflects the balance of quality and cost.

Kimi K3 is available as an open-source model under the Modified MIT license, meaning you can self-host it for predictable costs or use it through API providers like Swfte Connect.

Using Kimi K3 with Swfte

Access Kimi K3 through Swfte Connect, our unified LLM gateway. Connect gives you a single API for 50+ models, with automatic routing, cost optimization, and fallback handling. You can also try Kimi K3 in our AI Playground before integrating.

Deploy a model with Swfte Connect

One gateway, every provider, per-token cost visibility. Swap models without touching your code.