Fast and affordable small model for lightweight tasks and high-throughput use cases.
Context Window
128K
tokens
Max Output
16K
tokens
Input Price
$0.15
per 1M tokens
Output Price
$0.6
per 1M tokens
Speed
183
tokens/sec
Released
Jul 2024
2024-07-18
Blended Cost
$0.38
per 1M tokens
Value Score
192.0
quality per $
Capabilities
ChatVisionFunction CallingCode Generation
Benchmarks
Quality Index
72
MMLU Pro
82
HumanEval (Coding)
87.2
MATH
70.2
Arena ELO
1216
Price vs Quality
Cheaper is left, more capable is up. 5 of 62 priced models are both cheaper and score at least as high. Self-host models are left off the price axis — their cost depends on the hardware you run them on, not on tokens.
Context Window in Context
Grok 3131K
▸ GPT-4o Mini128K
GPT-4o128K
Mistral Large 2128K
DeepSeek V3128K
DeepSeek R1128K
Command R+128K
Gemma 4 27B128K
Bar length is logarithmic — each full step is 10× the tokens. Shown against the models with the nearest context windows, not the directory extremes.
Capped at 1024 output tokens and 20 queries an hour. The playground gives you full control.
About GPT-4o Mini
GPT-4o Mini is a fast AI model by OpenAI, released on July 18, 2024. It supports a context window of 128K tokens and can generate up to 16K output tokens.
At $0.15 per million input tokens and $0.6 per million output tokens, its blended cost of $0.38/1M tokens makes it one of the most affordable models available. Its value score of 192.0 reflects the balance of quality and cost.
Using GPT-4o Mini with Swfte
Access GPT-4o Mini through Swfte Connect, our unified LLM gateway. Connect gives you a single API for 50+ models, with automatic routing, cost optimization, and fallback handling. You can also try GPT-4o Mini in our AI Playground before integrating.
Deploy a model with Swfte Connect
One gateway, every provider, per-token cost visibility. Swap models without touching your code.