Fast and affordable small model for lightweight tasks and high-throughput use cases.
Context Window
128K
tokens
Max Output
16K
tokens
Input Price
$0.15
per 1M tokens
Output Price
$0.6
per 1M tokens
Speed
183
tokens/sec
Released
Jul 2024
2024-07-18
Blended Cost
$0.38
per 1M tokens
Value Score
192.0
quality per $
Capabilities
ChatVisionFunction CallingCode Generation
Benchmarks
Quality Index
72
MMLU Pro
82
HumanEval (Coding)
87.2
MATH
70.2
Arena ELO
1216
Price vs Quality
Cheaper is left, more capable is up. 5 of 62 priced models are both cheaper and score at least as high. Self-host models are left off the price axis — their cost depends on the hardware you run them on, not on tokens.
Context Window in Context
Grok 3131K
▸ GPT-4o Mini128K
GPT-4o128K
Mistral Large 2128K
DeepSeek V3128K
DeepSeek R1128K
Command R+128K
Gemma 4 27B128K
Bar length is logarithmic — each full step is 10× the tokens. Shown against the models with the nearest context windows, not the directory extremes.
출력 1024 토큰, 시간당 20회로 제한됩니다. 전체 제어는 playground에서 가능합니다.
About GPT-4o Mini
GPT-4o Mini is a fast AI model by OpenAI, released on July 18, 2024. It supports a context window of 128K tokens and can generate up to 16K output tokens.
At $0.15 per million input tokens and $0.6 per million output tokens, its blended cost of $0.38/1M tokens makes it one of the most affordable models available. Its value score of 192.0 reflects the balance of quality and cost.
Using GPT-4o Mini with Swfte
Access GPT-4o Mini through Swfte Connect, our unified LLM gateway. Connect gives you a single API for 50+ models, with automatic routing, cost optimization, and fallback handling. You can also try GPT-4o Mini in our AI Playground before integrating.
Deploy a model with Swfte Connect
One gateway, every provider, per-token cost visibility. Swap models without touching your code.