Model detail
gpt-4o-mini
GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more affordable than other recent frontier models, and more than 60% cheaper than [GPT-3.5 Turbo](/models/openai/gpt-3.5-turbo). It maintains SOTA intelligence, while being significantly more cost-effective. GPT-4o mini achieves an 82% score on MMLU and presently ranks higher than GPT-4 on chat preferences [common leaderboards](https://arena.lmsys.org/). Check out the [launch announcement](https://openai.com/index/gpt-4o-mini-advancing-cost-efficient-intelligence/) to learn more. #multimodal
Model specs
- Context length
- 128K
- Max output
- 16.384K
- I/O modalities
- Text / Image
- Released
- 2024-07
02
API endpoints
One MixRoute gateway, OpenAI-compatible
-
OpenAI-compatible
/v1/chat/completionsPOST
03
Pricing
Flat rate, unit: /1M Tokens
Input
$0.1500 /1M Tokens
Completion
$0.6000 /1M Tokens
Cache read
$0.0750 /1M Tokens
05
Selection summary
Quickly judge whether it fits your workload
What is this model good for?
Cost-effective multimodal tasks at scale
Use GPT-4o mini for high-volume multimodal tasks where cost efficiency is critical. It supports text and image inputs, achieving strong benchmark scores while being over 60% cheaper than GPT-3.5 Turbo.
What should you check before using it?
Confirm performance on your specific task class
GPT-4o mini is a small model optimized for cost. Test your specific multimodal tasks to confirm the accuracy meets your requirements compared to larger models like GPT-4o.
Why use it through MixRoute?
Use the OpenAI-compatible endpoint with stable model ID
Use the confirmed compatible endpoint with model ID gpt-4o-mini through MixRoute, then run the same integration tests used for the current client.
08
Token cost estimator
Live estimate from this page's pricing, not an actual bill
Share of the same prefix read repeatedly, up to 100%
09
FAQ
GPT-4o mini is OpenAI’s most advanced small model, supporting both text and image inputs. It is over 60% cheaper than GPT-3.5 Turbo while maintaining state-of-the-art intelligence, achieving an 82% score on MMLU.
One endpoint, a testable decision
Test this model and alternate routes with the same request format
Start from a real workload, then let quality, total cost, and failure conditions decide whether to send production traffic.