Model detail
gpt-4.1-nano-2025-04-14
For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size with its 1 million token context window, and scores 80.1% on MMLU, 50.3% on GPQA, and 9.8% on Aider polyglot coding – even higher than GPT‑4o mini. It’s ideal for tasks like classification or autocompletion.
Model specs
- Context length
- 1.047576M
- Max output
- 32.768K
- I/O modalities
- Text / Image
- Released
- 2025-04
02
API endpoints
One MixRoute gateway, OpenAI-compatible
-
OpenAI-compatible
/v1/chat/completionsPOST
03
Pricing
Flat rate, unit: /1M Tokens
Input
$0.1000 /1M Tokens
Completion
$0.2000 /1M Tokens
05
Selection summary
Quickly judge whether it fits your workload
What is this model good for?
Ultra-fast classification and autocompletion at scale
Use GPT-4.1 nano for high-volume tasks like classification, autocompletion, and data extraction. It delivers the 1M token context window at the lowest cost and latency in the GPT-4.1 series.
What should you check before using it?
Confirm accuracy for your task class
GPT-4.1 nano prioritizes speed and cost over deep reasoning. Test your classification and autocompletion tasks to confirm the accuracy meets your requirements before production deployment.
Why use it through MixRoute?
Use the OpenAI-compatible endpoint with stable model ID
Use the confirmed compatible endpoint with model ID gpt-4.1-nano through MixRoute, then run the same integration tests used for the current client.
08
Token cost estimator
Live estimate from this page's pricing, not an actual bill
This model has no cache-read price; caching is not counted
09
FAQ
GPT-4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance for its size with a 1 million token context window, making it ideal for classification, autocompletion, and other high-throughput tasks.
One endpoint, a testable decision
Test this model and alternate routes with the same request format
Start from a real workload, then let quality, total cost, and failure conditions decide whether to send production traffic.