Skip to content

Model detail

gpt-4.1-nano-2025-04-14

Provided by OpenAI
Pay-as-you-go

For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size with its 1 million token context window, and scores 80.1% on MMLU, 50.3% on GPQA, and 9.8% on Aider polyglot coding – even higher than GPT‑4o mini. It’s ideal for tasks like classification or autocompletion.

Model specs

Context length
1.047576M
Max output
32.768K
I/O modalities
Text / Image
Released
2025-04

02

API endpoints

One MixRoute gateway, OpenAI-compatible

  • OpenAI-compatible /v1/chat/completions POST

03

Pricing

Flat rate, unit: /1M Tokens

Input

$0.1000 /1M Tokens

Completion

$0.2000 /1M Tokens

05

Selection summary

Quickly judge whether it fits your workload

What is this model good for?

Ultra-fast classification and autocompletion at scale

Use GPT-4.1 nano for high-volume tasks like classification, autocompletion, and data extraction. It delivers the 1M token context window at the lowest cost and latency in the GPT-4.1 series.

What should you check before using it?

Confirm accuracy for your task class

GPT-4.1 nano prioritizes speed and cost over deep reasoning. Test your classification and autocompletion tasks to confirm the accuracy meets your requirements before production deployment.

Why use it through MixRoute?

Use the OpenAI-compatible endpoint with stable model ID

Use the confirmed compatible endpoint with model ID gpt-4.1-nano through MixRoute, then run the same integration tests used for the current client.

08

Token cost estimator

Live estimate from this page's pricing, not an actual bill

This model has no cache-read price; caching is not counted

09

FAQ

GPT-4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance for its size with a 1 million token context window, making it ideal for classification, autocompletion, and other high-throughput tasks.

One endpoint, a testable decision

Test this model and alternate routes with the same request format

Start from a real workload, then let quality, total cost, and failure conditions decide whether to send production traffic.