模型詳情
gpt-4.1-nano
For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size with its 1 million token context window, and scores 80.1% on MMLU, 50.3% on GPQA, and 9.8% on Aider polyglot coding – even higher than GPT‑4o mini. It’s ideal for tasks like classification or autocompletion.
模型規格
- 上下文長度
- 1.047576M
- 最大輸出
- 32.768K
- 輸入/輸出模態
- 文字 / 圖片
- 發布日期
- 2025-04
02
API 端點
單一 MixRoute 閘道,OpenAI 相容
-
OpenAI-compatible
/v1/chat/completionsPOST
03
價格
固定費率,單位:/1M Tokens
輸入
$0.1000 /1M Tokens
補全
$0.2000 /1M Tokens
快取讀取
$0.0250 /1M Tokens
05
選型摘要
快速判斷是否適合你的工作負載
What is this model good for?
Ultra-fast classification and autocompletion at scale
Use GPT-4.1 nano for high-volume tasks like classification, autocompletion, and data extraction. It delivers the 1M token context window at the lowest cost and latency in the GPT-4.1 series.
What should you check before using it?
Confirm accuracy for your task class
GPT-4.1 nano prioritizes speed and cost over deep reasoning. Test your classification and autocompletion tasks to confirm the accuracy meets your requirements before production deployment.
Why use it through MixRoute?
Use the OpenAI-compatible endpoint with stable model ID
Use the confirmed compatible endpoint with model ID gpt-4.1-nano through MixRoute, then run the same integration tests used for the current client.
08
Token 成本估算
基於本頁價格的即時估算,非實際帳單
同一前綴被重複讀取的輸入占比,最高 100%
單一端點,可驗證的決策
使用相同的請求格式測試此模型與替代路由
從真實工作負載開始,再由品質、總成本與故障條件決定是否導入正式環境流量。