模型详情
gpt-4.1-nano
For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size with its 1 million token context window, and scores 80.1% on MMLU, 50.3% on GPQA, and 9.8% on Aider polyglot coding – even higher than GPT‑4o mini. It’s ideal for tasks like classification or autocompletion.
模型规格
- 上下文长度
- 1.047576M
- 最大输出
- 32.768K
- 输入/输出模态
- 文本 / 图像
- 发布日期
- 2025-04
02
API 端点
单一 MixRoute 网关,OpenAI 兼容
-
OpenAI-compatible
/v1/chat/completionsPOST
03
价格
固定费率,单位:/1M Tokens
输入
$0.1000 /1M Tokens
补全
$0.2000 /1M Tokens
缓存读取
$0.0250 /1M Tokens
05
选型摘要
快速判断是否适合你的工作负载
What is this model good for?
Ultra-fast classification and autocompletion at scale
Use GPT-4.1 nano for high-volume tasks like classification, autocompletion, and data extraction. It delivers the 1M token context window at the lowest cost and latency in the GPT-4.1 series.
What should you check before using it?
Confirm accuracy for your task class
GPT-4.1 nano prioritizes speed and cost over deep reasoning. Test your classification and autocompletion tasks to confirm the accuracy meets your requirements before production deployment.
Why use it through MixRoute?
Use the OpenAI-compatible endpoint with stable model ID
Use the confirmed compatible endpoint with model ID gpt-4.1-nano through MixRoute, then run the same integration tests used for the current client.
08
Token 成本估算
基于本页价格的实时估算,非实际账单
同一前缀被重复读取的输入占比,最高 100%
单一端点,可验证的决策
使用相同的请求格式测试此模型和备用路由
从真实工作负载开始,再由质量、总成本和故障条件决定是否接入生产流量。