モデル詳細
gpt-5.4-nano
GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. It supports text and image inputs and is designed for low-latency use cases such as classification, data extraction, ranking, and sub-agent execution. The model prioritizes responsiveness and efficiency over deep reasoning, making it ideal for pipelines that require fast, reliable outputs at scale. GPT-5.4 nano is well suited for background tasks, real-time systems, and distributed agent architectures where minimizing cost and latency is essential.
モデル仕様
- コンテキスト長
- 400K
- 最大出力
- 128K
- 入出力モダリティ
- テキスト / 画像
- リリース
- 2026-03
02
API エンドポイント
MixRoute 単一ゲートウェイ、OpenAI 互換
-
OpenAI-compatible
/v1/chat/completionsPOST
03
料金
固定料金、単位: /1M Tokens
入力
$0.2000 /1M Tokens
生成
$1.2500 /1M Tokens
キャッシュ読み取り
$0.0200 /1M Tokens
05
選定サマリー
ワークロード適合性を素早く判断
What is this model good for?
Start with the documented use cases for gpt-5.4-nano
GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. It supports text and image inputs and is designed for low-latency use cases such as classification, data extraction, ranking, and sub-agent execution. The model prioritizes responsiveness and efficiency over deep reasoning, making it ideal for pipelines that require fast, reliable outputs at scale. GPT-5.4 nano is well suited for background tasks, real-time systems, and distributed agent architectures where minimizing cost and latency is essential.
What should you check before using it?
Validate limits, pricing, and a representative workload
Review the current model limits and pricing record before production use.
How do you call it through MixRoute?
Use the documented endpoint and exact model ID
The model record lists OpenAI-compatible access via POST /v1/chat/completions with model ID gpt-5.4-nano.
08
トークンコスト見積もり
このページの料金に基づく見積もりで、実請求ではありません
同一プレフィックスが繰り返し読み取られる入力の割合(最大100%)
単一エンドポイントで検証可能な判断
同じリクエスト形式でこのモデルと代替ルートをテスト
実際のワークロードから始め、品質、総コスト、障害条件を基に本番トラフィックを送るか判断します。