モデル詳細
deepseek-v4-flash
DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and high-throughput workloads, while maintaining strong reasoning and coding performance. The model includes hybrid attention for efficient long-context processing. Reasoning efforts `high` and `xhigh` are supported; `xhigh` maps to max reasoning. It is well suited for applications such as coding assistants, chat systems, and agent workflows where responsiveness and cost efficiency are important.
モデル仕様
- コンテキスト長
- –
- 最大出力
- –
- 入出力モダリティ
- –
- リリース
- 2026-04
02
API エンドポイント
MixRoute 単一ゲートウェイ、OpenAI 互換
-
OpenAI-compatible
/v1/chat/completionsPOST
03
料金
固定料金、単位: /1M Tokens
入力
$0.1400 /1M Tokens
生成
$0.2800 /1M Tokens
キャッシュ読み取り
$0.0028 /1M Tokens
08
トークンコスト見積もり
このページの料金に基づく見積もりで、実請求ではありません
同一プレフィックスが繰り返し読み取られる入力の割合(最大100%)
単一エンドポイントで検証可能な判断
同じリクエスト形式でこのモデルと代替ルートをテスト
実際のワークロードから始め、品質、総コスト、障害条件を基に本番トラフィックを送るか判断します。