モデル詳細
qwen3.5-flash
The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. Compared to the 3 series, these models deliver a leap forward in performance for both pure text and multimodal tasks, offering fast response times while balancing inference speed and overall performance.
モデル仕様
- コンテキスト長
- –
- 最大出力
- –
- 入出力モダリティ
- –
- リリース
- 2026-02
02
API エンドポイント
MixRoute 単一ゲートウェイ、OpenAI 互換
-
OpenAI-compatible
/v1/chat/completionsPOST
03
料金
固定料金、単位: /1M Tokens
入力
$0.1000 /1M Tokens
生成
$0.4000 /1M Tokens
キャッシュ読み取り
$0.0100 /1M Tokens
キャッシュ作成
$0.1250 /1M Tokens
08
トークンコスト見積もり
このページの料金に基づく見積もりで、実請求ではありません
同一プレフィックスが繰り返し読み取られる入力の割合(最大100%)
単一エンドポイントで検証可能な判断
同じリクエスト形式でこのモデルと代替ルートをテスト
実際のワークロードから始め、品質、総コスト、障害条件を基に本番トラフィックを送るか判断します。