モデル詳細
gemini-3.1-flash-lite
Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, and PDF inputs, and is designed for lightweight agentic workflows, simple data extraction, and applications where responsiveness and API cost are the primary constraints. Supports full thinking levels (minimal, low, medium, high) for fine-grained cost/performance trade-offs. Priced at half the cost of Gemini 3 Flash.
モデル仕様
- コンテキスト長
- –
- 最大出力
- –
- 入出力モダリティ
- –
- リリース
- 2026-05
02
API エンドポイント
MixRoute 単一ゲートウェイ、OpenAI 互換
-
Gemini
/v1beta/models/{model}:generateContentPOST -
OpenAI-compatible
/v1/chat/completionsPOST
03
料金
固定料金、単位: /1M Tokens
入力
$0.2500 /1M Tokens
生成
$1.5000 /1M Tokens
キャッシュ読み取り
$0.0250 /1M Tokens
08
トークンコスト見積もり
このページの料金に基づく見積もりで、実請求ではありません
同一プレフィックスが繰り返し読み取られる入力の割合(最大100%)
単一エンドポイントで検証可能な判断
同じリクエスト形式でこのモデルと代替ルートをテスト
実際のワークロードから始め、品質、総コスト、障害条件を基に本番トラフィックを送るか判断します。