跳至内容

模型详情

gemini-3-flash-preview

由 Google 提供
按量付费

Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers near Pro level reasoning and tool use performance with substantially lower latency than larger Gemini variants, making it well suited for interactive development, long running agent loops, and collaborative coding tasks. Compared to Gemini 2.5 Flash, it provides broad quality improvements across reasoning, multimodal understanding, and reliability. The model supports a 1M token context window and multimodal inputs including text, images, audio, video, and PDFs, with text output. It includes configurable reasoning via thinking levels (minimal, low, medium, high), structured output, tool use, and automatic context caching. Gemini 3 Flash Preview is optimized for users who want strong reasoning and agentic behavior without the cost or latency of full scale frontier models.

模型规格

上下文长度
最大输出
输入/输出模态
发布日期
2025-12

02

API 端点

单一 MixRoute 网关,OpenAI 兼容

  • Gemini /v1beta/models/{model}:generateContent POST
  • OpenAI-compatible /v1/chat/completions POST

03

价格

固定费率,单位:/1M Tokens

输入

$0.5000 /1M Tokens

补全

$3.0000 /1M Tokens

缓存读取

$0.0500 /1M Tokens

08

Token 成本估算

基于本页价格的实时估算,非实际账单

同一前缀被重复读取的输入占比,最高 100%

单一端点,可验证的决策

使用相同的请求格式测试此模型和备用路由

从真实工作负载开始,再由质量、总成本和故障条件决定是否接入生产流量。