Skip to content

Model detail

gpt-5-mini-2025-08-07

Provided by OpenAI
Pay-as-you-go

GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and safety-tuning benefits as GPT-5, but with reduced latency and cost. GPT-5 Mini is the successor to OpenAI's o4-mini model.

Model specs

Context length
400K
Max output
128K
I/O modalities
Text / Image
Released
2025-08

02

API endpoints

One MixRoute gateway, OpenAI-compatible

  • OpenAI-compatible /v1/chat/completions POST

03

Pricing

Flat rate, unit: /1M Tokens

Input

$0.2500 /1M Tokens

Completion

$2.0000 /1M Tokens

Cache read

$0.0250 /1M Tokens

05

Selection summary

Quickly judge whether it fits your workload

What is this model good for?

Cost-efficient reasoning for everyday tasks

Use GPT-5 Mini for lighter-weight reasoning tasks where you need GPT-5's instruction-following quality at reduced latency and cost. It is the successor to o4-mini, offering improved safety tuning.

What should you check before using it?

Confirm reasoning depth meets your task needs

GPT-5 Mini trades some reasoning depth for speed and cost. Test your specific tasks to confirm the accuracy meets your requirements compared to the full GPT-5 model.

Why use it through MixRoute?

Use the OpenAI-compatible endpoint with stable model ID

Use the confirmed compatible endpoint with model ID gpt-5-mini through MixRoute, then run the same integration tests used for the current client.

08

Token cost estimator

Live estimate from this page's pricing, not an actual bill

Share of the same prefix read repeatedly, up to 100%

09

FAQ

GPT-5 Mini is a compact version of GPT-5 designed for lighter-weight reasoning tasks. It provides the same instruction-following and safety-tuning benefits as GPT-5 with reduced latency and cost. It is the successor to OpenAI’s o4-mini model.

One endpoint, a testable decision

Test this model and alternate routes with the same request format

Start from a real workload, then let quality, total cost, and failure conditions decide whether to send production traffic.