Skip to content

Model detail

gpt-5-pro-2025-10-06

Provided by OpenAI
Pay-as-you-go

GPT-5 Pro is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and accuracy in high-stakes use cases. It supports test-time routing features and advanced prompt understanding, including user-specified intent like "think hard about this." Improvements include reductions in hallucination, sycophancy, and better performance in coding, writing, and health-related tasks.

Model specs

Context length
400K
Max output
272K
I/O modalities
Text / Image
Released
2025-10

02

API endpoints

One MixRoute gateway, OpenAI-compatible

  • OpenAI-compatible /v1/chat/completions POST

03

Pricing

Flat rate, unit: /1M Tokens

Input

$15.0000 /1M Tokens

Completion

$120.0000 /1M Tokens

05

Selection summary

Quickly judge whether it fits your workload

What is this model good for?

Complex reasoning with test-time routing

Use gpt-5-pro-2025-10-06 for complex tasks requiring step-by-step reasoning and high accuracy. It supports test-time routing and advanced prompt understanding with reduced hallucination.

What should you check before using it?

Confirm the snapshot date matches your stability needs

This is the 2025-10-06 snapshot of GPT-5 Pro. Verify that this version provides the stability you need for your high-stakes tasks.

Why use it through MixRoute?

Use the OpenAI-compatible endpoint with stable model ID

Use the confirmed compatible endpoint with model ID gpt-5-pro-2025-10-06 through MixRoute for a pinned version.

08

Token cost estimator

Live estimate from this page's pricing, not an actual bill

This model has no cache-read price; caching is not counted

09

FAQ

gpt-5-pro is intended for complex reasoning, high-stakes knowledge work, code quality, and tasks that need careful instruction following and accuracy.

One endpoint, a testable decision

Test this model and alternate routes with the same request format

Start from a real workload, then let quality, total cost, and failure conditions decide whether to send production traffic.