Skip to content

Model detail

gpt-4o-transcribe

Provided by OpenAI
Pay-as-you-go

Model specs

Context length
16K
Max output
2K
I/O modalities
Text
Released

02

API endpoints

One MixRoute gateway, OpenAI-compatible

  • OpenAI-compatible /v1/chat/completions POST

03

Pricing

Flat rate, unit: /1M Tokens

Input

$2.5000 /1M Tokens

Completion

$10.0000 /1M Tokens

Audio input

$6.0000 /1M Tokens

05

Selection summary

Quickly judge whether it fits your workload

What is this model good for?

High-accuracy speech-to-text transcription

Use GPT-4o Transcribe for speech-to-text tasks. Built on GPT-4o, it offers improved word error rate and better language recognition compared to the original Whisper model.

What should you check before using it?

Confirm supported audio formats and languages

Check supported audio formats, duration limits, and language coverage. Test your specific transcription tasks to verify accuracy improvements over Whisper.

Why use it through MixRoute?

Use the compatible endpoint with stable model ID

Use the confirmed compatible endpoint with model ID gpt-4o-transcribe through MixRoute.

08

Token cost estimator

Live estimate from this page's pricing, not an actual bill

This model has no cache-read price; caching is not counted

09

FAQ

GPT-4o Transcribe is a speech-to-text model that uses GPT-4o to transcribe audio. It offers improvements to word error rate and better language recognition compared to the original Whisper model.

One endpoint, a testable decision

Test this model and alternate routes with the same request format

Start from a real workload, then let quality, total cost, and failure conditions decide whether to send production traffic.