Model detail
gpt-4o-transcribe
Model specs
- Context length
- 16K
- Max output
- 2K
- I/O modalities
- Text
- Released
- –
02
API endpoints
One MixRoute gateway, OpenAI-compatible
-
OpenAI-compatible
/v1/chat/completionsPOST
03
Pricing
Flat rate, unit: /1M Tokens
Input
$2.5000 /1M Tokens
Completion
$10.0000 /1M Tokens
Audio input
$6.0000 /1M Tokens
05
Selection summary
Quickly judge whether it fits your workload
What is this model good for?
High-accuracy speech-to-text transcription
Use GPT-4o Transcribe for speech-to-text tasks. Built on GPT-4o, it offers improved word error rate and better language recognition compared to the original Whisper model.
What should you check before using it?
Confirm supported audio formats and languages
Check supported audio formats, duration limits, and language coverage. Test your specific transcription tasks to verify accuracy improvements over Whisper.
Why use it through MixRoute?
Use the compatible endpoint with stable model ID
Use the confirmed compatible endpoint with model ID gpt-4o-transcribe through MixRoute.
08
Token cost estimator
Live estimate from this page's pricing, not an actual bill
This model has no cache-read price; caching is not counted
09
FAQ
GPT-4o Transcribe is a speech-to-text model that uses GPT-4o to transcribe audio. It offers improvements to word error rate and better language recognition compared to the original Whisper model.
One endpoint, a testable decision
Test this model and alternate routes with the same request format
Start from a real workload, then let quality, total cost, and failure conditions decide whether to send production traffic.