Model detail
gpt-3.5-turbo
GPT-3.5 Turbo is OpenAI's fastest model. It can understand and generate natural language or code, and is optimized for chat and traditional completion tasks. Training data up to Sep 2021.
Model specs
- Context length
- 16.385K
- Max output
- 4.096K
- I/O modalities
- Text
- Released
- 2023-05
02
API endpoints
One MixRoute gateway, OpenAI-compatible
-
OpenAI-compatible
/v1/chat/completionsPOST
03
Pricing
Flat rate, unit: /1M Tokens
Input
$0.5000 /1M Tokens
Completion
$1.0000 /1M Tokens
05
Selection summary
Quickly judge whether it fits your workload
What is this model good for?
Fast, cost-effective chat and traditional completions
Use GPT-3.5 Turbo for chat and traditional completion tasks where speed and cost are prioritized over maximum accuracy. It is OpenAI's fastest model.
What should you check before using it?
Consider GPT-4o Mini for better performance
GPT-4o Mini offers significantly better intelligence at a comparable cost. Test your tasks to determine if GPT-3.5 Turbo still meets your needs.
Why use it through MixRoute?
Use the OpenAI-compatible endpoint with stable model ID
Use the confirmed compatible endpoint with model ID gpt-3.5-turbo through MixRoute.
08
Token cost estimator
Live estimate from this page's pricing, not an actual bill
This model has no cache-read price; caching is not counted
09
FAQ
GPT-3.5 Turbo is OpenAI’s fastest model, optimized for chat and traditional completion tasks. It can understand and generate natural language or code.
One endpoint, a testable decision
Test this model and alternate routes with the same request format
Start from a real workload, then let quality, total cost, and failure conditions decide whether to send production traffic.