Skip to content

Model detail

gpt-3.5-turbo

Provided by OpenAI
Pay-as-you-go

GPT-3.5 Turbo is OpenAI's fastest model. It can understand and generate natural language or code, and is optimized for chat and traditional completion tasks. Training data up to Sep 2021.

Model specs

Context length
16.385K
Max output
4.096K
I/O modalities
Text
Released
2023-05

02

API endpoints

One MixRoute gateway, OpenAI-compatible

  • OpenAI-compatible /v1/chat/completions POST

03

Pricing

Flat rate, unit: /1M Tokens

Input

$0.5000 /1M Tokens

Completion

$1.0000 /1M Tokens

05

Selection summary

Quickly judge whether it fits your workload

What is this model good for?

Fast, cost-effective chat and traditional completions

Use GPT-3.5 Turbo for chat and traditional completion tasks where speed and cost are prioritized over maximum accuracy. It is OpenAI's fastest model.

What should you check before using it?

Consider GPT-4o Mini for better performance

GPT-4o Mini offers significantly better intelligence at a comparable cost. Test your tasks to determine if GPT-3.5 Turbo still meets your needs.

Why use it through MixRoute?

Use the OpenAI-compatible endpoint with stable model ID

Use the confirmed compatible endpoint with model ID gpt-3.5-turbo through MixRoute.

08

Token cost estimator

Live estimate from this page's pricing, not an actual bill

This model has no cache-read price; caching is not counted

09

FAQ

GPT-3.5 Turbo is OpenAI’s fastest model, optimized for chat and traditional completion tasks. It can understand and generate natural language or code.

One endpoint, a testable decision

Test this model and alternate routes with the same request format

Start from a real workload, then let quality, total cost, and failure conditions decide whether to send production traffic.