Skip to content

Model detail

gpt-4.1-mini

Provided by OpenAI
Pay-as-you-go

GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million token context window and scores 45.1% on hard instruction evals, 35.8% on MultiChallenge, and 84.1% on IFEval. Mini also shows strong coding ability (e.g., 31.6% on Aider’s polyglot diff benchmark) and vision understanding, making it suitable for interactive applications with tight performance constraints.

Model specs

Context length
1.047576M
Max output
32.768K
I/O modalities
Text / Image
Released
2025-04

02

API endpoints

One MixRoute gateway, OpenAI-compatible

  • OpenAI-compatible /v1/chat/completions POST

03

Pricing

Flat rate, unit: /1M Tokens

Input

$0.4000 /1M Tokens

Completion

$0.8000 /1M Tokens

Cache read

$0.1000 /1M Tokens

05

Selection summary

Quickly judge whether it fits your workload

What is this model good for?

Cost-efficient long-context tasks with strong instruction following

Use GPT-4.1 Mini for interactive applications that need the 1M token context window at lower cost. It delivers competitive coding and instruction-following performance with substantially lower latency than GPT-4.1.

What should you check before using it?

Confirm performance meets your task requirements

Test your specific coding and instruction-following tasks. GPT-4.1 Mini trades some capability for speed and cost—verify it meets your accuracy thresholds before committing to production.

Why use it through MixRoute?

Use the OpenAI-compatible endpoint with stable model ID

Use the confirmed compatible endpoint with model ID gpt-4.1-mini through MixRoute, then run the same integration tests used for the current client.

08

Token cost estimator

Live estimate from this page's pricing, not an actual bill

Share of the same prefix read repeatedly, up to 100%

09

FAQ

GPT-4.1 Mini is a mid-sized model in the GPT-4.1 series, delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million token context window and shows strong coding and instruction-following ability.

One endpoint, a testable decision

Test this model and alternate routes with the same request format

Start from a real workload, then let quality, total cost, and failure conditions decide whether to send production traffic.