Skip to content

Model detail

gpt-4.1

Provided by OpenAI
Pay-as-you-go

GPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-context reasoning. It supports a 1 million token context window and outperforms GPT-4o and GPT-4.5 across coding (54.6% SWE-bench Verified), instruction compliance (87.4% IFEval), and multimodal understanding benchmarks. It is tuned for precise code diffs, agent reliability, and high recall in large document contexts, making it ideal for agents, IDE tooling, and enterprise knowledge retrieval.

Model specs

Context length
1.047576M
Max output
32.768K
I/O modalities
Text / Image
Released
2025-04

02

API endpoints

One MixRoute gateway, OpenAI-compatible

  • OpenAI-compatible /v1/chat/completions POST

03

Pricing

Flat rate, unit: /1M Tokens

Input

$2.0000 /1M Tokens

Completion

$4.0000 /1M Tokens

Cache read

$0.5000 /1M Tokens

05

Selection summary

Quickly judge whether it fits your workload

What is this model good for?

Long-context reasoning and software engineering

Use GPT-4.1 for tasks requiring a 1M token context window, including long-document analysis, complex codebases, and enterprise knowledge retrieval. It excels at instruction following and agent reliability.

What should you check before using it?

Plan for the 1M context window and instruction-following needs

Confirm your workload benefits from the 1M token context window. Test instruction-following accuracy on your specific prompts and verify agent reliability for your most demanding workflows.

Why use it through MixRoute?

Use the OpenAI-compatible endpoint with stable model ID

Use the confirmed compatible endpoint with model ID gpt-4.1 through MixRoute, then run the same integration tests used for the current client.

08

Token cost estimator

Live estimate from this page's pricing, not an actual bill

Share of the same prefix read repeatedly, up to 100%

09

FAQ

GPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-context reasoning. It supports a 1 million token context window and outperforms GPT-4o across coding and instruction compliance benchmarks.

One endpoint, a testable decision

Test this model and alternate routes with the same request format

Start from a real workload, then let quality, total cost, and failure conditions decide whether to send production traffic.