Skip to content

Model detail

rerank-v4.0-pro

Provided by Cohere
Pay-as-you-go Dynamic pricing

Cohere's AI search foundation model for enhancing the relevance of information surfaced within search and RAG systems. Features a 32K context window, multilingual support across 100+ languages, no data pre-processing required, and state of the art performance with low latency.

Model specs

Context length
32K
Max output
–
I/O modalities
Text
Released
2026-04

02

API endpoints

One MixRoute gateway, OpenAI-compatible

  • OpenAI-compatible /v1/chat/completions POST

03

Tiered pricing

Unit: /1M Tokens

Tier
base

05

Selection summary

Quickly judge whether it fits your workload

What is this model good for?

A multilingual model that allows for re-ranking English and non-english documents

A multilingual model that allows for re-ranking English and non-english documents and semi-structured data (JSON). This model is better suited for state-of-the-art quality and complex use-cases than its fast variant.

What should you check before using it?

32,000 context tokens

Model type: rerank. Verify supported features and quotas with the provider documentation.

Why use it through MixRoute?

One MixRoute key reaches rerank-v4.0-pro

One MixRoute key covers rerank-v4.0-pro, on the same account as the rest of the catalog. MixRoute serves rerank-v4.0-pro on the OpenAI-compatible endpoint /v1/chat/completions; point your existing client at https://api.mixroute.ai/v1 with the model ID rerank-v4.0-pro and re-run your integration tests.

09

FAQ

Before moving production traffic to rerank-v4.0-pro, confirm its context and token limits, price rows and current catalog status (stable). The record lists a 32,000-token context window; size the workload against those limits.

One endpoint, a testable decision

Test this model and alternate routes with the same request format

Start from a real workload, then let quality, total cost, and failure conditions decide whether to send production traffic.