Model detail
rerank-v4.0-pro
Cohere's AI search foundation model for enhancing the relevance of information surfaced within search and RAG systems. Features a 32K context window, multilingual support across 100+ languages, no data pre-processing required, and state of the art performance with low latency.
Model specs
- Context length
- 32K
- Max output
- –
- I/O modalities
- Text
- Released
- 2026-04
02
API endpoints
One MixRoute gateway, OpenAI-compatible
-
OpenAI-compatible
/v1/chat/completionsPOST
03
Tiered pricing
Unit: /1M Tokens
| Tier |
|---|
| base |
05
Selection summary
Quickly judge whether it fits your workload
What is this model good for?
A multilingual model that allows for re-ranking English and non-english documents
A multilingual model that allows for re-ranking English and non-english documents and semi-structured data (JSON). This model is better suited for state-of-the-art quality and complex use-cases than its fast variant.
What should you check before using it?
32,000 context tokens
Model type: rerank. Verify supported features and quotas with the provider documentation.
Why use it through MixRoute?
One MixRoute key reaches rerank-v4.0-pro
One MixRoute key covers rerank-v4.0-pro, on the same account as the rest of the catalog. MixRoute serves rerank-v4.0-pro on the OpenAI-compatible endpoint /v1/chat/completions; point your existing client at https://api.mixroute.ai/v1 with the model ID rerank-v4.0-pro and re-run your integration tests.
09
FAQ
Before moving production traffic to rerank-v4.0-pro, confirm its context and token limits, price rows and current catalog status (stable). The record lists a 32,000-token context window; size the workload against those limits.
One endpoint, a testable decision
Test this model and alternate routes with the same request format
Start from a real workload, then let quality, total cost, and failure conditions decide whether to send production traffic.