What Is OpenRouter? One API for Hundreds of Models

The model ID is the first thing people misread about OpenRouter: the anthropic/ in anthropic/claude-sonnet-4.5 names the developer that released the model, not the company that will run your request. OpenRouter is a hosted router. You send an OpenAI-compatible chat-completion call to https://openrouter.ai/api/v1, the model field carries a catalog ID, and OpenRouter selects a hosting provider that serves that model, returns the answer in chat-completion format, and deducts the cost from your prepaid credit balance. One key and one balance cover every model in the catalog, and the provider behind any single call gets decided at request time.
Four things decide whether it fits your stack: where a request goes, which key and base URL you configure, how you pick a model ID, and what the credits and platform fee cost.
What is OpenRouter?
OpenRouter is a hosted API router. One OpenAI-compatible interface provides model access, provider routing, and consolidated billing, and OpenRouter works the access layer between your code and the hosting providers instead of replacing the models themselves. Your client sends every request to the same endpoint, and a provider OpenRouter selects runs the model.
The catalog is where model selection happens. Each entry maps an ID string such as anthropic/claude-sonnet-4.5 to a concrete model and lists context length, supported input types, and price. The prefix in that ID identifies the developer that released the model; it does not by itself name the provider that will host your request. Models such as meta-llama/llama-3.3-70b-instruct are served by several hosting providers under the same ID, which is exactly why a routing layer has to exist.
Making the service work takes three jobs:
- The catalog keeps model names, capabilities, and prices current across a wide set of providers.
- Routing picks an eligible hosting endpoint for the model you named and reacts when that endpoint fails.
- Billing keeps per-vendor accounts out of your code, so one OpenRouter key and one credit balance replace one credential and one bill per provider.
If you already keep three provider accounts and reconcile three invoices, the billing layer is the part you notice first. If you want to compare models you have never run before, the catalog carries more of the value.
How does an OpenRouter request reach a model?
Three hops, and only the first one is yours: your application, OpenRouter, and a hosting provider that serves the model you selected. Everything after hop one happens behind the same OpenAI-compatible endpoint, which is why a working prototype needs no per-provider code.

What each hop does
Your application goes first. Your client POSTs a JSON body to https://openrouter.ai/api/v1/chat/completions using an OpenAI-compatible chat-completion shape: a model field, a messages array, and optional parameters such as temperature, max_tokens, or stream. Chat Completions is the route in the example, not the only inference route OpenRouter exposes.
OpenRouter handles the middle hop. Picture a switchboard operator who can see which lines are busy before connecting the call, and you have the right picture of the routing layer. It reads the catalog ID, checks which providers can serve that model at that moment, and selects an endpoint. The default strategy is price-based load balancing that prioritizes lower-cost providers while skipping providers with significant outages in the last 30 seconds. You can replace that default with per-request provider settings, covered in the model-selection section below. OpenRouter then forwards your messages to the chosen provider’s API.
The hosting provider runs the selected model on its infrastructure, and that is the last hop. Tokens return from that provider to OpenRouter, which returns them to your client in the OpenAI-compatible chat-completion format and adds usage fields such as prompt_tokens, completion_tokens, and total_tokens. For a plain-text, nonstreaming call like the one in the next section, the format is the one the OpenAI SDK parses, so existing parsing code handles this case. Tool calls, refusals, errors, empty choices, multimodal content, and streaming events each need their own handler.
If the upstream provider accepts the request, that path holds. If the provider is down or returns an error, OpenRouter can try another provider that serves the same model before the error reaches you. What happens next depends on the model and on the routing parameters covered in the model-selection section.
How do you make your first OpenRouter API call?
For credit-paid calls you need one OpenRouter account, one API key, and a credit balance on the account. The Keys area of the dashboard creates the key and lets you name it. OpenRouter also lets you set an optional per-key credit cap, and the GET /api/v1/key endpoint reports the remaining limit, so you can monitor it before requests start failing.
Use the OpenRouter key with the OpenRouter endpoint. Check the available balance, any per-key spending cap, and the catalog ID before the test. If the call fails, the response body tells you whether you hit a payment or key limit or are looking at transient throttling.
A minimal Python call with the OpenAI SDK
Install the openai package, put your own key in the OPENROUTER_API_KEY environment variable, then run the pattern that OpenRouter’s quickstart documents. Relative to a native OpenAI setup, two values change: base_url and api_key. The model string changes too, because catalog IDs carry a developer prefix such as anthropic/ or google/.
pip install openai
export OPENROUTER_API_KEY="your-key-here" # use the key from your dashboardfrom openai import OpenAI
import os
client = OpenAI(
base_url="https://openrouter.ai/api/v1",
api_key=os.environ["OPENROUTER_API_KEY"],
)
completion = client.chat.completions.create(
model="anthropic/claude-sonnet-4.5",
messages=[
{"role": "user", "content": "Explain model routing in one paragraph."}
],
)
print(completion.choices[0].message.content)The example prints a text response from a successful nonstreaming call. Before you depend on that output, handle SDK exceptions and confirm that a text choice is present. Error objects and partial streams get their own treatment in the prebuild checklist below.
Copy the full ID from the model’s catalog page, developer prefix included, into model. That page also lists input types and price. Credit-paid calls use the OpenRouter balance; BYOK uses your own provider credentials and a separate fee arrangement.
This compatibility note covers the plain-text, nonstreaming completion shown above. OpenRouter is a drop-in replacement for OpenAI for the endpoints it documents, but that does not make every native provider API interchangeable. If a model has provider-specific parameters or response behavior, read the model’s documentation instead of trusting an SDK default to transmit or parse it.
How do you choose which model to call?
You choose a model per request by sending its catalog ID in the model field. Each model page shows the exact ID, context length, supported input types, and current price. When several providers host the same ID, the routing layer decides which provider handles the call.

Request-level provider controls and fallbacks
OpenRouter’s provider-routing documentation defines request-level controls for the chat-completions route:
onlyrestricts the request to a list of provider slugs.ignoreremoves listed providers from consideration.ordernames the sequence of providers to try first, while other providers stay eligible as fallbacks unless disabled.allow_fallbacks, true by default, controls whether a backup provider may serve the request when the primary provider is unavailable.
Two other controls steer data and compatibility rather than speed. data_collection: deny limits routing to providers that do not store your data for training, and require_parameters: true limits routing to providers that support every parameter in the request. If no provider satisfies those constraints, the request can fail instead of silently ignoring the requirement.
Model-level fallbacks are separate from provider-level ones. A top-level models array lists the IDs to try in order, so when every provider for the first model is exhausted, the router can move to the next model you included.
A fallback can run only while an eligible backup remains: another provider serving the same model, or a fallback model you listed and that is eligible under the request’s own constraints. If none exists, the failure returns to your application. Test that path instead of assuming the router absorbs every outage.
What does calling OpenRouter cost?
OpenRouter is pay as you go and has no required subscription. You add credits to an account balance, and every paid request deducts its cost from that balance. For most models the deduction is a listed price per million tokens, and OpenRouter’s FAQ notes that some models also charge per request, per image, or for reasoning tokens; the model page shows which applies. The FAQ adds that each model and provider has its own price per million tokens, and that OpenRouter passes through provider inference pricing without markup.

The platform fee is separate from model consumption. On the pricing page checked September 20, 2026, Standard lists 5.5% and Business lists 8%; enterprise terms can differ. The FAQ places the credit-purchase fee at top-up. Compare the plan and checkout total, including any applicable taxes or payment charges.
Free models and a free tier exist for evaluation but carry documented limits far below paid usage. Credits also carry conditions: OpenRouter’s terms reserve the right to expire unused credits one year after purchase, refunds of unused credits are limited to the first 24 hours, and cryptocurrency purchases are never refundable. Read those three conditions before you deposit a large balance.
The fee structure, row by row
Each row separates what you consume from what you pay OpenRouter.
| Fee component | How it is charged | What to verify |
|---|---|---|
| Provider model rate | For most models, a listed price per million tokens with separate prompt and completion prices; the model page also shows per-request, per-image, and reasoning-token pricing where a model uses it. | The provider sets the rate and can reprice, so read the model page at call time rather than trusting a saved quote. |
| OpenRouter platform fee | A percentage charged when you purchase credits for pay-as-you-go accounts, separate from the model rate and not applied per inference. | Standard: 5.5%; Business: 8%, checked September 20, 2026. Confirm the checkout total for your plan and payment method. |
| Credit terms | Unused credits can expire one year after purchase under OpenRouter’s terms; refunds of unused credits are limited to the first 24 hours after purchase. | Cryptocurrency purchases are never refundable, so read the billing FAQ before choosing a payment method. |
If the purchase fee changes your unit economics, the listed model rates and the credit-purchase fee of each candidate belong on the same sheet. When you want MixRoute’s fee model on that sheet, compare MixRoute prepaid pricing and let concrete rates, not category labels, make the call.
Is OpenRouter an AI gateway?
OpenRouter is a hosted model API router, and it performs the functions people usually group under an AI gateway: one interface for many models, catalog-based model selection, provider routing with fallbacks, and consolidated billing. Whether the gateway label applies depends on where the person using the term draws the boundary, and vendors draw it differently.
MixRoute describes itself as an AI API gateway focused on unified access, routing, reliability, and cost governance. OpenRouter describes itself as one API for any model. Vercel AI Gateway describes itself as a unified API to access hundreds of AI models. Three true descriptions with different labels is normal in this market, so a category name proves little about capabilities.
Ask functional questions instead:
- Does one OpenAI-compatible client reach the models you need?
- How is the fee expressed, and at what moment is it charged?
- Which provider controls exist per request?
- What happens when an upstream provider fails?
OpenRouter answers those with a model catalog, one credit balance, and request-level routing parameters. Before you sort products into categories, read how an AI API gateway is defined and apply the same yardstick to each vendor.
What should you check before you build on OpenRouter?
Much of the multi-provider integration work disappears once you route through OpenRouter, but a few checks separate a clean production path from surprises. None of them take long.
Check errors and partial output
With raw HTTP, a 200 status can precede a later generation error. Check a nonstreaming body for a top-level error and verify that choices contains the output you need. In streaming mode, handle error events as well as tokens. A retry after partial output starts a new generation; replace or clearly separate the partial result instead of appending fresh text as though it resumed the old response.
A short prebuild checklist
- Confirm the models exist in the catalog today. Search for the exact IDs you plan to ship with and note each model’s context length and supported inputs. If an ID is missing, that capability is not available through OpenRouter at the moment.
- Price a realistic month, not a single request. Estimate prompt and completion tokens per request, multiply each by the model’s listed per-million rates, and remember models that charge per request, per image, or for reasoning tokens. The platform fee lands when you buy credits instead of per request, so confirm the checkout total for your payment method.
- Decide what the application does when a provider fails. Prefer a model that multiple providers serve, or configure a fallback
modelsarray, then test the failure case instead of assuming the router absorbs it. - Do not size production from free models or the free tier. Those are for evaluation and carry tighter documented limits. A paid workload needs the credit balance and the purchase fee in the operating budget from day one.
- Compare before hardening an integration. The first code path tends to become the permanent one, and OpenRouter is one of several hosted routers, alongside direct provider contracts and self-hosted gateways. Review the OpenRouter alternatives breakdown against the same checklist before you commit.
It earns its place when model variety and one bill matter more than the credit-purchase fee.
FAQ
What is OpenRouter in one sentence?
OpenRouter is a hosted API router: a common interface to many models, with provider routing and consolidated billing behind it. The part that trips people up is the naming. A request’s model ID names the model, and routing picks the provider that serves it, so the developer in the ID is not always the company running your inference.
How is OpenRouter different from OpenAI?
OpenAI is one model provider, while OpenRouter is an access layer that can forward your request to OpenAI models and to models from other providers through the same endpoint. OpenAI’s API returns OpenAI models; OpenRouter’s catalog includes those alongside models from other developers. The difference shows up in billing, where an OpenAI model becomes one catalog entry among many, paid from the same OpenRouter balance as everything else, with the purchase fee applied when you top up.
Do I need separate accounts with OpenAI, Anthropic or Google to use OpenRouter?
No, for credit-paid calls an OpenRouter key and balance are enough, because OpenRouter holds the upstream billing relationship. The clearest way to see where the split falls is bring-your-own-key mode. There you connect direct provider credentials, and OpenRouter routes some or all of your traffic through your own accounts under a separate, plan-dependent fee, so read that mode’s terms before assuming it bills like the credit path.
How does OpenRouter charge for a request?
Two charges land at different moments. Model usage consumes prepaid credits at the listed rate, while buying those credits carries a separate platform fee. On September 20, 2026, Standard listed 5.5% and Business 8%. That split matters when you compare routers, because a per-request rate on its own hides the fee you pay at top-up.
Is OpenRouter free?
There are free models and a free tier, both aimed at evaluation. New accounts get a small free allowance, and the request limits on all of it sit far below paid usage, so anything near production needs credits. The catch is what credits are: OpenRouter’s terms allow unused credits to expire one year after purchase, refunds of unused credits are limited to the first 24 hours, and cryptocurrency purchases are never refundable.
Can I keep using the OpenAI SDK with OpenRouter?
Yes, for the calls OpenRouter documents. Set the base URL and key, send an available model ID, and check parameter and response handling; the quickstart shows the pattern with the OpenAI Python SDK on the chat-completion route. What does not carry over is the edge of the API. Tool calls, streaming, refusals, multimodal content and provider-specific parameters each need a handler that matches the actual response, and an SDK default may not match it.
What happens when the provider behind my model goes down or rate-limits me?
OpenRouter can move the request to another provider serving the same model, or to a fallback model you listed, but only while a usable backup exists. Request-level parameters govern that, and allow_fallbacks defaults to true, so a provider error can shift the request to another provider for the same model. With no eligible backup, the error reaches your application. Retries with backoff cover transient rate limits; persistent failures need to surface.