Free AI APIs Compared: Recurring Free Tiers, Trial Credits, and Local Models

You are looking at a provider’s pricing page, or at a dashboard that just answered with 429, and free is doing three different jobs. A recurring free tier resets on a timer, and its allowance may be counted in requests, tokens, or credits. A one-time trial credit is a balance that only moves one direction. Local inference drops the per-request bill into hardware, power, and your own time. Which contract you are holding decides whether your prototype still answers next week, so here is what each one means, who offers one right now, and what the paid stage costs.
What does a free AI API actually mean?
Two questions settle it: what gets counted, and can the count go back up? Ask both before you compare model quality, because an offer that fails the second question is not the same kind of deal as one that passes it.

A recurring free tier resets on a timer
Each request, token, or credit is counted against a cap, and the counter starts over when the provider’s timer expires. The cap and the rhythm belong to the provider, so treat the published number as a ceiling and keep headroom for bursts.
Free tiers do not all count in the same unit:
- Requesty’s free tier allows 200 requests per day.
- Groq publishes per-model limits, such as 1,000 requests per day for specific free models.
- Hugging Face gives free accounts $0.10 per month in routed Inference Providers credit, subject to change.
A $0.10 monthly credit is still a recurring allowance, not a trial, because it comes back. What changes between those three is the measurement unit, not the category.
Which number bites depends on the shape of your traffic. If your prototype fires a few hundred small calls a day, a per-day request cap is what stops you; if it sends a few heavy prompts, a token or credit cap gets there first.
One-time trial credit spends down to zero
Trial credit is a balance, not an allowance. The provider grants a fixed amount once, and every call subtracts from it. When the balance reaches zero, the account stops accepting calls until you add funds.
Picture a signup offer that grants $5 of trial credit. That balance is finite no matter how lightly you use it, and nothing about it resets. Whether it fits comes down to the amount, the expiry date, and how fast your workload consumes it. A two-day evaluation can live on trial credit. A job that runs every night cannot.
Local inference swaps the bill for your own hardware
Nothing about local inference deletes the bill; the cost moves. A tool such as Ollama runs an open-weight model on a computer you control, so there is no hosted API key and no per-call fee for localhost.
The expense shows up elsewhere: a GPU or a large enough CPU, electricity while the model runs, license terms, and the engineering time spent installing, tuning, and troubleshooting a setup that a hosted endpoint would have returned in one request. Compare only locally installed models, because Ollama’s hosted cloud routes are APIs with their own billing.
Local is not an automatic privacy guarantee either. The tools and extensions around the model can still send data elsewhere unless you verify their settings.
Low cost is not the same as free
Some offerings sit between free and paid and should not be filed under free. DeepSeek’s API, for example, appears in provider lists as low cost or near-free rather than as a bounded free allowance with a reset, so treat it as a cheap paid option.
That distinction matters when you choose a prototype stack, because a small per-token price still generates a bill you did not plan for.
Which providers have a usable free AI API right now?
Five current offerings carry the comparison: OpenRouter free-variant models, Google AI Studio’s Gemini free tier, Groq’s Free plan, Requesty’s free tier, and Hugging Face’s monthly Inference Providers credit. Each is free under documented conditions, each resets or expires on a different rule, and none of them promises that the same model lineup will still be free next quarter.
A dated free AI API comparison list
Every row below names the provider page behind its limits and payment conditions. Model availability and account limits can change, so check that source before you commit a model ID to your prototype.
| Offer | What is free | Allowance and reset | Key and payment | Best fit for a prototype | At the limit | Source |
|---|---|---|---|---|---|---|
| OpenRouter free models | Free-variant chat and completion models whose model IDs carry a :free suffix, plus the openrouter/free router that selects an eligible free model for a request. |
OpenRouter’s pricing page listed its Free plan at 50 requests per day. Free-model limits also depend on account history: OpenRouter’s docs describe different free-model request limits for accounts that have purchased credits versus accounts that have not. Check the current plan limits against your account before estimating capacity. | OpenRouter account and API key. Key creation and free-variant calls work without a card, and buying credits changes the applicable free-model limit. | Chat and coding prototypes that want to compare several model families through one OpenAI-compatible key, accepting that model capability varies per variant. | A 429 response can come from the OpenRouter platform rate cap or from the upstream provider, and OpenRouter retries alternate providers for the same model before a provider-side error reaches you. A 402 response signals an account balance problem that funding resolves. | OpenRouter pricing page (openrouter.ai/pricing) |
| Google AI Studio Gemini free tier | Free-of-charge input and output tokens for the Gemini models that list a Free tier on the Gemini API pricing page, such as several Gemini Flash text models. Some models are paid only, including Gemini 3.1 Pro Preview and the video and music generators, so free access is not universal across the catalog. Google documents that free-tier calls may be used to improve its products. | Rate limits are per model and per project, with requests-per-day quotas resetting at midnight Pacific time. The exact RPM, TPM and RPD values for your account are shown in Google AI Studio rather than summarized on the pricing page. | Google account and a project in AI Studio. Google’s rate-limit guide describes the move from the Free tier to a paid tier as the step where billing is set up. | Chat, coding, vision and embedding experiments on current Gemini models that still carry a free tier, when the free-tier data-use term is acceptable. | Exceeding a model quota returns HTTP 429 RESOURCE_EXHAUSTED. Waiting for the reset window is the documented recovery, and the account’s own limits page in AI Studio is authoritative. | Gemini API pricing and rate limits (ai.google.dev/gemini-api/docs/pricing) |
| Groq Free plan | Hosted inference on the Groq models covered by the Free plan, including open models such as openai/gpt-oss-20b for text and code tasks plus speech-to-text models at their own listed rates. |
Per-model, per-organization limits. openai/gpt-oss-20b on the Free plan shows 30 requests per minute, 1,000 requests per day, 8K tokens per minute and 200K tokens per day. Other models have different numbers, so the Limits page for your organization is the authoritative source. |
Groq account and API key. Groq’s billing FAQ places the payment-method requirement at the upgrade from the Free tier to the Developer tier. | Chat and code generation prototypes where the needed model is in Groq’s free catalog and low-latency responses matter more than model variety. | Exceeding a limit returns HTTP 429 Too Many Requests. Groq sets the Retry-After header only when a rate limit is hit, and rate-limit headers report remaining daily requests and per-minute tokens. | Groq rate limits docs (console.groq.com/docs/rate-limits) |
| Requesty free tier | Free models served through an OpenAI-compatible endpoint, with routing, caching and EU data residency described as included by the provider’s free-model page. | 200 requests per day on free models. The provider describes the offer as a recurring daily allowance with no credit card and no trial timer. | Account key from the Requesty dashboard. The provider’s free-model page states that no credit card is required. | Steady, low-volume chat, coding and agent prototypes that call an OpenAI-compatible endpoint and expect to use the allowance every day. | The provider’s stated path past 200 requests per day is the paid plan on the same key with access to paid models. | Requesty free models page (requesty.ai/free-models) |
| Hugging Face Inference Providers credit | Every free Hugging Face account receives $0.10 in monthly credits, subject to change, that apply to requests Hugging Face routes to participating Inference Providers, with no separate provider account needed. | $0.10 per month, renewing on a monthly rhythm for routed requests. Free users can purchase additional credits for extra usage. A custom provider key setup is billed by that provider directly and receives no monthly credit. | Hugging Face account and a User Access Token. | Small routed experiments across open models for chat completion and feature extraction when the workload fits a $0.10 monthly credit. | Routed requests consume the credit first. Once it is spent, continuing on the same account requires purchased credits, and the monthly allowance returns on the next cycle. | Hugging Face Inference Providers pricing (huggingface.co/docs/inference-providers/pricing) |
For a local model, swap the allowance column for hardware, power, and maintenance costs, including the time to install the model, keep it available, and serve requests at your expected load.
What happens when a free allowance runs out?
How a free offer fails tells you whether a retry belongs in your code.
- Exhausted recurring allowance: the endpoint refuses further requests until its window resets, so retrying in the same minute does nothing useful.
- Trial balance at zero: it stays at zero, and no retry repairs it.
- Paid account out of funds: it needs a top-up, not a waiting game.
Requests and tokens can be capped separately, and Groq’s rate-limit guide walks through those dimensions. Take a worked example with hypothetical limits: an API allows 60 requests per minute and 120,000 counted tokens per minute. One complete task takes 4 API calls and 12,000 counted tokens across those calls. Requests alone allow 15 tasks per minute (60 / 4); tokens alone allow 10 (120,000 / 12,000). The lower number, 10, is the theoretical ceiling. A retry spends extra capacity, and daily caps, latency, concurrency and provider counting rules can pull real throughput below that ceiling.

Read the status code and the limit window before you react
An AI API error response is an incident report in miniature, so read it in order. The status code names the family: 429 for rate limiting, 402 for a payment problem in credit-based systems, 401 for an authentication failure. The message and headers name the member. OpenRouter’s docs explicitly separate a 429 that comes from its own platform rate cap from one the upstream provider returned, and an exhausted account balance in a credit system can surface as 402 even on a request that would otherwise be free.
Those three failures need three different reactions. Respect the reset or the Retry-After header for a transient or daily cap. Add funds, or move to a real free route, when the balance is spent. Fix the key when it is an auth error. A 429 is not always a daily quota counter, so read the headers instead of assuming.
Before you wire an automatic retry loop, read the 429 rate-limit fix that maps a status code and its headers to the real cause. That guide is the right next step when a free allowance, an upstream provider or a misconfigured client could each be producing the same status.

What will the same prototype cost once you leave the free tier?
Three inputs get you there: calls per day, input tokens per call, and output tokens per call. Multiply the model’s per-million-token input price by the input tokens you send, multiply the output price by the tokens the model returns, divide each product by one million, add them for the per-call cost, then multiply by your daily call volume.

A worked example with one consistent set of numbers
This example is hypothetical, including the token prices. Suppose the model page lists $2 per million input tokens and $10 per million output tokens. Your chatbot sends 400 input tokens per user turn and the model returns 700 output tokens. The input portion is 400 multiplied by $2 divided by 1,000,000, which is $0.0008. The output portion is 700 multiplied by $10 divided by 1,000,000, which is $0.007. One call costs $0.0078. At 1,000 calls per day that is $7.80 per day, and at the same volume for a 30-day month the bill is $234.
Provider token prices change, sometimes monthly, so the rates you plug in are only as fresh as the page you took them from. When you are ready to project against real rates, use the cost-estimation walkthrough that captures current token prices into a monthly projection instead of trusting a remembered number.
Will you rewrite code when you switch from a free API to a paid one later?
Often no, for compatible Chat Completions text calls. A client written against the OpenAI SDK keeps its message structure and its completion loop. The visible change is configuration: the base URL, the API key and the model name. The move becomes a config edit rather than a rewrite only after you verify the destination model and the parameters your code depends on.
What still differs after the switch
Compatibility does not mean identical. Model IDs are provider-specific, so a model slug from a free catalog may not exist on the paid destination. Context windows, structured-output support, vision handling, streaming behavior, optional headers and error formats can all differ between endpoints. Confirm that your model is in the target catalog, test the features the prototype actually uses, and only then flip the base URL and key in a test environment.
The same caution applies to the SDK in play. Requesty advertises Claude Code as a supported client, but Claude Code connects through its own Anthropic-style configuration using ANTHROPIC_BASE_URL and ANTHROPIC_AUTH_TOKEN, not through the OpenAI SDK that Requesty also offers. A tool that can consume an endpoint is not proof that it speaks the same client protocol as your code.
If you have no credit card, how do you keep building past the free tier?
Several current free tiers issue a key without a card:
- OpenRouter’s free-variant models hand out a key without a card.
- Requesty’s free tier states on its free-model page that no credit card is required.
- Groq’s Free tier needs only an account and an API key, with the payment-method requirement landing at the upgrade from Free to Developer.
Once a prototype outgrows those allowances, a prepaid account lets you keep building without attaching a card to a recurring charge. If the provider you choose accepts stablecoins, follow the crypto payment guide for funding an API account without a card.
MixRoute as the funded stage after free access
MixRoute sits at the paid end of this decision. It advertises no free tier. What it sells is prepaid credit, and its billing FAQ documents pay-as-you-go charging with model-specific token rates. A single top-up is not a monthly subscription; each request deducts from the balance, and credits do not expire, so money you load stays available until you spend it on tokens. Payment works by card or by USDT and USDC, which is the relevant path when a card is the obstacle. Once your prototype has real usage, compare the current prepaid tiers before you decide how much to fund, because the tier you choose changes the credit bonus on the same deposit.

FAQ
Is any AI API truly free with no credit card?
Yes. OpenRouter’s free-variant models and Requesty’s free tier both issue a key without a card, and Groq’s Free tier asks only for an account and an API key, with the payment-method requirement appearing at the upgrade to Developer. Hugging Face works differently, since the $0.10 monthly routed credit needs a User Access Token rather than a card. The pattern worth checking on any offer is where the card request shows up. If it appears at key creation, that is the offer’s anti-abuse check, and a small number of promotional programs do exactly that.
What is the difference between a free tier, trial credits, and local inference?
The reset rule separates them. A recurring free allowance counts back up on a schedule, so it is the only one of the three that survives a daily job. Trial credit only goes down, and zero is the end of it, which makes it a fit for a bounded evaluation rather than a long-running prototype. Local inference removes the per-request bill and puts the cost into hardware, electricity, and the hours you spend keeping the setup alive.
Do free AI API credits and allowances expire?
It depends on which contract you are holding. A recurring free tier generally lasts while the program and the model lineup continue, so what usually expires is a model variant, not your account. Promotional and trial credit carries its own expiry date, and only the offer page states it. Look for two things on that page: the quota and the valid-until date. Model lineups also rotate when a provider retires a variant, which is why every row in the comparison table above names its own source page.
What are the actual limits on OpenRouter free models?
OpenRouter’s pricing page listed its Free plan at 50 requests per day, and the :free suffix in a model ID is what marks a free variant. The limit that applies to your account can be different, because OpenRouter’s docs describe separate free-model limits for accounts that have purchased credits and accounts that have not. When a 429 comes back, the response headers say whether the platform cap or an upstream provider produced it, and that detail decides your next move.
Which free AI API should a prototype start with?
Match the offer to your call pattern. A workload that calls every day wants a recurring, OpenAI-compatible tier, which points at OpenRouter’s free variants or Requesty’s; a one-week evaluation can spend trial credit. Google AI Studio is the pick when you need current Gemini models for chat, vision, or embeddings and the free-tier data-use term is acceptable, and Groq fits when your model is already in its free catalog. Hugging Face covers small experiments across open models on $0.10 a month. Local inference earns its place once you can absorb the hardware, power, and operating hours. If data has to stay on your machine, check the surrounding tools too.
What happens when I hit a free AI API limit?
Read the status code and the window before you change anything. A 429 against a daily allowance means the reset time is the answer, or the Retry-After header when the response carries one. A 429 that an upstream provider produced can clear on its own. A 402 in a credit-based system means the balance has to go above zero, and one-time trial credit that reached zero stays there until you add funds. Retrying the same request body without reading the response repairs none of these.
Can I move from a free API to a paid account without rewriting my code?
For compatible Chat Completions calls, often yes. An OpenAI-SDK client keeps its message structure and its completion loop, and what changes is the base URL, the key, and the model name. The check that matters happens before the switch: confirm the destination catalog still contains your model and supports the parameters your code actually uses. Model IDs, context windows, structured-output support, streaming behavior, and error formats differ between providers, and a tool that can consume your endpoint is not proof that it speaks your client’s protocol.