Last updated: 11 September 2026. Every limit on this page was checked against the provider's own docs that day.
Seven LLM APIs are free in September 2026 with no credit card: Google Gemini, Groq, OpenRouter, Cloudflare Workers AI, Mistral, Cohere and Hugging Face. Two more, NVIDIA's API catalog and Z.ai, give free access for prototyping without asking for payment on their pages. Everyone else is either a one-time trial or wants money before the first request.
The list moved a lot since June. Cerebras turned its free tier into a $5 trial that needs a card, GitHub Models shut down, Groq dropped Llama from its free plan, Together stopped giving trial credit, and every model OpenRouter offered free in June is paid now. The timeline further down has the dates. Here is the map as it stands.
Free LLM APIs with no credit card
| Provider | What you get free | Main limit | The catch |
|---|---|---|---|
| Google Gemini API | Gemini 3.x Flash, 2.5 Flash, Gemma 4 | Shown per project in AI Studio | Free-tier prompts may be used for training |
| Groq | gpt-oss-120b, gpt-oss-20b, Qwen 27B | 30 req/min, 1,000 req/day, 200K tokens/day | Small 8K tokens/min cap |
| OpenRouter | 19 :free models | 20 req/min, 50 req/day | Upstream providers return 429s |
| Cloudflare Workers AI | Llama 3.3 70B, gpt-oss-120b, Qwen3, Kimi K2.5 and more | 10,000 Neurons/day | Budget is in Neurons, not tokens |
| Mistral | Mistral API models | $10/month of credits | Training on by default in Free mode |
| Cohere | Every Cohere model, Command A included | 1,000 calls/month, 20 req/min | No production or commercial use |
| Hugging Face | Models from 18 inference partners | $0.10/month of credits | 40K to 700K output tokens, by model |
"No card" here means the provider's own pages say you can call the API without adding a payment method. Limits reset daily unless the table says monthly.
All 17 providers compared
| Provider | Free offer today | Card | Status vs June |
|---|---|---|---|
| Google Gemini API | Free tier, limits in AI Studio | No | Gemini 2.0 retired, 3.x Flash added |
| Groq | 30 RPM, 1K RPD on gpt-oss and Qwen | No | Llama removed from free plan |
| OpenRouter | 19 free models, 50 or 1,000 req/day | No | Whole free roster replaced |
| Cloudflare Workers AI | 10,000 Neurons/day | No | New to this list |
| Mistral | $10/month Free-plan credits | No | Was listed as card-required trial |
| Cohere | Trial key, 1,000 calls/month | No | Was listed as card-required trial |
| Hugging Face | $0.10/month credits | No | Smaller than "free tier" suggests |
| NVIDIA API catalog | Up to 40 RPM on free endpoints | Not asked on its pages | Prototyping only |
| Z.ai | GLM-4.7-Flash, GLM-4.5-Flash free | Not stated | New to this list |
| Alibaba Model Studio | ~1M tokens per model, 90 days | Not stated | New to this list |
| Fireworks AI | $1 one-time credit | No, trial only | New to this list |
| Cerebras | $5 for 30 days | Yes | Was free with no card |
| OpenAI | None; data-sharing tokens if eligible | Yes, $5 prepay | No signup credit |
| Anthropic Claude | None; OSS program gives Max 20x | Yes | No documented signup credit |
| DeepSeek | None, very cheap per token | Top-up | Now DeepSeek V4.1 Flash |
| xAI Grok | None in official docs | Prepaid | Grok 4.6 |
| Together AI | None, $5 minimum purchase | Yes | Trial credit ended |
Heads up, these free tiers change almost monthly. Since June alone, Cerebras, GitHub Models, Together and SambaNova dropped their free offers and Groq cut its model list. I keep this list updated. Subscribe with your email and I'll send new free APIs as they launch.
Google Gemini API
- Free models: Gemini 3.8, 3.7, 3.6 and 3.5 Flash, 3.5 and 3.1 Flash-Lite, Gemini 2.5 Flash, Gemma 4 (31B and 26B A4B), embeddings, plus live, TTS and transcription models. Gemma 4 has no paid tier at all. The pricing page also lists 2.5 Pro and 2.5 Flash-Lite as free, but on 11 September both told our key they are "no longer available to new users".
- Not free: Gemini 3.1 Pro preview, the Nano Banana image models, Veo and Lyria.
- Rate limits: Google no longer publishes free-tier numbers in its docs. The rate limits page now sends you to AI Studio, where each project sees its own RPM, TPM and daily caps. The daily quota resets at midnight Pacific. The "1,500 requests a day at 10 RPM" figures quoted around the web, including by this page in June, are no longer official.
- Grounding with Google Search: free up to 500 requests a day, shared between 2.5 Flash and Flash-Lite.
- Card required: No. A Google account is enough; billing is only needed to move to a paid tier.
- The $300 Google Cloud trial no longer pays for the Gemini API in AI Studio on accounts opened after 2 March 2026, and the paid tier now starts with a $5 prepayment. Details in Gemini Free Credits.
- Your data: on the free tier Google may use prompts and responses to improve its products, and human reviewers may read them. The paid tier does not.
- Regions: the free tier works in the EEA, UK and Switzerland, but the terms allow only paid services if you serve end users there.
- Start: aistudio.google.com -> Get API key.
Still the best general-purpose free baseline: current Flash models, a large context window, and no card. Details on credits and the Google Cloud side are in Gemini Free Credits.
Groq
Groq's free plan now runs on OpenAI's open-weight gpt-oss models and Qwen. Llama 3.3 70B and Llama 3.1 8B left the free and developer tiers on 16 August 2026 and are enterprise-only now.
| Model | Req/min | Req/day | Tokens/min | Tokens/day |
|---|---|---|---|---|
openai/gpt-oss-120b | 30 | 1,000 | 8,000 | 200,000 |
openai/gpt-oss-20b | 30 | 1,000 | 8,000 | 200,000 |
qwen/qwen3.6-27b, qwen/qwen3.8-27b | 30 | 1,000 | 8,000 | 200,000 |
groq/compound, groq/compound-mini | 30 | 250 | 70,000 | - |
whisper-large-v3, -turbo (speech to text) | 20 | 2,000 | - | - |
Source: Groq's rate limits page, checked 11 September 2026.
- Card required: No. A payment method is only asked for when you upgrade to the Developer tier.
- Your data: not retained by default; logs may be kept up to 30 days for abuse cases, with an opt-out.
- The catch: 8,000 tokens a minute means one long prompt can use the whole minute. Groq is for fast, short, interactive calls, not bulk jobs.
- Start: console.groq.com
More on paid tiers in Groq Pricing.
OpenRouter
- Free models: 19 text models with the
:freesuffix (NVIDIA Nemotron 3, Google Gemma 4, Poolside Laguna, inclusionAI Ling and others), plus theopenrouter/freerouter that picks one for you. - Limits: 20 requests a minute; 50 a day, or 1,000 a day once you have bought $10 of credits at any point.
- Card required: No.
- The catch: each free model is served by an upstream provider with its own capacity. When it is busy you get a 429 even if you made one request all day. In our test of all 19 on 11 September, 4 were refused this way.
Not one of the models OpenRouter offered free in June (DeepSeek R1, Llama 3.3 70B, Qwen3 Coder...) is free there today. We re-test the free list every day in OpenRouter Free Tier, with today's result for each model.
Cloudflare Workers AI
- Free allowance: 10,000 Neurons a day on both the free and paid Workers plans, reset at 00:00 UTC. A Neuron is Cloudflare's compute unit; each model has a Neuron price per million tokens.
- What that buys, counting output tokens only: about 49,000 tokens a day on Llama 3.3 70B, about 147,000 on gpt-oss-120b, about 287,000 on Llama 3.1 8B. Our arithmetic from Cloudflare's pricing page.
- Free models: Llama 3.3 70B, Llama 3.1 8B, Llama 4 Scout, gpt-oss-120b and 20b, Qwen3 30B, QwQ 32B, Qwen2.5 Coder 32B, Gemma 3 and Gemma 4, Mistral Small 3.1, Nemotron 3 120B, Kimi K2.5, GLM-4.7-Flash and more. Kimi K2.6 and K2.7 Code, GLM-5.x and DeepSeek V4 need a paid plan.
- Rate limits: 300 requests a minute by default for text generation; 20 a minute on the larger frontier models.
- Card required: No.
- Your data: Cloudflare does not use your content to train models.
- Start: a free Cloudflare account, then the REST API or a Worker binding. There is an OpenAI-compatible endpoint.
This is where Llama 3.3 70B is still free after Groq and OpenRouter dropped it.
Mistral
- Free offer: new accounts start in Free mode, "no credit card required", and the Free plan includes $10 a month of credits shared between the API, Studio and Mistral's apps.
- Rate limits: requests per second, tokens per minute and tokens per month apply, and Free mode has the lowest. Mistral no longer publishes the numbers; you see them in the admin panel.
- Models: the API list includes Mistral Large, Medium, Small and Codestral; which of them Free mode covers is shown in your console.
- Your data: in Free mode Mistral may use your inputs and outputs for training unless you opt out.
- Start: console.mistral.ai
Our June version said Mistral needed a card. Its own docs say otherwise today.
Cohere
| Endpoint | Trial key limit |
|---|---|
| All calls | 1,000 per month |
| Chat, any model | 20 requests/min |
| Embed | 2,000 inputs/min |
| Rerank | 10 requests/min |
- Models: a trial key reaches every Cohere model: Command A+, Command A, Command A Reasoning, Vision and Translate, Command R+ and R, and North Mini Code.
- Card required: No. The trial key is created at signup.
- The catch: trial keys may not be used for production or commercial purposes.
- Best for: trying Rerank and Embed for RAG before paying.
Hugging Face Inference Providers
- Free offer: $0.10 of credit a month on a free account, spent at the partner providers' own prices (Cerebras, Groq, Together, Fireworks, Novita, Scaleway, Z.ai and others). PRO costs $9 a month and includes $2.
- Card required: No for the monthly $0.10; yes to buy more.
- The catch: at partner prices, ten cents is about 590,000 output tokens on gpt-oss-120b at the cheapest provider, or 38,000 on DeepSeek V4 Pro. Enough to check that a model works in your code, not to run anything.
Our June version called this a "rate-limited free tier". The real limit is the ten cents. Full breakdown in Hugging Face Inference API.
NVIDIA API catalog (build.nvidia.com)
- Free offer: models tagged "Free Endpoint", at up to 40 requests a minute. NVIDIA says limits vary by model and your exact numbers are shown in your account.
- Models: a large catalog that includes DeepSeek V4 Pro and V4 Flash, Kimi K3, MiniMax M3, Nemotron 3 Ultra and Nemotron 3.5, among about 80 hosted models.
- Conditions: you join the free NVIDIA Developer Program, and the endpoints are for "prototyping, research, development and testing purposes only". Production use needs an NVIDIA AI Enterprise license.
- Card required: not mentioned on any of NVIDIA's pages; you create and verify an account.
- Start: build.nvidia.com
The easiest free way to try frontier open models such as DeepSeek V4 Pro before you pick a paid host.
Z.ai
Z.ai's pricing page lists GLM-4.7-Flash, GLM-4.5-Flash and GLM-4.6V-Flash (vision) as free for input and output. It does not state the rate limits or whether a card is needed, so treat it as a model you can try rather than a quota you can plan on.
One-time trial credits
These give you something free once, then stop.
- Alibaba Cloud Model Studio: about 1,000,000 free tokens per model for new users, each model with its own quota, valid 90 days (for activations from 8 September 2026). Singapore region with International deployment only, real-time inference only. The quota page implies you can use it before completing billing details. Source.
- Fireworks AI: $1 of credit with no card, at 10 requests a minute until you add a payment method. When the dollar is spent the account is suspended until you add one.
- Cerebras: since 16 July 2026 there is no permanent free tier. You get a $5 credit, valid 30 days, only after adding a verified payment method. On the trial,
gpt-oss-120bandqwen-3.8-27brun at 5 requests a minute and 1 million tokens a day, with context capped at about 64K. Still very fast; no longer free in the sense this page means.
Also checked and left out of the tables: Vercel AI Gateway (a monthly free credit, but it needs a payment method), Scaleway (1 million free tokens, card required), SambaNova (its docs still describe a free tier, but new accounts are told to buy credits), Nebius ($1 for 30 days with a card), Chutes and Kimi (no free tier).
Pay first: OpenAI, Anthropic, DeepSeek, xAI, Together
OpenAI. No free credit for new API accounts; you prepay at least $5. What remains is the data-sharing program: eligible organizations that agree to share API traffic get free tokens every day, 250,000 on flagship models and 2.5 million on small ones at usage tiers 1-2, 1 million and 10 million at tiers 3-5. It covers the GPT-5.x line (including gpt-5.6-sol, terra and luna) but not the new flagship gpt-6-astra, and it needs a positive balance. Startup programs are in OpenAI Free Credits.
Anthropic Claude. No documented starter credit: Anthropic's help center says you buy credits before using the API. Current models are Claude Fable 5.1 ($10 / $50 per million input / output tokens), Opus 5 ($5 / $25), Sonnet 5 ($2 / $10) and Haiku 4.5 ($1 / $5). The free route is Claude for Open Source: six months of Claude Max 20x for qualifying maintainers, up to 10,000 recipients, still open. It is a subscription, not API credit, and eligibility widened this year (merged PRs in other projects, external contributors, OpenSSF criticality score). See Claude Free Credits and Claude API Pricing.
DeepSeek. No free grant, but the cheapest serious model here: deepseek-flash (V4.1 Flash) costs $0.15 per million input tokens and $0.60 output off-peak, twice that at peak hours; V4-Pro is $0.66 / $1.98 off-peak. Both have 1M context. You top up by card, PayPal, Alipay or WeChat Pay.
xAI Grok. xAI's current docs mention only prepaid credits and promo codes. Third-party sites report a $25 signup credit and $150 a month for sharing data; we could not find either on xAI's own pages. The flagship is grok-4.6 at $2 / $6.
Together AI. "Together AI does not currently offer free trials." A $5 minimum purchase unlocks the API; purchased credits do not expire.
What changed since June 2026
| Date | Change |
|---|---|
| 1 June | Google shuts down Gemini 2.0 Flash and Flash-Lite; 3.6 Flash and 3.1 Flash-Lite replace them |
| 16 June | GitHub Models closes to new customers |
| 17 June | Groq announces the end of Llama 3.3 70B and Llama 3.1 8B |
| 16 July | Cerebras replaces its free tier with a $5, 30-day trial that needs a card |
| 30 July | GitHub Models is retired; API, playground and BYOK are gone |
| 16 August | Llama 3.3 70B and 3.1 8B leave Groq's free and developer tiers |
| 17 August | Cerebras deprecates GLM-4.7; Qwen 3.8 27B takes its trial slot |
| August | SambaNova stops giving credits to new accounts |
| 8 September | Alibaba's new-user quota becomes valid for 90 days |
| By September | None of OpenRouter's June free models are free any more; the roster is 19 new ones |
Who went which way, compared with this page's June version:
| Provider | June | September | Direction | Full review |
|---|---|---|---|---|
| Cerebras | Free, no card, 1M tokens/day | $5 trial, card required | Went paid | - |
| GitHub Models | Free with a GitHub account | Retired | Gone | - |
| Together AI | Small starter credit | $5 minimum purchase | Went paid | - |
| Groq | Free Llama 3.x | Free gpt-oss and Qwen only | Fewer models | Groq Pricing |
| OpenRouter | ~28 free models | 19 free models, all different | Replaced | OpenRouter Free Tier |
| Anthropic | ~$5 starter credit | None documented | Went paid | Claude Free Credits |
| OpenAI | Inconsistent ~$5 credit | None; $5 prepay | Went paid | OpenAI Free Credits |
| Hugging Face | "Rate-limited free tier" | $0.10/month | Clarified: tiny | Hugging Face Inference API |
| Mistral | Trial credits, card | $10/month, no card | Better | - |
| Cohere | Trial credits, card | 1,000 calls/month, no card | Better | - |
| Cloudflare Workers AI | Not listed | 10,000 Neurons/day, no card | Added | - |
| Google Gemini | 2.5 Flash free | 3.x Flash and 2.5 Flash free, limits unpublished | New models | Gemini Free Credits |
Some rows reflect a real change on a known date; others (Mistral, Cohere, Hugging Face) mean our June version was wrong or the provider changed quietly. We re-check this page monthly.
How to stack free LLM APIs
Each provider counts its limits separately, so a small product can run on several free tiers at once:
- Default to Gemini Flash for general chat, extraction and instruction following.
- Groq for latency-sensitive calls, such as voice or streaming chat, with short prompts that fit its 8K tokens a minute.
- Cloudflare Workers AI when you want Llama 3.3 70B or gpt-oss inside a Worker, with no card.
- OpenRouter as the fallback: when a primary provider rate-limits you, send the same request to
openrouter/free. - Mistral and Cohere for evaluation: their free allowances are monthly, so save them for tests, not traffic.
- Pay for the hard 10 percent. Keep a small balance on OpenAI, Anthropic or DeepSeek for the calls where quality matters.
For the routing layer, see LLM Gateway 2026, which compares OpenRouter, LiteLLM, Portkey and Helicone. For ranking models on output quality, see Best Free LLM 2026. To stop depending on free tiers altogether, Best Local LLM covers running open models on your own hardware.
Frequently asked questions
Is there a truly free LLM API with no credit card? Yes. In September 2026, Google Gemini, Groq, OpenRouter, Cloudflare Workers AI, Mistral, Cohere and Hugging Face all let you call their API without adding a card. Gemini, Groq and Cloudflare give the most daily usage; Mistral and Cohere give a monthly allowance; Hugging Face gives $0.10 a month.
Is there a free LLM API with no limits? No. Every free tier has a cap: requests per minute, requests or tokens per day, or a monthly credit. The way to get more is to use several free tiers side by side, since each one counts its limits separately, or to run an open model locally.
Which free LLM API is the most generous in 2026? For daily volume without a card: Groq (1,000 requests and 200,000 tokens a day per model on gpt-oss-120b), Cloudflare Workers AI (10,000 Neurons a day, about 147,000 output tokens on gpt-oss-120b), and Gemini, whose limits Google now shows only inside AI Studio. Cerebras used to lead with 1 million tokens a day but became a card-required trial in July 2026.
Can I combine multiple free LLM APIs to get more inference? Yes. Each provider has independent rate limits, so routing across Gemini, Groq, Cloudflare and OpenRouter multiplies your free capacity. A gateway such as OpenRouter or LiteLLM, or a small model registry of your own, handles the fallback when one of them returns a 429.
Does OpenAI give free API credits? Not to new accounts: you prepay at least $5. Eligible organizations can get free daily tokens by sharing their API traffic with OpenAI, up to 1 million tokens a day on flagship GPT-5.x models and 10 million on small ones at higher usage tiers.
Is the Claude API free? No. Anthropic's help center says you buy credits before using the API, and no signup credit is documented. Open source maintainers can apply to Claude for Open Source, which gives six months of Claude Max 20x; that is a subscription, not API credit.
Which free LLM APIs shut down or went paid in 2026? GitHub Models was retired on 30 July 2026. Cerebras replaced its free tier with a $5, 30-day trial that needs a card on 16 July. Together AI now requires a $5 purchase, SambaNova stopped giving credits to new accounts, and Groq removed Llama 3.3 70B and 3.1 8B from its free plan on 16 August.
Related guides
- OpenRouter Free Tier 2026 - rate limits and the free models, re-tested daily
- Free AI API Credits - startup and grant programs (OpenAI, Anthropic and more)
- Best Local LLM 2026 - top open models by hardware
- LM Studio vs Ollama 2026 - run LLMs locally
- Claude API Pricing 2026 - all models, caching, batch
- Claude Code Pricing 2026 - plans, API cost, and the Agent SDK credit
- Best MCP Servers 2026 - the ones worth using and how to add them
- What Is Vibe Coding? Meaning, origin, tools, and how to start
- Free GPU Compute - running your own model when free APIs are not enough
- Free Cloud Credits for Developers - infra to host your inference
- Free Startup Credits 2026: Complete Guide - every credits program in one place
Which free tier are you on right now? I update this list as providers change quotas. Reply if a quota has shifted or a new provider deserves the table.
Wiring up free tiers is the easy part. If you want to automate your business with AI, that is what I build.