Free LLM API 2026: 17 Tested, 7 Need No Card

Last updated: 11 September 2026. Every limit on this page was checked against the provider's own docs that day.

Seven LLM APIs are free in September 2026 with no credit card: Google Gemini, Groq, OpenRouter, Cloudflare Workers AI, Mistral, Cohere and Hugging Face. Two more, NVIDIA's API catalog and Z.ai, give free access for prototyping without asking for payment on their pages. Everyone else is either a one-time trial or wants money before the first request.

The list moved a lot since June. Cerebras turned its free tier into a $5 trial that needs a card, GitHub Models shut down, Groq dropped Llama from its free plan, Together stopped giving trial credit, and every model OpenRouter offered free in June is paid now. The timeline further down has the dates. Here is the map as it stands.

Free LLM APIs with no credit card

ProviderWhat you get freeMain limitThe catch
Google Gemini APIGemini 3.x Flash, 2.5 Flash, Gemma 4Shown per project in AI StudioFree-tier prompts may be used for training
Groqgpt-oss-120b, gpt-oss-20b, Qwen 27B30 req/min, 1,000 req/day, 200K tokens/daySmall 8K tokens/min cap
OpenRouter19 :free models20 req/min, 50 req/dayUpstream providers return 429s
Cloudflare Workers AILlama 3.3 70B, gpt-oss-120b, Qwen3, Kimi K2.5 and more10,000 Neurons/dayBudget is in Neurons, not tokens
MistralMistral API models$10/month of creditsTraining on by default in Free mode
CohereEvery Cohere model, Command A included1,000 calls/month, 20 req/minNo production or commercial use
Hugging FaceModels from 18 inference partners$0.10/month of credits40K to 700K output tokens, by model

"No card" here means the provider's own pages say you can call the API without adding a payment method. Limits reset daily unless the table says monthly.

All 17 providers compared

ProviderFree offer todayCardStatus vs June
Google Gemini APIFree tier, limits in AI StudioNoGemini 2.0 retired, 3.x Flash added
Groq30 RPM, 1K RPD on gpt-oss and QwenNoLlama removed from free plan
OpenRouter19 free models, 50 or 1,000 req/dayNoWhole free roster replaced
Cloudflare Workers AI10,000 Neurons/dayNoNew to this list
Mistral$10/month Free-plan creditsNoWas listed as card-required trial
CohereTrial key, 1,000 calls/monthNoWas listed as card-required trial
Hugging Face$0.10/month creditsNoSmaller than "free tier" suggests
NVIDIA API catalogUp to 40 RPM on free endpointsNot asked on its pagesPrototyping only
Z.aiGLM-4.7-Flash, GLM-4.5-Flash freeNot statedNew to this list
Alibaba Model Studio~1M tokens per model, 90 daysNot statedNew to this list
Fireworks AI$1 one-time creditNo, trial onlyNew to this list
Cerebras$5 for 30 daysYesWas free with no card
OpenAINone; data-sharing tokens if eligibleYes, $5 prepayNo signup credit
Anthropic ClaudeNone; OSS program gives Max 20xYesNo documented signup credit
DeepSeekNone, very cheap per tokenTop-upNow DeepSeek V4.1 Flash
xAI GrokNone in official docsPrepaidGrok 4.6
Together AINone, $5 minimum purchaseYesTrial credit ended
Heads up, these free tiers change almost monthly. Since June alone, Cerebras, GitHub Models, Together and SambaNova dropped their free offers and Groq cut its model list. I keep this list updated. Subscribe with your email and I'll send new free APIs as they launch.

Google Gemini API

  • Free models: Gemini 3.8, 3.7, 3.6 and 3.5 Flash, 3.5 and 3.1 Flash-Lite, Gemini 2.5 Flash, Gemma 4 (31B and 26B A4B), embeddings, plus live, TTS and transcription models. Gemma 4 has no paid tier at all. The pricing page also lists 2.5 Pro and 2.5 Flash-Lite as free, but on 11 September both told our key they are "no longer available to new users".
  • Not free: Gemini 3.1 Pro preview, the Nano Banana image models, Veo and Lyria.
  • Rate limits: Google no longer publishes free-tier numbers in its docs. The rate limits page now sends you to AI Studio, where each project sees its own RPM, TPM and daily caps. The daily quota resets at midnight Pacific. The "1,500 requests a day at 10 RPM" figures quoted around the web, including by this page in June, are no longer official.
  • Grounding with Google Search: free up to 500 requests a day, shared between 2.5 Flash and Flash-Lite.
  • Card required: No. A Google account is enough; billing is only needed to move to a paid tier.
  • The $300 Google Cloud trial no longer pays for the Gemini API in AI Studio on accounts opened after 2 March 2026, and the paid tier now starts with a $5 prepayment. Details in Gemini Free Credits.
  • Your data: on the free tier Google may use prompts and responses to improve its products, and human reviewers may read them. The paid tier does not.
  • Regions: the free tier works in the EEA, UK and Switzerland, but the terms allow only paid services if you serve end users there.
  • Start: aistudio.google.com -> Get API key.

Still the best general-purpose free baseline: current Flash models, a large context window, and no card. Details on credits and the Google Cloud side are in Gemini Free Credits.

Groq

Groq's free plan now runs on OpenAI's open-weight gpt-oss models and Qwen. Llama 3.3 70B and Llama 3.1 8B left the free and developer tiers on 16 August 2026 and are enterprise-only now.

ModelReq/minReq/dayTokens/minTokens/day
openai/gpt-oss-120b301,0008,000200,000
openai/gpt-oss-20b301,0008,000200,000
qwen/qwen3.6-27b, qwen/qwen3.8-27b301,0008,000200,000
groq/compound, groq/compound-mini3025070,000-
whisper-large-v3, -turbo (speech to text)202,000--

Source: Groq's rate limits page, checked 11 September 2026.

  • Card required: No. A payment method is only asked for when you upgrade to the Developer tier.
  • Your data: not retained by default; logs may be kept up to 30 days for abuse cases, with an opt-out.
  • The catch: 8,000 tokens a minute means one long prompt can use the whole minute. Groq is for fast, short, interactive calls, not bulk jobs.
  • Start: console.groq.com

More on paid tiers in Groq Pricing.

OpenRouter

  • Free models: 19 text models with the :free suffix (NVIDIA Nemotron 3, Google Gemma 4, Poolside Laguna, inclusionAI Ling and others), plus the openrouter/free router that picks one for you.
  • Limits: 20 requests a minute; 50 a day, or 1,000 a day once you have bought $10 of credits at any point.
  • Card required: No.
  • The catch: each free model is served by an upstream provider with its own capacity. When it is busy you get a 429 even if you made one request all day. In our test of all 19 on 11 September, 4 were refused this way.

Not one of the models OpenRouter offered free in June (DeepSeek R1, Llama 3.3 70B, Qwen3 Coder...) is free there today. We re-test the free list every day in OpenRouter Free Tier, with today's result for each model.

Cloudflare Workers AI

  • Free allowance: 10,000 Neurons a day on both the free and paid Workers plans, reset at 00:00 UTC. A Neuron is Cloudflare's compute unit; each model has a Neuron price per million tokens.
  • What that buys, counting output tokens only: about 49,000 tokens a day on Llama 3.3 70B, about 147,000 on gpt-oss-120b, about 287,000 on Llama 3.1 8B. Our arithmetic from Cloudflare's pricing page.
  • Free models: Llama 3.3 70B, Llama 3.1 8B, Llama 4 Scout, gpt-oss-120b and 20b, Qwen3 30B, QwQ 32B, Qwen2.5 Coder 32B, Gemma 3 and Gemma 4, Mistral Small 3.1, Nemotron 3 120B, Kimi K2.5, GLM-4.7-Flash and more. Kimi K2.6 and K2.7 Code, GLM-5.x and DeepSeek V4 need a paid plan.
  • Rate limits: 300 requests a minute by default for text generation; 20 a minute on the larger frontier models.
  • Card required: No.
  • Your data: Cloudflare does not use your content to train models.
  • Start: a free Cloudflare account, then the REST API or a Worker binding. There is an OpenAI-compatible endpoint.

This is where Llama 3.3 70B is still free after Groq and OpenRouter dropped it.

Mistral

  • Free offer: new accounts start in Free mode, "no credit card required", and the Free plan includes $10 a month of credits shared between the API, Studio and Mistral's apps.
  • Rate limits: requests per second, tokens per minute and tokens per month apply, and Free mode has the lowest. Mistral no longer publishes the numbers; you see them in the admin panel.
  • Models: the API list includes Mistral Large, Medium, Small and Codestral; which of them Free mode covers is shown in your console.
  • Your data: in Free mode Mistral may use your inputs and outputs for training unless you opt out.
  • Start: console.mistral.ai

Our June version said Mistral needed a card. Its own docs say otherwise today.

Cohere

EndpointTrial key limit
All calls1,000 per month
Chat, any model20 requests/min
Embed2,000 inputs/min
Rerank10 requests/min
  • Models: a trial key reaches every Cohere model: Command A+, Command A, Command A Reasoning, Vision and Translate, Command R+ and R, and North Mini Code.
  • Card required: No. The trial key is created at signup.
  • The catch: trial keys may not be used for production or commercial purposes.
  • Best for: trying Rerank and Embed for RAG before paying.

Hugging Face Inference Providers

  • Free offer: $0.10 of credit a month on a free account, spent at the partner providers' own prices (Cerebras, Groq, Together, Fireworks, Novita, Scaleway, Z.ai and others). PRO costs $9 a month and includes $2.
  • Card required: No for the monthly $0.10; yes to buy more.
  • The catch: at partner prices, ten cents is about 590,000 output tokens on gpt-oss-120b at the cheapest provider, or 38,000 on DeepSeek V4 Pro. Enough to check that a model works in your code, not to run anything.

Our June version called this a "rate-limited free tier". The real limit is the ten cents. Full breakdown in Hugging Face Inference API.

NVIDIA API catalog (build.nvidia.com)

  • Free offer: models tagged "Free Endpoint", at up to 40 requests a minute. NVIDIA says limits vary by model and your exact numbers are shown in your account.
  • Models: a large catalog that includes DeepSeek V4 Pro and V4 Flash, Kimi K3, MiniMax M3, Nemotron 3 Ultra and Nemotron 3.5, among about 80 hosted models.
  • Conditions: you join the free NVIDIA Developer Program, and the endpoints are for "prototyping, research, development and testing purposes only". Production use needs an NVIDIA AI Enterprise license.
  • Card required: not mentioned on any of NVIDIA's pages; you create and verify an account.
  • Start: build.nvidia.com

The easiest free way to try frontier open models such as DeepSeek V4 Pro before you pick a paid host.

Z.ai

Z.ai's pricing page lists GLM-4.7-Flash, GLM-4.5-Flash and GLM-4.6V-Flash (vision) as free for input and output. It does not state the rate limits or whether a card is needed, so treat it as a model you can try rather than a quota you can plan on.

One-time trial credits

These give you something free once, then stop.

  • Alibaba Cloud Model Studio: about 1,000,000 free tokens per model for new users, each model with its own quota, valid 90 days (for activations from 8 September 2026). Singapore region with International deployment only, real-time inference only. The quota page implies you can use it before completing billing details. Source.
  • Fireworks AI: $1 of credit with no card, at 10 requests a minute until you add a payment method. When the dollar is spent the account is suspended until you add one.
  • Cerebras: since 16 July 2026 there is no permanent free tier. You get a $5 credit, valid 30 days, only after adding a verified payment method. On the trial, gpt-oss-120b and qwen-3.8-27b run at 5 requests a minute and 1 million tokens a day, with context capped at about 64K. Still very fast; no longer free in the sense this page means.

Also checked and left out of the tables: Vercel AI Gateway (a monthly free credit, but it needs a payment method), Scaleway (1 million free tokens, card required), SambaNova (its docs still describe a free tier, but new accounts are told to buy credits), Nebius ($1 for 30 days with a card), Chutes and Kimi (no free tier).

Pay first: OpenAI, Anthropic, DeepSeek, xAI, Together

OpenAI. No free credit for new API accounts; you prepay at least $5. What remains is the data-sharing program: eligible organizations that agree to share API traffic get free tokens every day, 250,000 on flagship models and 2.5 million on small ones at usage tiers 1-2, 1 million and 10 million at tiers 3-5. It covers the GPT-5.x line (including gpt-5.6-sol, terra and luna) but not the new flagship gpt-6-astra, and it needs a positive balance. Startup programs are in OpenAI Free Credits.

Anthropic Claude. No documented starter credit: Anthropic's help center says you buy credits before using the API. Current models are Claude Fable 5.1 ($10 / $50 per million input / output tokens), Opus 5 ($5 / $25), Sonnet 5 ($2 / $10) and Haiku 4.5 ($1 / $5). The free route is Claude for Open Source: six months of Claude Max 20x for qualifying maintainers, up to 10,000 recipients, still open. It is a subscription, not API credit, and eligibility widened this year (merged PRs in other projects, external contributors, OpenSSF criticality score). See Claude Free Credits and Claude API Pricing.

DeepSeek. No free grant, but the cheapest serious model here: deepseek-flash (V4.1 Flash) costs $0.15 per million input tokens and $0.60 output off-peak, twice that at peak hours; V4-Pro is $0.66 / $1.98 off-peak. Both have 1M context. You top up by card, PayPal, Alipay or WeChat Pay.

xAI Grok. xAI's current docs mention only prepaid credits and promo codes. Third-party sites report a $25 signup credit and $150 a month for sharing data; we could not find either on xAI's own pages. The flagship is grok-4.6 at $2 / $6.

Together AI. "Together AI does not currently offer free trials." A $5 minimum purchase unlocks the API; purchased credits do not expire.

What changed since June 2026

DateChange
1 JuneGoogle shuts down Gemini 2.0 Flash and Flash-Lite; 3.6 Flash and 3.1 Flash-Lite replace them
16 JuneGitHub Models closes to new customers
17 JuneGroq announces the end of Llama 3.3 70B and Llama 3.1 8B
16 JulyCerebras replaces its free tier with a $5, 30-day trial that needs a card
30 JulyGitHub Models is retired; API, playground and BYOK are gone
16 AugustLlama 3.3 70B and 3.1 8B leave Groq's free and developer tiers
17 AugustCerebras deprecates GLM-4.7; Qwen 3.8 27B takes its trial slot
AugustSambaNova stops giving credits to new accounts
8 SeptemberAlibaba's new-user quota becomes valid for 90 days
By SeptemberNone of OpenRouter's June free models are free any more; the roster is 19 new ones

Who went which way, compared with this page's June version:

ProviderJuneSeptemberDirectionFull review
CerebrasFree, no card, 1M tokens/day$5 trial, card requiredWent paid-
GitHub ModelsFree with a GitHub accountRetiredGone-
Together AISmall starter credit$5 minimum purchaseWent paid-
GroqFree Llama 3.xFree gpt-oss and Qwen onlyFewer modelsGroq Pricing
OpenRouter~28 free models19 free models, all differentReplacedOpenRouter Free Tier
Anthropic~$5 starter creditNone documentedWent paidClaude Free Credits
OpenAIInconsistent ~$5 creditNone; $5 prepayWent paidOpenAI Free Credits
Hugging Face"Rate-limited free tier"$0.10/monthClarified: tinyHugging Face Inference API
MistralTrial credits, card$10/month, no cardBetter-
CohereTrial credits, card1,000 calls/month, no cardBetter-
Cloudflare Workers AINot listed10,000 Neurons/day, no cardAdded-
Google Gemini2.5 Flash free3.x Flash and 2.5 Flash free, limits unpublishedNew modelsGemini Free Credits

Some rows reflect a real change on a known date; others (Mistral, Cohere, Hugging Face) mean our June version was wrong or the provider changed quietly. We re-check this page monthly.

How to stack free LLM APIs

Each provider counts its limits separately, so a small product can run on several free tiers at once:

  1. Default to Gemini Flash for general chat, extraction and instruction following.
  2. Groq for latency-sensitive calls, such as voice or streaming chat, with short prompts that fit its 8K tokens a minute.
  3. Cloudflare Workers AI when you want Llama 3.3 70B or gpt-oss inside a Worker, with no card.
  4. OpenRouter as the fallback: when a primary provider rate-limits you, send the same request to openrouter/free.
  5. Mistral and Cohere for evaluation: their free allowances are monthly, so save them for tests, not traffic.
  6. Pay for the hard 10 percent. Keep a small balance on OpenAI, Anthropic or DeepSeek for the calls where quality matters.

For the routing layer, see LLM Gateway 2026, which compares OpenRouter, LiteLLM, Portkey and Helicone. For ranking models on output quality, see Best Free LLM 2026. To stop depending on free tiers altogether, Best Local LLM covers running open models on your own hardware.

Frequently asked questions

Is there a truly free LLM API with no credit card? Yes. In September 2026, Google Gemini, Groq, OpenRouter, Cloudflare Workers AI, Mistral, Cohere and Hugging Face all let you call their API without adding a card. Gemini, Groq and Cloudflare give the most daily usage; Mistral and Cohere give a monthly allowance; Hugging Face gives $0.10 a month.

Is there a free LLM API with no limits? No. Every free tier has a cap: requests per minute, requests or tokens per day, or a monthly credit. The way to get more is to use several free tiers side by side, since each one counts its limits separately, or to run an open model locally.

Which free LLM API is the most generous in 2026? For daily volume without a card: Groq (1,000 requests and 200,000 tokens a day per model on gpt-oss-120b), Cloudflare Workers AI (10,000 Neurons a day, about 147,000 output tokens on gpt-oss-120b), and Gemini, whose limits Google now shows only inside AI Studio. Cerebras used to lead with 1 million tokens a day but became a card-required trial in July 2026.

Can I combine multiple free LLM APIs to get more inference? Yes. Each provider has independent rate limits, so routing across Gemini, Groq, Cloudflare and OpenRouter multiplies your free capacity. A gateway such as OpenRouter or LiteLLM, or a small model registry of your own, handles the fallback when one of them returns a 429.

Does OpenAI give free API credits? Not to new accounts: you prepay at least $5. Eligible organizations can get free daily tokens by sharing their API traffic with OpenAI, up to 1 million tokens a day on flagship GPT-5.x models and 10 million on small ones at higher usage tiers.

Is the Claude API free? No. Anthropic's help center says you buy credits before using the API, and no signup credit is documented. Open source maintainers can apply to Claude for Open Source, which gives six months of Claude Max 20x; that is a subscription, not API credit.

Which free LLM APIs shut down or went paid in 2026? GitHub Models was retired on 30 July 2026. Cerebras replaced its free tier with a $5, 30-day trial that needs a card on 16 July. Together AI now requires a $5 purchase, SambaNova stopped giving credits to new accounts, and Groq removed Llama 3.3 70B and 3.1 8B from its free plan on 16 August.


Which free tier are you on right now? I update this list as providers change quotas. Reply if a quota has shifted or a new provider deserves the table.


Wiring up free tiers is the easy part. If you want to automate your business with AI, that is what I build.

enjoyed this? follow me!

X / Twitter LinkedIn GitHub

share this!

← Back to blog