Groq Free Tier 2026: 1,000 Requests a Day, Llama Is Gone

Last updated: 11 September 2026. Limits and prices below are from Groq's own docs, checked that day.

Short answer: Groq's free tier in September 2026 gives 30 requests per minute, 1,000 requests per day, 8,000 tokens per minute and 200,000 tokens per day on each of its main chat models: OpenAI's gpt-oss-120b and gpt-oss-20b, and Qwen 3.6 27B and Qwen 3.8 27B. No credit card is needed. Llama is no longer on it: Llama 3.1 8B and Llama 3.3 70B left the free and Developer tiers on 16 August 2026 and are now enterprise-only.

Paid pricing starts at $0.075 per million input tokens (gpt-oss-20b) and goes up to $0.80 (Qwen 3.8 27B). The Batch API halves that. The June version of this page, like most pages you will find, still lists Llama 3.1 8B at $0.05 and a 14,400-requests-a-day free tier; neither exists for new accounts now.

For free tiers at every provider, see Free LLM APIs 2026.

Groq free tier rate limits (official, September 2026)

ModelReq/minReq/dayTokens/minTokens/day
openai/gpt-oss-120b301,0008,000200,000
openai/gpt-oss-20b301,0008,000200,000
openai/gpt-oss-safeguard-20b301,0008,000200,000
qwen/qwen3.6-27b301,0008,000200,000
qwen/qwen3.8-27b301,0008,000200,000
groq/compound, groq/compound-mini3025070,000-
meta-llama/llama-prompt-guard-2-22m, -86m3014,40015,000500,000
whisper-large-v3, whisper-large-v3-turbo202,000--
canopylabs/orpheus-v1-english, -arabic-saudi (TTS)101001,2003,600

Whisper is also capped at 7,200 audio seconds per hour and 28,800 per day. Source: Groq's rate limits page, checked 11 September 2026. Groq calls this "a high level summary", and your organization's exact numbers are on the Limits page in the console.

How the limits behave in practice:

  • Limits are per organization, not per API key. Ten keys in one org share one set of limits.
  • Whichever limit you hit first stops you. 200,000 tokens a day across 1,000 requests is an average of 200 tokens per request. With 2,000-token prompts you run out of tokens after about 100 requests, long before the request cap.
  • 8,000 tokens per minute is the tight one. One request with a 6,000-token prompt and a 2,000-token answer uses the whole minute.
  • Cached tokens do not count toward your limits (prompt caching works on the gpt-oss models, see below).
  • When you hit a limit you get HTTP 429 with a retry-after header in seconds. Every response carries x-ratelimit-remaining-requests (per day) and x-ratelimit-remaining-tokens (per minute), so you can back off before you hit it.

The daily cap on gpt-oss-120b is worth between 3 and 12 cents a day at paid prices, depending on the input/output mix. The free tier is for building and testing, not for traffic.

What did Groq remove in 2026?

ShutdownModelGroq's recommended replacement
5 MarchLlama Guard 4 12Bgpt-oss-safeguard-20b
9 MarchLlama 4 Maverick 17Bgpt-oss-120b
15 AprilKimi K2 (moonshotai/kimi-k2-instruct-0905)gpt-oss-120b
17 JulyQwen 3 32Bgpt-oss-120b
17 JulyLlama 4 Scout 17Bgpt-oss-120b or Qwen 3.6 27B
16 AugustLlama 3.1 8B Instantgpt-oss-20b
16 AugustLlama 3.3 70B Versatilegpt-oss-120b or Qwen 3.6 27B

The two August shutdowns apply to free and Developer-tier usage only; enterprise customers with a committed-spend contract keep Llama 3.1 8B and 3.3 70B. Source: Groq's deprecations page. If you still need Llama 3.3 70B for free, Cloudflare Workers AI serves it within its 10,000 Neurons a day, see Free LLM APIs.

Groq API pricing per model

ModelInput / 1M tokensOutput / 1M tokensGroq's speed figureContext
openai/gpt-oss-20b$0.075$0.30~1,000 tok/s131K
openai/gpt-oss-safeguard-20b$0.075$0.30~1,000 tok/s131K
openai/gpt-oss-120b$0.15$0.60~500 tok/s131K
qwen/qwen3.6-27b (preview)$0.60$3.00~500 tok/s131K
qwen/qwen3.8-27b (preview)$0.80$4.00~450 tok/s131K
Llama Prompt Guard 2 22M / 86M$0.03 / $0.04same-512
Whisper Large V3 Turbo$0.04 per hour of audio---
Whisper Large V3$0.111 per hour of audio---
Orpheus V1 English (TTS)$22 per 1M characters---
Orpheus Arabic Saudi (TTS)$40 per 1M characters---
Llama 3.1 8B, Llama 3.3 70B, MiniMax M2.7Enterprise onlycontact sales--

Source: Groq's models page, 11 September 2026. The old groq.com/pricing page now redirects to the home page; the models page is where prices live. Qwen models are marked preview, which Groq says means "evaluation purposes only" and possible discontinuation at short notice. groq/compound has no token price of its own: it is billed by the models and tools it calls.

Two things stand out. gpt-oss-20b is half the price of gpt-oss-120b and twice as fast by Groq's figures, so start there. And Qwen on Groq is expensive: $0.80 / $4.00 for Qwen 3.8 27B, where other providers sell the same model at $0.15 to $0.45 input on OpenRouter.

How fast is Groq, measured?

Groq's own speed figures are best-case. OpenRouter publishes what each provider actually delivers on real traffic, so here is gpt-oss-120b across the providers that serve it, from OpenRouter's endpoint stats for the 30 minutes before 16:18 UTC on 11 September 2026:

ProviderThroughput, medianTime to first token, medianPrice in / out per 1M
Cerebras669 tok/s274 ms$0.35 / $0.75
SambaNova353 tok/s729 ms$0.14 / $0.95
Amazon Bedrock316 tok/s404 ms$0.15 / $0.60
Groq256 tok/s234 ms$0.15 / $0.60
Nebius218 tok/s306 ms$0.15 / $0.60
BaseTen148 tok/s281 ms$0.10 / $0.50
Google146 tok/s386 ms$0.09 / $0.36
Together79 tok/s260 ms$0.15 / $0.60
CoreWeave51 tok/s359 ms$0.03 / $0.17

What that says:

  • Groq had the fastest first token of every provider listed, which is what makes a chat UI feel instant. Its uptime over the window was 99.99%.
  • Throughput was about half of Groq's own "~500 tok/s" figure and well behind Cerebras. Groq is fast, but no longer the fastest at raw generation for this model.
  • Groq does not undercut on price. $0.15 / $0.60 is the same as Bedrock, Nebius and Together; GPU hosts sell gpt-oss-120b for about a quarter of that if you can accept 50 tok/s.

A 30-minute window moves with load, so treat these as a snapshot, not a ranking. For gpt-oss-20b in the same window, Groq's median throughput was 329 tok/s with the 90th percentile at 717.

Free, Developer or Enterprise: which plan do I need?

FreeDeveloperEnterprise
CardNoCard, US bank account or SEPA debitContract
Billing-Pay-as-you-go, monthly in arrearsCommitted spend
gpt-oss-120b limits30 RPM, 1K RPD, 8K TPM1,000 RPM, 250K TPMCustom
Flex tier (10x limits, same price)NoYesYes
Batch API (50% off)NoYesYes
Spend limits and budget alertsNoYesYes
Llama 3.1 8B / 3.3 70BNoNoYes
Performance tier with 99.9% SLANoNoYes

Upgrading to Developer costs nothing up front. Groq bills new accounts progressively: an invoice is triggered when lifetime usage first crosses $1, $10, $100, $500 and $1,000, then monthly. Invoices under $0.50 are not charged. You can downgrade back to Free at any time after paying the final invoice. There is no general discount on Developer: the June version of this page claimed 25% off on-demand prices, and nothing in Groq's billing docs supports that.

Batch, Flex and prompt caching

  • Batch API: 50% off. Upload a JSONL file of requests, choose a processing window from 24 hours to 7 days, collect results. Batch jobs do not use your normal rate limits. Available on the gpt-oss models and Whisper. Developer tier and above.
  • Flex: same price, 10x the limits. Pass service_tier: "flex" and your requests run with ten times the on-demand limits while capacity allows. When it does not, you get a fast HTTP 498 capacity_exceeded and should retry with jittered backoff. Paid tiers only. service_tier: "auto" picks the best tier you have.
  • Prompt caching: 50% off cached input tokens. Automatic, no code change, on gpt-oss-120b, gpt-oss-20b and gpt-oss-safeguard-20b only. A cache hit needs an exact prefix match, so put the system prompt, tools and examples first and the variable part last. Cached prefixes expire after 2 hours without use.

The two discounts do not stack. Groq's batch docs say it plainly: "All batch tokens are billed at the 50% batch rate regardless of cache status." The "75% off with Batch plus caching" figure repeated around the web, including by this page in June, is wrong. The best you can do is half price.

How does Groq compare with other APIs on price?

Output price is what dominates most bills:

ModelProviderInput / 1MOutput / 1M
gpt-oss-20bGroq$0.075$0.30
gpt-oss-120bGroq$0.15$0.60
deepseek-flash (V4.1 Flash), off-peakDeepSeek$0.15$0.60
gpt-5.6-lunaOpenAI$0.20$1.20
Claude Haiku 4.5Anthropic$1.00$5.00
Claude Sonnet 5Anthropic$2.00$10.00
gpt-5.6-terraOpenAI$2.00$12.00

Prices from each provider's pricing page, 11 September 2026. Groq runs open-weight models only: no GPT-5.x, Claude or Gemini. The case for it is speed to first token at a mainstream open-model price, not the lowest price and not frontier quality. If a task needs a frontier model, pay the frontier provider; if gpt-oss-120b is good enough, Groq is one of the quickest places to run it. Claude API Pricing covers the Anthropic side in detail.

How do I start with Groq?

  1. Sign up at console.groq.com. No card. You land on the Free plan.
  2. Create an API key in the console.
  3. Use the OpenAI-compatible endpoint https://api.groq.com/openai/v1 with any OpenAI SDK:
from openai import OpenAI

client = OpenAI(base_url="https://api.groq.com/openai/v1", api_key="YOUR_GROQ_KEY")
reply = client.chat.completions.create(
    model="openai/gpt-oss-20b",
    messages=[{"role": "user", "content": "One sentence on what an LPU is."}],
)
print(reply.choices[0].message.content)
  1. Start with gpt-oss-20b. Move to gpt-oss-120b when quality is the bottleneck.
  2. Check your limits on the console's Limits page, or read the x-ratelimit-* headers on every response.
  3. Upgrade to Developer when you need more than 1,000 requests a day or want Batch.

What are the common mistakes with Groq?

  • Still calling llama-3.1-8b-instant or llama-3.3-70b-versatile. On free and Developer accounts these stopped working on 16 August 2026. Switch to openai/gpt-oss-20b and openai/gpt-oss-120b.
  • Planning on the request cap and forgetting the token cap. 200,000 tokens a day and 8,000 a minute run out first on any real prompt.
  • Expecting Batch and caching to stack. They do not; batch is always billed at 50%.
  • Shipping on a preview model. The Qwen models are preview and can be withdrawn at short notice, as Qwen 3 32B was in July.
  • Single-provider dependency. Groq has shut down seven models since March 2026. Put a gateway in front (OpenRouter or LiteLLM) so a model shutdown is a config change, not an outage.

Frequently asked questions

What are the Groq API free tier rate limits in 2026?

On gpt-oss-120b, gpt-oss-20b, Qwen 3.6 27B and Qwen 3.8 27B: 30 requests per minute, 1,000 requests per day, 8,000 tokens per minute and 200,000 tokens per day, per organization, with no credit card. Whisper gets 20 requests a minute and 2,000 a day. These are Groq's published figures as of September 2026; your exact limits are on the Limits page of the console.

Is Groq free to use?

Yes, the Free plan needs no card and gives every account the limits above. It is meant for building and testing: the daily token cap on gpt-oss-120b is worth only a few cents at paid prices. For more, the Developer plan is pay-as-you-go with no upfront charge.

Which models are free on Groq?

The chat models on the free plan are openai/gpt-oss-120b, openai/gpt-oss-20b, openai/gpt-oss-safeguard-20b, qwen/qwen3.6-27b and qwen/qwen3.8-27b, plus the groq/compound agent systems, Whisper speech-to-text, Orpheus text-to-speech and Llama Prompt Guard. Llama 3.1 8B and Llama 3.3 70B are no longer available on free or Developer accounts.

How much does Groq cost per million tokens?

gpt-oss-20b costs $0.075 input and $0.30 output per million tokens, gpt-oss-120b $0.15 and $0.60, Qwen 3.6 27B $0.60 and $3.00, and Qwen 3.8 27B $0.80 and $4.00. Whisper Large V3 Turbo is $0.04 per hour of audio. The Batch API takes 50% off.

Is Llama still available on Groq?

Not on free or Developer accounts. Groq shut down Llama 3.1 8B Instant and Llama 3.3 70B Versatile for those tiers on 16 August 2026 and recommends gpt-oss-20b and gpt-oss-120b instead. Enterprise customers with a committed-spend contract still have them.

Do Groq's batch and prompt caching discounts stack?

No. Batch is 50% off and prompt caching is 50% off cached input tokens, but batch tokens are always billed at the batch rate regardless of cache hits. The lowest you can pay is half of the on-demand price.

How fast is Groq compared with other providers?

On OpenRouter's measurements for gpt-oss-120b on 11 September 2026, Groq had the lowest median time to first token (234 ms) and a median throughput of 256 tokens per second. Cerebras was faster at generation (669 tokens per second) at a higher price, and SambaNova and Amazon Bedrock were also ahead on throughput.

enjoyed this? follow me!

X / Twitter LinkedIn GitHub

share this!

← Back to blog