Groq offers free API access to a set of LLMs with no credit card required. The rate limits are tighter than the paid tier, but for prototyping, testing, or low-volume automation, the free tier covers most of what you need. Here is how to get a key, what the limits actually are, and what to do when you hit 429.
For the paid tier rates and per-model costs once you outgrow the free quota, see Groq pricing. For free API access across providers, see free LLM API credits.
How to get a Groq API key
Go to console.groq.com. Sign up with an email address or continue with a GitHub or Google account. Once you are in the console, click "API Keys" in the left sidebar and create a new key. Copy it immediately: Groq only shows the full key once.
The key starts with gsk_. Set it as an environment variable in your project:
export GROQ_API_KEY=gsk_...
No billing setup is needed for the free tier. Your key is active as soon as you create it.
Free tier rate limits
As of October 2026, Groq's free tier limits are split by model size. The limits apply at the organization level: if you create multiple API keys under the same account, they all draw from the same pool.
Smaller models (llama-3.1-8b-instant, openai/gpt-oss-20b):
- 30 requests per minute
- 14,400 requests per day
- 6,000 tokens per minute
- 500,000 tokens per day
Larger models (llama-3.3-70b-versatile, openai/gpt-oss-120b, qwen/qwen3-32b, meta-llama/llama-4-scout-17b-16e-instruct):
- 30 requests per minute
- 1,000 requests per day
- 8,000 tokens per minute
- 200,000 tokens per day
The per-minute limit resets on a rolling window, not at a fixed clock boundary. The per-day limit resets at midnight UTC.
The 429 error
HTTP 429 Too Many Requests is the most common error on the Groq free tier. It fires in two distinct situations.
Rate limit hit (RPM). You sent more than 30 requests in the past 60 seconds. The response includes a Retry-After header with the number of seconds to wait. Handle it with exponential backoff.
Daily quota hit (RPD). You have used all 14,400 or 1,000 requests for the current day. The error message will say something like "Daily request limit reached." Waiting until midnight UTC resets the counter. No amount of retrying within the same day will work.
In client code, 429 sometimes appears as a connection failure rather than a clean HTTP response, especially when the SDK retries internally and logs it as "connection failed status 429." Check the status code directly, not just the exception message.
A quick way to test: send a single request, read the x-ratelimit-remaining-requests header in the response, and log it. You will see how many requests you have left in the current window before writing any retry logic.
Free model lineup
The free tier model catalog as of October 2026 includes:
llama-3.1-8b-instant(fast, low latency)llama-3.3-70b-versatilemeta-llama/llama-4-scout-17b-16e-instructopenai/gpt-oss-20bopenai/gpt-oss-120bqwen/qwen3-32bgroq/compoundwhisper-large-v3(speech to text)whisper-large-v3-turbo(speech to text)
Groq adjusts the free catalog without much notice. Check the Models page in your console for the current list before building a dependency on a specific model name.
The API endpoint
Groq's API is fully OpenAI-compatible. The base URL is:
https://api.groq.com/openai/v1
If you are using the OpenAI Python or JavaScript SDK, set the base_url to the Groq endpoint and pass your GROQ_API_KEY as the API key. No other changes are needed.
from openai import OpenAI
client = OpenAI(
base_url="https://api.groq.com/openai/v1",
api_key=os.environ["GROQ_API_KEY"]
)
When the free tier is not enough
Adding a credit card to your Groq account (zero minimum spend required) immediately increases your rate limits by up to 10 times across all models. You are not charged until you exceed the free tier.
If you need higher limits still, Groq's paid tiers remove the daily caps and increase the per-minute ceiling further. See Groq pricing for a breakdown of the paid tier rates and per-model costs.