Gemini API Free Tier: Models, Limits, and Paid Pricing (October 2026)

Google AI Studio is where you get a Gemini API key. The free tier requires no credit card and covers most prototyping and low-volume use cases. When you outgrow it, you upgrade to a paid Google Cloud billing account and pay per token. This article covers the current model catalog, what limits apply on the free tier, how the endpoint works, and how the paid tier is structured.

Getting a Gemini API key

Go to aistudio.google.com. Sign in with a Google account. From the top navigation, click "Get API key" and create a new key. No billing setup is required for the free tier.

The key works immediately with the Gemini API endpoint. Store it as an environment variable:

export GEMINI_API_KEY=AI...

Current free-tier models

As of October 2026, the Gemini API free tier is built around the Gemini 3.x generation. New projects are directed to these models:

Gemini 3.8 Flash is the current flagship. It is the model Google recommends for most new production-ready applications and has the broadest free-tier access.

Gemini 3.5 Flash and Gemini 3.5 Flash-Lite are the stable mid-generation options. Flash-Lite is the lightest and cheapest to run; Flash is capable enough for most development work.

Gemini 2.5 Flash remains broadly available on the free tier and continues to work for existing integrations.

Gemma 4 is Google's open-weight model available through the same API key. It runs under the same free-tier limits as the Flash lineup.

Two older models, Gemini 2.5 Pro and Gemini 2.5 Flash-Lite, have restricted access as of 2026. Google limits them to users who were already running projects on them; new projects are routed to 3.x models instead. If you are starting fresh, plan on 3.8 Flash or 2.5 Flash.

Free-tier limits

Google no longer publishes per-model rate limits in a central table. The limits are per-project and visible in the AI Studio rate-limit console at aistudio.google.com/rate-limit. They vary by usage tier and, for preview models, are more restricted than for stable releases.

Community tracking as of October 2026 (benchlm.ai) puts the 3.x Flash free-tier limits at roughly 10 requests per minute and 1,500 requests per day. These are indicative, not guaranteed, and your project's actual limits may differ. Check the console.

When you hit the per-minute limit, the API returns HTTP 429 immediately. When you hit the daily limit, further requests are blocked until the counter resets at midnight Pacific Time.

The free tier has no billing fallback. There is no automatic charge if you exceed the limit; requests just fail.

The endpoint

The Gemini API provides two endpoint styles.

The native Gemini endpoint:

https://generativelanguage.googleapis.com/v1/

An OpenAI-compatible endpoint for teams that want to reuse the OpenAI SDK without rewiring their client:

https://generativelanguage.googleapis.com/v1beta/openai/

Using the OpenAI-compatible path with the Python SDK:

from openai import OpenAI

client = OpenAI(
    base_url="https://generativelanguage.googleapis.com/v1beta/openai/",
    api_key=os.environ["GEMINI_API_KEY"]
)

The same API key works on both endpoints.

Upgrading links your Google AI Studio project to a Google Cloud billing account. Once linked, rate limits increase significantly and billing switches to per-token consumption.

Google publishes a three-tier model hierarchy for pricing:

Flash-Lite is the cheapest tier, designed for high-volume tasks where cost per call matters more than peak capability.

Flash is the mid tier, capable enough for most production tasks and significantly cheaper than Pro.

Pro is the most expensive tier, intended for complex reasoning and long-context tasks. Pro models carry a context-window premium: prompts that exceed 200,000 tokens are billed at a higher input rate than shorter prompts.

Current per-token rates for each model are at ai.google.dev/gemini-api/docs/pricing and change with new model releases. The structural hierarchy (Lite, Flash, Pro, with Pro context premium) has been stable across the 2026 model generations; specific dollar figures change.

Free vs. paid: what to use

For prototyping and scripts that run a few dozen requests a day, the free tier on 3.8 Flash or 2.5 Flash is sufficient. The daily limit is enough for active development.

For production traffic with real user volume, you will hit the free tier ceiling quickly. Upgrading to paid removes the daily cap and raises the per-minute ceiling substantially. Flash is the standard production choice; Flash-Lite for cost-sensitive high-volume pipelines; Pro only when the task genuinely needs it.

If you are migrating from another provider, note that the 3.x Flash models are not available under the 2.5 model names. Check your model identifier strings before migrating an existing integration.

If you want a broader survey of no-card inference options including OpenRouter, Groq and the GitHub Models shutdown migration paths, those are covered in the companion posts. For a different angle on Gemini specifically, see Gemini free credits for new Google Cloud accounts.

enjoyed this? follow me!

X / Twitter LinkedIn GitHub

share this!

← Back to blog