Cheapest LLM API in 2026: 9 Models Ranked by Price per Token

Last verified: September 25, 2026. LLM prices change often; confirm on each provider's pricing page before relying on a figure below.

The cheapest widely-used LLM API in September 2026 is Gemini 2.5 Flash-Lite at $0.10 per million input tokens and $0.40 per million output tokens. DeepSeek V4 Flash is close on input at $0.14 (output rose to $0.60 off-peak in Q3 2026), and OpenAI's cheapest current-generation model is GPT-5.4 nano at $0.20 / $1.25. Older-family OpenAI models such as gpt-5-nano ($0.05 / $0.40) and gpt-4.1-nano ($0.10 / $0.40) run cheaper still if capability trade-offs are acceptable.

This guide ranks nine widely-used paid LLM APIs by verified per-token price as of September 25, 2026, separates the genuinely cheap tier from mid-priced options, and shows three levers that cut a real bill in half. If your volume is small, the cheapest option may be free entirely, which is covered at the end.

Price per million tokens (verified September 25, 2026)

Prices are approximate and change frequently. Input and output are billed separately; output usually costs several times more. Always confirm on the provider's own pricing page.

ModelInput $/1MOutput $/1MNotesPricing
OpenAI GPT-5.4 nano~$0.20~$1.25Cheapest from OpenAI, capable for simple tasksopenai.com/api/pricing
Gemini 2.5 Flash-Lite~$0.10~$0.40Very cheap, fast, multimodalai.google.dev/pricing
DeepSeek V4 Flash~$0.14~$0.60 off-peak, ~$1.20 peakOutput doubled in Q3 2026; peak = 01:00-10:00 UTC Mon-Fri. Cache-hit input near $0.003api-docs.deepseek.com
GPT-5.4 mini~$0.75~$4.50Step up in quality, still cheapopenai.com/api/pricing
Gemini 2.5 Flash~$0.30~$2.50Workhorse balance of cost and qualityai.google.dev/pricing
Mistral Small 4~$0.15~$0.60EU-hosted; both input and output doubled from earlier Small tiermistral.ai/pricing
Llama (via Groq)Enterprise pricingEnterprise pricingPublic per-token tier removed in Q3 2026. GPT OSS 20B on Groq: $0.075 / $0.30groq.com/pricing
Qwen 3.8 Flash (Together.ai)~$0.09~$0.28Cheapest verified Qwen option in Q3 2026together.ai/pricing
Claude Haiku 4.5~$1.00~$5.00Pricier, but strong quality per tokenanthropic.com/pricing

The cheapest tier: GPT-5.4 nano, Gemini Flash-Lite, DeepSeek

At the bottom of the price range, three models stand out. Gemini 2.5 Flash-Lite is the outright cheapest at $0.10 / $0.40, fast, and multimodal. GPT-5.4 nano is OpenAI's cheapest current-generation model at $0.20 / $1.25, good for classification, routing, and simple extraction. DeepSeek V4 Flash sits close on input at $0.14 and punches far above its price on reasoning and code, though its output price rose to $0.60 off-peak in Q3 2026. DeepSeek's deep prompt-cache discount (cached input drops to a few thousandths of a cent per million) still makes repeated context nearly free. For most high-volume, low-difficulty tasks, any of these three is the right default.

The mid tier: when cheap is not capable enough

When the cheapest models miss, the next tier up is still inexpensive. GPT-5.4 mini and Gemini 2.5 Flash are the workhorses: a few times the price of the cheapest tier, but markedly better at multi-step instructions and longer context. Mistral Small 4 fills the same niche with an EU-hosting advantage, though both its input and output doubled from earlier Small versions in Q3 2026. Llama on Groq moved to enterprise-only pricing in Q3 2026; if you need Llama on a public per-token tier, check Together.ai or Fireworks. The jump from "cheapest" to "workhorse" is usually a 3 to 8 times price increase, which is still trivial next to frontier-model pricing.

How to actually cut your LLM bill

Choosing a cheap model is only the first lever. Three more cut real costs:

  1. Prompt caching. If you send the same long system prompt on every request, caching lets you pay for it once instead of every call. On a chat app with a big system prompt, this alone can cut input cost dramatically.
  2. Batch APIs. For non-urgent jobs (overnight processing, bulk extraction), most providers offer a batch endpoint at roughly half price in exchange for slower turnaround.
  3. Model routing. Send each request to the cheapest model that can handle it, and escalate to an expensive model only when needed. A router like OpenRouter or LiteLLM makes this one line of config. The full pattern is in the guide on LLM gateways and routers.

Stacked together, caching plus batch plus routing routinely cut a bill by more than half without changing what the app does.

Cheapest of all: free first

For small volume, the cheapest LLM API is no API bill at all. Gemini, Groq, Cerebras, and several OpenRouter models have free tiers that cost nothing within their rate limits, which covers a lot of real applications. Start there, and move to a cheap paid model like GPT-5.4 nano or DeepSeek only when you outgrow the free rate limits or need guaranteed throughput. The complete list of free options is in the guide on free LLM APIs in 2026.

Frequently asked questions

What is the cheapest LLM API in 2026? As of September 2026, Gemini 2.5 Flash-Lite is the cheapest widely-used paid LLM API at $0.10 per million input tokens and $0.40 per million output tokens. DeepSeek V4 Flash follows at $0.14 input / $0.60 output (off-peak), and OpenAI's GPT-5.4 nano is $0.20 input / $1.25 output. Older-family models like gpt-5-nano ($0.05 / $0.40) run cheaper still if capability trade-offs are acceptable.

Is it cheaper to use a free LLM API instead? For low volume, yes. Free tiers from Gemini, Groq, and OpenRouter cost nothing within their rate limits, so a small app may never need to pay. Paid APIs become worthwhile when you exceed free rate limits or need guaranteed throughput. Many builders start free and move to a cheap paid model only when they scale.

How do input and output token prices differ? Output tokens almost always cost more than input, often 3 to 5 times more. A model priced at $0.30 per million input tokens may charge $2.50 per million output tokens. For cost estimates, weight your expected output volume heavily, because generation is where the bill grows.

How can I cut my LLM API costs further? Three levers: use prompt caching to avoid re-paying for a repeated system prompt, use batch APIs for non-urgent jobs (often half price), and route each request to the cheapest model that can handle it instead of sending everything to a frontier model. Together these can cut a bill by more than half.

Is the cheapest LLM API good enough for production? Often yes. Models like Gemini Flash, GPT-5.4 mini, and DeepSeek handle classification, extraction, summarization, and most chat at a fraction of frontier-model cost. Reserve expensive models for hard reasoning and code. Routing cheap models for easy tasks and expensive ones only when needed is the standard production pattern.


Which model do you ship on? I keep this table current as providers change pricing. Reply if a price has shifted or a model belongs on the list.

enjoyed this? follow me!

X / Twitter LinkedIn GitHub

share this!

← Back to blog