✍️Writing & Content23🎨Image Generation32🎬Video & Animation64🎵Audio & Music49💬Chatbots & Assistants36💻Coding & Development155📈Marketing & SEO52Productivity139🎯Design & UI/UX52📊Data & Analytics36📚Education & Research24💼Business & Finance50🏥Healthcare & Wellness18🔍Search & Knowledge14🤖AI Agent Infrastructure25🛡️AI Security & Testing2🧊3D & Spatial14🔎SEO Tools5🏡Real Estate4🗃️Data Extraction3🧠ADHD & Focus Tools9
← Back to AI Tools Directory

LLM API Pricing Comparison 2026: 41 Models, Ranked by Workload

Most comparisons sort providers by headline rate and declare a winner. We ran all 41 verified models against five real workload shapes instead — and got 1 different winners. Which API is cheapest is a question about your traffic, not about the rate card.

Prices verified July 28, 2026 against each provider's own pricing page. Sources linked in the table below.

The short answer

  • Support chatbot: GPT-5 nano (OpenAI) at $9.88/mo — next cheapest is DeepSeek V4 Flash at $12.11
  • RAG question answering: GPT-5 nano (OpenAI) at $11.38/mo — next cheapest is Gemini 2.5 Flash-Lite at $18.76
  • Document summarization: GPT-5 nano (OpenAI) at $5.35/mo — next cheapest is GPT-4.1 nano at $9.10
  • Coding agent: GPT-5 nano (OpenAI) at $17.24/mo — next cheapest is DeepSeek V4 Flash at $19.51
  • Bulk classification: GPT-5 nano (OpenAI) at $50.40/mo — next cheapest is Gemini 2.5 Flash-Lite at $88.80

Those are absolute cheapest-per-shape, across every tier. Scroll down for the comparison that usually matters more — the same five shapes restricted to frontier-tier models, where the winner changes less than you would expect.

Compare LLM API Costs Side by Side

Put your own numbers in. Tick models from any of the five providers and rank them on the identical workload — including cache hit rate and the batch discount where it exists.

Start from a typical workload
Models to compare (5 selected)
OpenAI
Anthropic
Google
xAI
DeepSeek
ModelInput costOutput costPer requestMonthly total
DeepSeek V4 ProDeepSeekCHEAPEST$62.77$8.70$0.0036$71.47
GPT-5OpenAI$184.50$100.00$0.01$284.50
Gemini 2.5 ProGoogle$184.50$100.00$0.01$284.50
Grok 4.5xAI$298.80$60.00$0.02$358.80
Claude Sonnet 5 (intro rate)Anthropic$295.20$100.00$0.02$395.20
On this exact workload, Claude Sonnet 5 (intro rate) costs 5.5x what DeepSeek V4 Pro does — $323.73 more per month for the same number of tokens. Model choice is usually a bigger lever on your bill than any amount of prompt golfing.

You picked a model. The harder question is whether AI picks you.

Buyers increasingly ask an assistant instead of a search box. Our free scan puts your category to ChatGPT across 5 prompt angles and scores how often your product is the one it names.

Cheapest model, per workload shape

Five realistic application shapes, every model priced on each, standard synchronous rates with the stated cache hit rate and no batch discount. Top five shown per shape, plus the most expensive model on the same job for scale.

Support chatbot

A ~1,200-token system prompt plus a short conversation history, answered in a paragraph or two. 50,000 conversations a month. Most of the input repeats every turn, so caching bites hard.

2,500 in · 350 out · 50,000 req/mo · 60% cached

#ModelProviderMonthly
1GPT-5 nanoOpenAI$9.88
2DeepSeek V4 FlashDeepSeek$12.11
3Gemini 2.5 Flash-LiteGoogle$12.75
4GPT-4.1 nanoOpenAI$13.88
5Gemini 2.0 FlashGoogle$13.88
lastGPT-5.5 ProOpenAI$6,900

Spread on this shape: 699x between the cheapest and the most expensive model.

RAG question answering

Eight retrieved chunks (~1,000 tokens each) stuffed into the prompt, answered with a cited paragraph. Input-heavy: the retrieved context dwarfs the answer.

9,000 in · 500 out · 20,000 req/mo · 20% cached

#ModelProviderMonthly
1GPT-5 nanoOpenAI$11.38
2Gemini 2.5 Flash-LiteGoogle$18.76
3GPT-4.1 nanoOpenAI$19.30
4Gemini 2.0 FlashGoogle$19.30
5DeepSeek V4 FlashDeepSeek$23.06
lastGPT-5.5 ProOpenAI$7,200

Spread on this shape: 633x between the cheapest and the most expensive model.

Document summarization

A 20-page document (~15,000 tokens) reduced to a one-page summary. Almost pure input cost, and a textbook batch-API candidate because nobody is waiting on the response.

15,000 in · 800 out · 5,000 req/mo · 0% cached

#ModelProviderMonthly
1GPT-5 nanoOpenAI$5.35
2GPT-4.1 nanoOpenAI$9.10
3Gemini 2.5 Flash-LiteGoogle$9.10
4Gemini 2.0 FlashGoogle$9.10
5DeepSeek V4 FlashDeepSeek$11.62
lastGPT-5.5 ProOpenAI$2,970

Spread on this shape: 555x between the cheapest and the most expensive model.

Coding agent

A large repo context plus tool definitions, and long generated diffs. Output-heavy and expensive: this is the shape where the input/output price gap really shows up.

30,000 in · 4,000 out · 8,000 req/mo · 70% cached

#ModelProviderMonthly
1GPT-5 nanoOpenAI$17.24
2DeepSeek V4 FlashDeepSeek$19.51
3Gemini 2.5 Flash-LiteGoogle$21.68
4GPT-4.1 nanoOpenAI$24.20
5Gemini 2.0 FlashGoogle$24.20
lastGPT-5.5 ProOpenAI$12,960

Spread on this shape: 752x between the cheapest and the most expensive model.

Bulk classification

Tag a short record with one of 12 labels. Tiny prompts, tiny answers, enormous volume. Per-call cost is a rounding error; the monthly bill is entirely about volume and unit price.

600 in · 15 out · 2,000,000 req/mo · 40% cached

#ModelProviderMonthly
1GPT-5 nanoOpenAI$50.40
2Gemini 2.5 Flash-LiteGoogle$88.80
3GPT-4.1 nanoOpenAI$96.00
4Gemini 2.0 FlashGoogle$96.00
5DeepSeek V4 FlashDeepSeek$110.54
lastGPT-5.5 ProOpenAI$41,400

Spread on this shape: 821x between the cheapest and the most expensive model.

The comparison you probably want: frontier tier, head to head

The table above is dominated by nano, lite and flash models, which is correct arithmetic and slightly unhelpful advice — most teams have already decided roughly what capability tier they need. So here is the same five workloads restricted to one flagship-ish model per provider: GPT-5, Claude Sonnet 5, Gemini 2.5 Pro, Grok 4.5 and DeepSeek V4 Pro.

WorkloadClaude Sonnet 5 (intro rate)DeepSeek V4 ProGemini 2.5 ProGPT-5Grok 4.5
Support chatbot$290.00$37.25$246.88$246.88$227.50
RAG question answering$395.20$71.47$284.50$284.50$358.80
Document summarization$190.00$36.11$133.75$133.75$174.00
Coding agent$497.60$59.77$431.00$431.00$386.40
Bulk classification$1,836$341.04$1,260$1,260$1,764

Green cells are the cheapest per row. Across five workload shapes at the frontier tier there are 1 distinct winners. That is the whole argument of this page in one number: an article that tells you "X is the cheapest LLM API" has, at best, answered a question about one workload shape and not told you which.

Five providers, and any of them can reprice next week

This comparison is only useful while it's current. Leave your email and we'll send the delta whenever one of the five changes a published rate.

Already on the calendar: Claude Sonnet 5's introductory $2 / $10 per 1M tokens expires Aug 31, 2026 and goes to $3 / $15 — a 50% increase on the same workload. Prices here were last re-verified July 28, 2026.

How the five providers differ structurally

Rates are one thing; pricing structure is what makes a provider expensive or cheap for you specifically. Everything in this table is computed from the verified rate data.

ProviderModelsInput rangeTier spreadOutput ÷ inputAvg cache discountBatch
OpenAI20$0.05$30.00600x4.08.0x8x50% off
Anthropic8$1.00$10.0010x5.05.0x10x50% off
Google8$0.10$2.0020x4.08.3x9x50% off
xAI3$1.00$2.002x2.03.0x6xNone
DeepSeek2$0.14$0.433x2.02.0x85xNone

Tier spread is a routing lever

OpenAI's range spans 600x from cheapest to dearest model, which makes "route the easy 80% down a tier" a real strategy. xAI's spans only 2x, so the same strategy barely exists there.

Output ratio decides workload fit

Anthropic is a flat 5x on every model. OpenAI runs 4x–8x. xAI and DeepSeek sit at 2x–3x, which is why they punch above their rate card on generation-heavy work and below it on retrieval-heavy work.

Discounts do not stack equally

A GPT-5 workload that is fully cached and batched reaches $0.06 per million input. Grok 4.5's best case is $0.30. Same ballpark list price, very different floor.

Two providers charge by prompt length

Google and xAI roughly double their rate above a 200,000-token prompt, on the whole prompt. OpenAI, Anthropic and DeepSeek are flat at any length. For long-context work this outranks the headline rate.

Sources. Every rate above was read on July 28, 2026 from: OpenAI · Anthropic · Google · xAI · DeepSeek.

Our recommendations

Price is a shortlisting tool, not a decision. These are the calls we would make on cost grounds, to be confirmed with an evaluation on your own traffic.

Prototyping with no budget → Gemini

It is the only provider of the five with a standing free tier rather than expiring credit. Build on free Flash, and move to the paid tier before any real customer data touches it — free-tier content is used to improve Google's products, which is documented on their pricing page.

Very high volume, mechanical work → DeepSeek, or a nano/lite tier

For classification, tagging, extraction and routing at millions of calls a month, DeepSeek V4 Flash and its 50x cache discount are hard to beat — but check GPT-5 nano and Gemini 2.5 Flash-Lite first, because with batch and caching applied they get remarkably close and carry none of the data-jurisdiction question.

Long-form generation and coding agents → Grok deserves a look

Output-heavy workloads are the one shape where xAI's 2x–3x output-to-input ratio genuinely beats the 4x–8x norm. The caveat is that there is no batch discount and a shallow cache discount, so your estimate at architecture time is close to your final bill.

Anything with a large fixed prefix → Claude or OpenAI

Both stack a 10x cache discount with a 50% batch discount cleanly. Note Anthropic's cache-write premium (1.25x for a 5-minute TTL, 2x for an hour) makes caching a loss if the prefix is reused fewer than two or three times — and that Claude Sonnet 5's introductory rate rises 50% on September 1, 2026.

Mixed interactive and offline traffic → route, don't choose

The largest saving available to most teams is not a provider switch. It is splitting traffic: cheap tier for the easy majority, frontier tier on an explicit escalation signal, and everything with no human waiting moved to a batch endpoint at half price. That routinely beats any single-vendor choice on this page.

What actually moves your bill, ranked

Taken from the spreads computed above, largest lever first.

  1. Capability tier. The spread between the cheapest and dearest model on a single workload above reaches 821x. Nothing else comes close. Most production traffic does not need a frontier model.
  2. Batch, where it exists. A flat 50% off both directions on three of five providers, at zero quality cost, for anything with no human waiting.
  3. Prompt caching. 10x on most providers, up to 120x on DeepSeek, as low as 5x on xAI. Put stable content at the front of the prompt so the cacheable prefix is as long as possible.
  4. Output length. Output costs 2x–8x input depending on provider, so every token you stop generating is worth several you stop sending. Ask for structured output, set a maximum length, and stop asking the model to restate the question.
  5. Staying under 200k tokens on Google and xAI. Crossing that line roughly doubles the rate on the entire prompt.
  6. Retrieved-context discipline. RAG systems routinely send twenty chunks where five would do. Rerank, then truncate — usually the single largest line on a RAG bill.
  7. A hard spend cap. Not an optimisation, a seatbelt. Every provider console supports one, and a retry loop shipped on a Friday is the classic way a $50 estimate becomes a $5,000 invoice.

LLM API pricing comparison FAQ

Which LLM API is cheapest overall?

Overall, DeepSeek V4 Flash at $0.14 input / $0.28 output has the lowest published rates of the 41 models compared here. But "cheapest overall" is close to a meaningless question, because the winner changes with workload shape, and because two of the five providers sell discounts deep enough to reorder the whole table. Once you apply batch and caching, GPT-5 nano reaches an effective input rate below any DeepSeek rate. The comparison worth running is per workload, which is what the table on this page does.

Why does the cheapest model change depending on the workload?

Because providers price input and output very differently relative to each other, and because their discount structures do not overlap. Anthropic charges exactly 5x output-to-input on every model. OpenAI ranges from 4x to 8x. xAI is as low as 2x. So a model that is expensive on an input-dominated RAG job can be the cheapest on an output-dominated generation job at the same total token count. Layer on that OpenAI, Anthropic and Google offer 50% batch discounts while xAI and DeepSeek do not, and cache discounts that range from 5x to over 100x, and there is no ordering that survives a change of workload.

Should I just pick the cheapest model and move on?

No, and this is the most expensive mistake in the category. Token price is only one term. A cheaper model that needs two attempts costs 2x its rate. One that needs a longer prompt to reach the same quality costs more per call. One that burns more reasoning tokens bills the difference at the output rate. And a wrong answer that reaches a customer costs engineering time and goodwill that no rate card captures. Use price to narrow a shortlist to two or three candidates, then decide between them with an evaluation on your own traffic — never the other way round.

Which provider has the best prompt caching?

DeepSeek by a wide margin: cache hits run at roughly 50x to 120x below the miss rate, which effectively removes cached tokens from the cost model. OpenAI, Anthropic and Google cluster at about 10x. xAI is the shallowest at roughly 5x to 7x. The structures also differ beyond the ratio — Anthropic charges a premium of 1.25x or 2x to write to the cache, Google charges $1.00 per million tokens per hour to keep a cache alive, and OpenAI charges neither. Copying a caching strategy from one provider to another without re-checking those details will quietly cost you money.

Which providers offer a batch discount?

OpenAI, Anthropic and Google all take 50% off both input and output for asynchronous work returned within 24 hours. xAI and DeepSeek publish none. Google additionally sells a Priority tier in the other direction, charging a premium for lower latency. If a meaningful share of your workload is offline, the batch discount is usually a larger lever than switching providers, and it is the only one on this page that costs you nothing in quality.

Which LLM API has a free tier?

Only Google, among these five. Google's pricing page lists "Free of charge" input and output on several Gemini Flash and Flash-Lite models, subject to rate limits, with the documented trade-off that free-tier content is used to improve Google's products. The others may offer promotional starting credit but publish no standing free allowance. If free-to-build matters, that is a one-provider decision.

Do any providers charge more for longer prompts?

Yes — Google and xAI, and it catches people out constantly. Gemini 2.5 Pro and 3.1 Pro, and several Grok models, roughly double their rate above a 200,000-token prompt, with the higher rate applying to the whole prompt rather than the excess. OpenAI, Anthropic and DeepSeek charge one flat rate at any length. If you run long-context workloads, this single structural difference can matter more than the headline rates.

How current are these prices and where did they come from?

Every rate on this page was read directly off the provider's own pricing page on July 28, 2026, and each provider's source URL is linked in the comparison table. No aggregator, no third-party summary, and no model that could not be verified against a first-party page. LLM API prices move often — and at least one price change is already scheduled, with Claude Sonnet 5's introductory rate expiring August 31, 2026 — so confirm against the linked source before committing a budget.

Per-token pricing for the other providers

Every page in this set reads from the same verified price table, so the numbers never disagree with each other.

Related

ChatGPT already recommends your AI app. Does it recommend yours?

If you're building in LLM APIs, run a free AI-visibility scan on your own product — we ask ChatGPT across 5 prompt angles and score how often you get named. ~30 seconds, no signup, no card.

📬 Get the best new AI tools delivered weekly

One concise email with fresh launches, trending picks, and featured standouts.