DeepSeek's New Model Runs a Full Benchmark for 3 Cents. Claude Costs $3.15 for the Same Work.
Independent firm Artificial Analysis says DeepSeek's V4-Flash is by far the cheapest well-known AI model to actually run — about $0.03 per test battery versus $0.86 for Kimi K3, $1.86 for GPT-5.6 Sol and $3.15 for Claude Fable 5. The catch: it scores 50 on intelligence where the frontier scores 66. Here's why the ratio still matters.

Run the same battery of benchmark tests through DeepSeek's newest model and it costs about three cents.
Run it through Anthropic's Claude Fable 5 and it costs $3.15.
That's the finding from independent research firm Artificial Analysis, published today: DeepSeek's V4-Flash is now by a wide margin the cheapest well-known AI model to actually run. Not marginally cheaper. Roughly 105 times cheaper than Claude Fable 5 for the same work.
The numbers
Artificial Analysis measures something more useful than sticker price: what it costs to push a model through its full Intelligence Index test battery, start to finish. Here's how the field landed:
- DeepSeek V4-Flash — ~$0.03
- Moonshot Kimi K3 — $0.86 (~29× more)
- OpenAI GPT-5.6 Sol — $1.86 (~62× more)
- Anthropic Claude Fable 5 — $3.15 (~105× more)
DeepSeek's published API rates are $0.14 per million input tokens and $0.28 per million output tokens — and if your request hits its cache, input drops to $0.0028 per million. That is not a typo.
Why "cost per test" is the number that matters
Here's the part most coverage skips, and it's the whole reason this benchmark exists.
Per-token pricing is a misleading way to compare models, because models don't spend tokens equally. A reasoning model that thinks out loud for 4,000 tokens before answering can easily cost more to finish a task than a "pricier" model that gets there in 600. Sticker price tells you the rate; it tells you nothing about the meter.
Cost per test folds both together — the rate and how efficiently a model spends its way to an answer. As Artificial Analysis frames it, cost per test turns on token efficiency as much as on the price card. V4-Flash wins on both ends at once: cheap tokens, and it doesn't ramble to get to the point.
That's why the gap blows out from "a few times cheaper" on the price sheet to two orders of magnitude in practice.
The catch: cheapest is not smartest
Now the honest part, and it's a real limitation.
On the same Intelligence Index that generated those costs, V4-Flash scores 50 out of 100. Here's where everyone else sits:
- Claude Opus 5, Claude Fable 5, GPT-5.6 — ~66
- Moonshot Kimi K3 — 57
- Meta Muse Spark 1.1, Z.ai GLM-5.2 — 51
- Google Gemini 3.6 Flash — 50
- DeepSeek V4-Flash — 50
So this is not a story about DeepSeek beating the frontier. A 16-point gap between 50 and 66 is a meaningful capability gap, and on the hardest reasoning work the expensive models are still the ones you want.
The real story is the ratio. V4-Flash delivers roughly 76% of frontier-model intelligence for under 1% of the frontier-model cost. For an enormous share of production workloads — classification, extraction, summarization, routing, chat, bulk code edits — that trade is not close. You do not need a 66 to tag support tickets.
What's under the hood
V4-Flash is a Mixture-of-Experts model with 284 billion total parameters but only about 13 billion active at any one time — which is precisely how it stays cheap. You pay to run a fraction of the model on each token, not the whole thing. It carries a 1M-token context window, and entered public beta on July 31, 2026.
It's not weak across the board, either. It posts 82.7 on Terminal Bench — reportedly about ten points above DeepSeek's own Pro model — and 90.8% on GPQA Diamond. The averaged Intelligence Index score hides some genuinely strong individual results.
Why this is a problem for someone
DeepSeek isn't stumbling into these prices. It made a 75% discount permanent earlier this year and is sitting on more than $7 billion in funding. This is a deliberate play for the high-volume floor of the market — the chatbots, coding assistants and automation pipelines where token bills compound fast and nobody is paying for the last 16 points of reasoning.
That's uncomfortable for a specific group of people. A number of Western AI companies have valuations — and IPO plans — built on the assumption that inference stays a premium product with premium margins. Analysts have already flagged that this kind of pricing pressure undercuts exactly that assumption. It's hard to defend a premium tier when a competitor delivers three-quarters of the capability at one-hundredth of the price, and does it in the open.
Put this alongside the other big number from this week — OpenAI's Astra producing new mathematical proofs for about $2,000 — and a theme comes into focus. The interesting AI story in mid-2026 isn't only what these systems can do. It's how fast the price of having them do it is collapsing.
Frequently asked questions
How much does DeepSeek V4-Flash cost?
Published API pricing is $0.14 per million input tokens and $0.28 per million output tokens, dropping to $0.0028 per million for cache-hit input. In Artificial Analysis's benchmark run, completing the full Intelligence Index test battery cost roughly $0.03.
Is DeepSeek V4-Flash better than GPT-5.6 or Claude?
No — it's cheaper, not better. V4-Flash scores 50 on Artificial Analysis's Intelligence Index, while GPT-5.6, Claude Fable 5 and Claude Opus 5 sit around 66. It's competitive with Google's Gemini 3.6 Flash (50) in the efficiency tier, and behind Kimi K3 (57).
Why is it so much cheaper to run than its price sheet suggests?
Two reasons stack. Its per-token rates are low, and it's token-efficient — it uses fewer tokens to finish the same task. Since real cost is rate multiplied by tokens consumed, a model that is both cheap and concise pulls far ahead of one that is merely cheap.
What is V4-Flash's architecture?
A Mixture-of-Experts design with roughly 284 billion total parameters and about 13 billion activated per token, plus a 1M-token context window. Activating only a small slice of the network per token is the main source of its cost advantage.
When was it released?
The DeepSeek-V4-Flash API entered public beta on July 31, 2026.
Should I switch my app to it?
It depends entirely on the workload. For high-volume, well-defined tasks — extraction, classification, summarization, routing, chat — the cost difference is difficult to argue with. For hard multi-step reasoning, agentic work or anything where a subtle mistake is expensive, the 16-point capability gap against frontier models is the thing to weigh. Many teams end up routing between both.
Reporting based on Artificial Analysis's Intelligence Index cost findings published August 3, 2026, and coverage from TNW, Reuters and others. Benchmark scores are point-in-time and shift as models are updated. The Bot Post will update this story as new figures land.
About the author
UbedullaFounder & Editor
Founder and editor of The Bot Post, covering AI news and technology.


