Flash Models Compared: Gemini 3.8, DeepSeek V4.1, Qwen3.8-Omni and MiMo V2.6
Four fast-tier models were released in September 2026 and their published prices differ by roughly eight times on a typical workload. MiMo V2.6 Flash from Xiaomi is the cheapest at $0.14 per million input tokens and $0.28 per million output tokens. Qwen3.8-Omni-Flash is $0.15 and $0.47, and accepts audio and video as well as text and images. DeepSeek V4.1 Flash is $0.15 and $0.60 off-peak but doubles to $0.30 and $1.20 during peak hours, which are 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays. Gemini 3.8 Flash is the most expensive at $0.75 and $3.75, and those are introductory rates that Google's own pricing page says become $1.50 and $7.50 on 1 January 2027. On a month of 10 million input and 2 million output tokens, that is $1.96 for MiMo against $15.00 for Gemini today, and $30.00 for Gemini from January. Neither OpenAI nor Anthropic shipped a new fast-tier model this month. We have not benchmarked quality, so treat this as a pricing and capability comparison rather than a ranking.
Four cheap models in twenty days
Between 2 and 22 September 2026, four labs shipped a model into the same slot: fast, cheap, high volume, aimed at the work that sits underneath an application rather than the work that impresses on a benchmark chart.
They are not equivalent products, and the published prices differ by roughly eight times on a realistic workload. Two of them carry billing conditions that a single per-token figure cannot express at all.
How we compared
Every price below was read from the provider's own pricing page or model page on 29 September 2026 and linked once from each model's heading. We have not run these models against a benchmark harness, and we make no claims about relative quality. Where a specification is reported by a third party rather than the lab, we say so.
The published rates
| Model | Released | Input / 1M | Output / 1M | Cached input / 1M |
|---|---|---|---|---|
| --- | --- | --- | --- | --- |
| MiMo V2.6 Flash | 22 Sep | $0.14 | $0.28 | $0.0028 |
| Qwen3.8-Omni-Flash | 18 Sep | $0.15 | $0.47 | $0.016 |
| DeepSeek V4.1 Flash | 10 Sep | $0.15 off-peak / $0.30 peak | $0.60 off-peak / $1.20 peak | $0.003 off-peak / $0.006 peak |
| Gemini 3.8 Flash | 2 Sep | $0.75 ($1.50 from 1 Jan) | $3.75 ($7.50 from 1 Jan) | $0.075 ($0.15 from 1 Jan) |
Sources: MiMo V2.6 Flash, Qwen3.8-Omni-Flash on Alibaba Cloud Model Studio, DeepSeek pricing, Gemini API pricing.
What a month actually costs
Take a workload of 10 million input tokens and 2 million output tokens a month, which is a modest production feature rather than a toy:
| Model | Monthly cost | From 1 January 2027 |
|---|---|---|
| --- | --- | --- |
| MiMo V2.6 Flash | $1.96 | $1.96 |
| Qwen3.8-Omni-Flash | $2.44 | $2.44 |
| DeepSeek V4.1 Flash (off-peak) | $2.70 | $2.70 |
| DeepSeek V4.1 Flash (peak) | $5.40 | $5.40 |
| Gemini 3.8 Flash | $15.00 | $30.00 |
Those are arithmetic from the table above, not vendor estimates. The headline is not that Gemini is expensive in absolute terms; at fifteen dollars a month nobody is filing a variance report. It is that the ratio holds as you scale, and at a hundred times this volume the same choice is fifteen hundred dollars against two hundred.
The two billing traps
Gemini's price is introductory and Google says so plainly. The pricing page carries the rate "through December 31, 2026" and the replacement rate "starting January 1, 2027" in the same cell. Comparison articles quoting $0.75 without that condition are describing a price with a deadline attached.
DeepSeek charges by the clock. Peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, excluding Chinese public holidays, and everything doubles inside them. For a US-facing application most traffic falls outside those windows, so the effective rate is near the off-peak figure. For a European product, the 06:00 to 10:00 UTC window is the morning rush. The same model can be the cheapest or the second most expensive here depending on where your users wake up.
What you get beyond the price
Cheap is a property; capability is a different one.
- Qwen3.8-Omni-Flash accepts text, images, audio and video, and returns text, with a 1 million token context window, according to TechNode's report of Alibaba's announcement. If your pipeline currently transcribes audio before sending it to a model, this collapses two steps into one. Alibaba also lists a realtime variant with separate audio rates of $0.93 input and $1.87 output per million tokens.
- MiMo V2.6 Flash publishes a 1 million token context window on its model page. Third-party listings describe it as a mixture-of-experts model with 309 billion total parameters and 15 billion active per token, released under an MIT licence; we could not confirm those figures on Xiaomi's own page, so treat them as reported rather than verified.
- DeepSeek V4.1 Flash is the volume workhorse of the group, and its cache-hit price of $0.003 per million tokens off-peak is the lowest number in this article by an order of magnitude.
- Gemini 3.8 Flash is the only fast-tier release here from a US lab, and the only one that arrives inside an enterprise cloud with the contractual apparatus that implies.
Who is missing, and why that matters
Neither OpenAI nor Anthropic shipped a fast-tier model in September 2026. OpenAI's cheapest current option is GPT-5.6 Luna, released in July 2026, and Anthropic's is Claude Haiku 4.5, from October 2025.
That is worth stating because a reader scanning release trackers could reasonably conclude the entire cheap tier refreshed this month. It did not. Three quarters of this cohort came from Chinese labs, and the competitive pressure on price is coming from there rather than from San Francisco. Our frontier model comparison covers the other end of the same market.
Is the cheapest model the right default?
Not automatically, and the reason is not quality.
Three of these four models come from Chinese labs. For a hobby project that is irrelevant. For a company with data residency obligations, a customer contract naming processing locations, or a security review to pass, it is the first question rather than the last one. Alibaba prices this model from its Singapore deployment; DeepSeek and Xiaomi publish from China. None of that makes a model unsafe to use, and we are not suggesting it does. It makes the choice a procurement decision as much as an engineering one, and the person who signs it should know which jurisdiction they are signing into.
The second consideration is operational. A price that doubles on a calendar date and a price that doubles at 06:00 UTC are both fine if you know about them, and both nasty if you discover them in an invoice. If you route across several providers, a gateway makes switching a configuration change rather than a migration, and our guide to AI evals covers how to check that a cheaper model still does the job before you move traffic onto it.
Which should you choose?
If cost per token is the deciding factor: MiMo V2.6 Flash, at $0.14 and $0.28. Confirm it can do your task at acceptable quality first, because there is no independent benchmark in this article and the cheapest model is only cheap if it does not need three attempts.
If your input includes audio or video: Qwen3.8-Omni-Flash. It is the only model here that takes those natively, and folding a transcription step into the model call usually saves more than the token price difference.
If your traffic is concentrated outside 01:00 to 10:00 UTC on weekdays: DeepSeek V4.1 Flash, where you will mostly pay off-peak rates, with cache hits at $0.003 per million tokens if your prompts share a stable prefix.
If you need a US provider, enterprise contracts or Google Cloud adjacency: Gemini 3.8 Flash, budgeted at the January rates of $1.50 and $7.50 rather than today's introductory figures.
If you are choosing on benchmark charts: be careful. Every model here was released this month, most published benchmark comparisons are vendor-run, and fast-tier models are unusually sensitive to prompt format and task shape. Measure on your own workload. Our comparison of inference providers covers the hosting side of the same question, and the open-source model guide covers running weights yourself.
Conclusion
The fast tier is now a price war, and the interesting part is not who is cheapest this week. It is that two of the four prices in this article are conditional: one on a calendar date, one on the hour of the day. A single number in a comparison table cannot carry either condition, which is why most tables covering this cohort are already wrong.
Work out your monthly token volume, apply the rates above to it, and then apply them again as they will be on 1 January. That second number is the one to budget from.
This comparison is an editorial synthesis of vendor pricing pages and model documentation read on 29 September 2026 and linked inline. We did not run benchmarks, measure latency, or receive briefings from these labs. Specifications reported only by third parties are labelled as such. Prices in this category move quickly and one of them changes on a known date; verify before committing spend.
Key Takeaways
- Gemini 3.8 Flash costs $0.75 per million input tokens and $3.75 per million output tokens, but Google's pricing page states those rates apply only through 31 December 2026 and become $1.50 and $7.50 on 1 January 2027. Any budget built on today's figure needs a January line.
- DeepSeek V4.1 Flash is the only model here whose price depends on the clock. Off-peak it is $0.15 input and $0.60 output; during peak hours, 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, both double.
- MiMo V2.6 Flash from Xiaomi publishes the lowest rates in the group at $0.14 and $0.28, with cache hits at $0.0028 per million tokens.
- Qwen3.8-Omni-Flash is the only one of the four that takes audio and video as input alongside text and images, at $0.15 input and $0.47 output, with a 1 million token context window.
- On 10 million input and 2 million output tokens a month, the four cost $1.96, $2.44, $2.70 and $15.00 respectively. The gap is not rounding; it is roughly eight times, and it widens to fifteen times when Gemini's introductory pricing ends.
- Three of the four come from Chinese labs, which makes data residency a procurement question rather than a technical one. Alibaba prices this model from its Singapore deployment.
- Neither OpenAI nor Anthropic released a fast-tier model in September 2026. Their cheapest current options date from July 2026 and October 2025 respectively, which is worth knowing before you assume the whole market moved.
Frequently Asked Questions
Which fast model is cheapest right now?
MiMo V2.6 Flash, at $0.14 per million input tokens and $0.28 per million output tokens on Xiaomi's own page. Qwen3.8-Omni-Flash is close on input at $0.15 but higher on output at $0.47. DeepSeek V4.1 Flash matches Qwen on input off-peak and doubles during peak hours.
Why is Gemini 3.8 Flash so much more expensive than the others?
It is priced as a Google product with Google's infrastructure and support behind it, and it is the only fast-tier model here from a US lab. The more important point is timing: its listed rates are introductory and double on 1 January 2027, which moves it from roughly six times MiMo's cost to about twelve.
What are DeepSeek's peak hours and why do they matter?
DeepSeek's pricing page states peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday, excluding Chinese public holidays. Rates double during those windows. If your traffic is US business hours, most of it lands off-peak; if you run European mornings, much of it does not.
Do these models replace frontier models?
No. Fast-tier models are built for high volume and low latency work such as classification, extraction, routing and summarisation. Complex reasoning, long agentic chains and difficult code still belong on frontier models. Most production systems route between both rather than choosing one.
Does cached input pricing actually save money?
It can be the largest saving available. DeepSeek charges $0.003 per million tokens for cache hits off-peak against $0.15 for a miss, and MiMo charges $0.0028 against $0.14. If your prompts share a long stable prefix, such as a system prompt or a document, caching is worth engineering for.
Where is my data processed with these models?
That depends on the provider and the deployment you use. Alibaba lists this model under its Singapore deployment, and DeepSeek and Xiaomi are Chinese companies. If your organisation has data residency requirements, confirm the processing location in writing before you send anything sensitive.
About the Author
Aisha Patel
AI Editorial Desk
AI Editorial Desk · Web3AIBlog
Aisha Patel is a pen name for our AI editorial desk. Posts under this byline are written and reviewed by our team of contributors with backgrounds in machine learning, large language models, AI infrastructure, and applied research. The desk covers frontier model releases, agent architectures, retrieval-augmented generation, on-device inference, and the engineering tradeoffs that matter when shipping AI in production. Every technical claim is verified against primary sources before publication.