Xiaomi’s MiMo V2.6 Pro is a free, MIT-licensed open-weight model released September 21, 2026, that jumped from 26 to 46 on Artificial Analysis’s Intelligence Index in five months β the highest score of any open-weight model. It’s genuinely the cheapest capable option on the API ($0.435/$0.87 per million tokens), it’s strong at agentic automation specifically, and despite being “open-weight,” almost nobody reading this will actually run it themselves.
A 1.02-trillion-parameter model just became the top-ranked open-weight system in the world, and the headline number is real. What’s less obvious from the announcement posts is that the model ranks differently depending on which benchmark you check, takes several seconds just to start responding, and is sized for a datacenter, not a desktop.
The Jump From V2.5 Is Real, Not Marketing
MiMo V2.5 Pro scored 26 on Artificial Analysis’s Intelligence Index in April 2026. MiMo V2.6 Pro scores 46, released five months later β ahead of GLM-5.3, Kimi K3, and DeepSeek V4.1 Flash, and well above the median score of 18 for open-weight models its size. Both the Pro (1.02T total parameters, 42B active) and Flash (310B total, 15B active) variants ship under a clean MIT license with no revenue threshold or MaaS carve-out, unlike GLM-5.3’s bespoke license, which only restricts operators over $10B in revenue.
The architecture is genuinely new for this line: MiMo V2.6 is natively omnimodal, meaning text, image, video and audio all feed into the same model rather than a vision adapter bolted onto a text backbone, with a 1,048,576-token context window across all of it.
One Composite Score, Two Different Rankings
Artificial Analysis puts Pro ahead of Flash, which is the expected order. Vals AI’s independent index doesn’t agree: Flash ranks #16 at 59.58%, with Pro right behind it at #17, 59.47% β a smaller model beating its own bigger sibling by a hair on that particular index. Neither number is wrong; they’re testing different things, the same pattern this site keeps finding whenever the same model gets different scores from different trackers.
The category split matters more than the overall number. On AutomationBench, a benchmark reproduced by OpenLM, Pro scores 53.1 against Claude Opus 5’s 50.3 and GPT-5.6 Sol’s 45.8 β a real win on agentic, tool-using tasks. But coverage of the same release describes Pro as weaker on ProgramBench and Terminal Bench 4.0 than closed competitors. A single “46” or “#17” tells you nothing about which of those two situations you’re in.
| Benchmark | MiMo V2.6 Pro result | What it means |
|---|---|---|
| Artificial Analysis Intelligence Index | 46 (highest of any open-weight model) | Composite score across many task types |
| Vals AI Index | #17 overall, behind its own Flash variant (#16) | Pro isn’t uniformly the stronger pick even within its own family |
| AutomationBench | 53.1, ahead of Claude Opus 5 (50.3) and GPT-5.6 Sol (45.8) | Genuine strength: agentic, tool-using workflows |
| ProgramBench / Terminal Bench 4.0 | Weaker than some closed competitors | Not the pick for pure coding-correctness tasks |
Fast Throughput, Slow to Start
Once it’s generating, MiMo V2.6 Pro moves at a reasonable clip β Artificial Analysis clocks it around 41 to 130 tokens per second depending on the measurement window. Getting to that first token is the problem: Artificial Analysis measured a 3.93-second time to first token, and LLM Stats’ trailing seven-day p95 figure puts it at 8.55 seconds. For a chat interface where someone is watching the screen, that’s a genuinely noticeable pause before anything appears β a different kind of latency than the per-token pricing games we’ve covered in our reasoning-token pricing breakdown, but a real cost to interactive use either way.
The Price Is the Actual Headline
| Model | Input $/M tokens | Output $/M tokens | License |
|---|---|---|---|
| MiMo V2.6 Pro | $0.435 | $0.87 | MIT (open weight) |
| MiMo V2.6 Pro (cached input) | $0.0036 | β | β |
| Kimi K3 (for comparison) | ~7x MiMo’s rate | ~17x MiMo’s rate | Modified MIT |
| DeepSeek V4 | Volatile, surge-priced by time of day | Volatile | MIT |
At roughly a seventh of Kimi K3’s input cost and a seventeenth of its output cost, MiMo V2.6 Pro is currently the cheapest capable model on the API by a wide margin, with a 99% cache-hit discount on top of that for repeated system prompts. Unlike DeepSeek V4’s pricing, which moved through six changes in under a year including a surge-pricing scheme tied to time of day, Xiaomi says these rates carry over unchanged from the V2.5 series β though given how this category has behaved all year, that’s worth rechecking before you build a budget around it rather than assuming it holds.
“Open-Weight” Doesn’t Mean You’re Running This at Home
The MIT license means anyone can download and self-host MiMo V2.6 Pro. Almost nobody actually can. At 1.02 trillion total parameters, this sits well past the top of the RAM-tier table in our local AI models guide, which tops out at 128GB systems running quantized 100-120B-class models like GLM-5.2 or DeepSeek-V4-Flash. MiMo V2.6 Pro’s 42B active parameters per token don’t change the total weight footprint that needs to fit somewhere, and that puts it in the same “open-weight in name, datacenter-only in practice” category we flagged for GLM-5.3-Flash, which still needs four H200 GPUs to run at all.
The Flash variant (310B total, 15B active) is meaningfully more reachable, but still nowhere close to a single consumer GPU. For this site’s readers, “open-weight” here mainly means license terms and API pricing, not something to download tonight.
Who This Actually Fits
MiMo V2.6 Pro makes sense for cost-sensitive API workloads that are agentic or automation-heavy β the kind of multi-step tool-calling tasks covered in our coding-agent harness piece β where the price gap versus Kimi K3 or a closed frontier model compounds fast across thousands of calls, and where a few extra seconds before the first token doesn’t matter because nobody’s watching it type. It’s a weaker fit for interactive chat products where latency is felt immediately, and for anyone drawn in by “open-weight” expecting to self-host without enterprise-grade hardware.
Frequently Asked Questions
Is MiMo V2.6 Pro better than GLM-5.3 or Kimi K3?
On Artificial Analysis’s composite Intelligence Index, yes β it scores higher than both. On agentic automation tasks specifically, it also leads. On ProgramBench and Terminal Bench coding tests, it’s reported weaker than some closed competitors, so the honest answer depends on the task.
Can I run MiMo V2.6 Pro on my own computer?
Not realistically. At 1.02 trillion total parameters, it’s well beyond what any consumer or typical small-business hardware setup can handle, even quantized. The MIT license permits self-hosting, but the practical requirement is enterprise-grade multi-GPU infrastructure.
How much does MiMo V2.6 Pro cost to use via API?
$0.435 per million input tokens and $0.87 per million output tokens on Xiaomi’s own API, with cached input at $0.0036 per million. That’s roughly a seventh of Kimi K3’s input price and a seventeenth of its output price.
Why does MiMo V2.6 Flash outrank MiMo V2.6 Pro on some benchmarks?
On Vals AI’s index, Flash edges out Pro by a fraction of a point (59.58% vs. 59.47%). Different benchmarks weight different task types differently, so a smaller sibling model beating a larger one on one particular index isn’t a contradiction, just a different measurement.
Is MiMo V2.6 Pro good for real-time chat applications?
Its time to first token is slow for interactive use, measured between roughly 4 and 8.5 seconds depending on the source and time window, even though its per-token generation speed once running is reasonable. It’s a better fit for background or agentic workloads than for a live chat interface.
Shurah is the founder of AI Tools Daily, tracking pricing, licensing and policy changes across AI tools so readers can make decisions without wading through marketing claims themselves.