It feels suspicious that MiMo-V2.6
Pro gets 46 in de index while DeepSeek-V4.1 (https://artificialanalysis.ai/models/deepseek-v4-1-flash) gets 39. According to the appendix at the bottom of https://mimo.xiaomi.com/mimo-v2-6 the deepseek model sometimes surpasses mimo and it's not so far behind in capabilities. A week ago opus 5 appeared 1 points ahead of fable 5 despite fable being a much smarter model (this has been corrected already)
The main AA benchmark keeps changing, and had to be radically changed when Astra came out and showed zero improvement over GPT 5.6 Sol in their benchmark. Opus 5 is still 1 point ahead of Fable 5.0 on the index, if you manually add Fable 5.0 back into the list, so it hasn't actually been "corrected". It's only Fable 5.1 that is shown as ahead of Opus 5.
The AA benchmark is a weighted average of other benchmarks and some internal ones. I think the difficult part is finding benchmarks that reflect your own use of the models.
The way Artificial Analysis keeps changing their weights feels kind of like deciding who the winner should be and making the weights reflect that. They’ve been changing their weights to add more weight to improved long-running agentic capabilities, but doing so means they’re reducing the relative importance of world knowledge and of writing ability.
I’ll grant that maybe world knowledge isn’t that important for these models. But writing ability is important for human understanding, and I think the weird turns of phrase and word choices reflect the labs’ underweighting of the importance of human understanding.
DS 4.1 is good but it’s clearly not as “smart” as non-flash models- it just doesn’t have the training data. Without a solid plan, it goes off the rails pretty regularly.
All of the chinese labs have been overfitting on benchmark data to game the results for a while now - MiMo and Deepseek are not anywhere near frontier and mostly compete with models like Luna - which they are still worse than.
There isn't much compelling reason to use these unless you are just averse to giving money to openai/altman. A $20 codex sub gives you ~$150 of luna use per weekly limit, while there isn't any good subsidized options for chinese models at all (and the few who were subsidizing, like opencode, rugpulled by reducing monthly limit to $60 to $15 with no notice to users).
I mean, why would you? Chinese models through an API are more expensive than any subsidized subscription and the only one offering a transparent subsidy (opencode) only gives you $15 of credit on any relevant models, and still has a 5 hour limit.
rugpull is a very strong word... I agree that they are not great with their pricing/credit communications, but they do state the multipliers quite clearly in various places. And plenty still have $60 or $30.
Changing pricing buried in docs (which they didn't actually do when they started switching every model to $15 limit 2 months ago BTW, they only did this once people noticed and called it out) without notifying users is the definition of a rugpull.
It is an impressive model. Agreed on most that is written on this page, with the exception of it being fast. I ran it on my own LLM benchmark suite[1] and it is faster than DeepSeek but still much slower than leading models. But it's pricing is where it really shines.
KillSwitch-Bench 1.0
Claude Opus 5 66.9
GPT-6 Astra 57.9
Claude Fable 5.1 46.7
MiMo-V2.6-Pro 38.8
Muse Spark 1.3 36.5
Why does Luna have a score of 0? I will say that in my limited experience, I use Luna and DSv4 Flash (haven't tried 4.1 yet) and Luna is wayyyy faster. They both output at the same speed but DS has an endless thought process
OpenAI usage limits have been severely cut, and intelligence appears to be markedly declining, so I'm going to start trying these Chinese models seriously now. I don't mind if it takes longer. I just need the intelligence to predictably work the same way from day to day.
I strongly agree. Check out the Codex subreddit. Many empirical examples of Astra silently downgrading the models. One found Astra was silently using Luna Max (but still billing for Astra).
Even when I try to stick with Sol X/High, my limits are at best half of what they were before Astra launched, and the intelligence has declined markedly.
I cancelled my $100 plan. This is absolutely absurd and frankly unusable now.
Entirely depends on the work asked and the repo involved.
When I’m doing work on a repo where I’m implementing a standard and the agents have to read the standard to keep from hallucinating my usage skyrockets.
Hell this changes depending on which language I’m working with.
Weird. That's definitely not my experience. On the $100 plan and I can burn my entire week's budget in a few hours easily with Astra. It's borderline unusable.
People often use a bunch of subagents, poor context management (though Codex's tight context limits and constant compaction tend to mitigate this), a bunch of projects at once, etc., as well as not using workflows that do heavy planning once up front and then consult it rather than thinking endlessly about what to do during the implementation part.
It's also the case that working on massive codebases is just a different beast. If they've been slopmining a monorepo for months with 200x, then their codebase is probably Lovecraftian at that point and requiring extensive effort to iterate on.
For me Sol is less efficient than Astra - Sol makes many avoidable mistakes and has issues with context compaction - sometimes it goes haywire after a couple compactions.
Agreed on the “being dumbed down” observation. It appears they’re most powerful at release time and then are gradually “optimized” so every new model feels more powerful. But there’s no evidence on routing to a deployment with other weights. It would be plausible to do so though at least at peak times.
What really is fun is when the opposite of the public outcry about purported distillation happens - when Astra in Codex suddenly responds in Chinese. Now that is fun. Wonder what it’s routing to.
I have a 20x sub and it feels like the 20$ sub eight months ago. I can easily burn it in an afternoon, and I don't have many projects.
According to API usage, they cut you off at around the equivalent of 900$ of API usage, whatever that means. It's very difficult to track all this, and very subjective. What validates me is that of all my friends I am not the only one.
I can only imagine the 100$ users must be feeling the rug being pulled even harder.
Anyway, this has led me to get a Spark, and a second is on the way.
> and intelligence appears to be markedly declining
Serious question: does anyone have evidence of this?
It’s something that’s constantly asserted, and has been since 2023. Every time someone posts a site that tries to track this though, I look at it and it’s just a flat line.
By and large they don’t. I have seen this drop a few times, eg before fable came out opus dropped a lot probably due to less compute available.
My guess is it’s a combination of getting used to the new cliff models fall off on and forgetting that model performance drops significantly when context fills up.
So new model comes out, people try it and it’s amazing on a task or two. Then they start using it, context window fills up and it gets a lot worse.
Sol 5.6 xhigh had been a very reliable workhorse for coding for me via the 200 bucks sub.
But this week they seem to have tweaked the system to a point at which all models (Astra, Sol, Luna) hit rate limits all_the_time without me being anywhere close to the weekly limit.
Early results with MiMo 2.6pro are quite encouraging for anything that's non-UI work so likely switching spend for the time being
I have done the same. I hesitated for way too long. I shouldn’t have.
I get way more usage for way less money without any quality or performance degradation. My $200 Codex Pro plan allowance is depleted in 2-3 days. Sometimes Tibo announces a usage reset. But GPT-5.6 models are really not good for coding. Sol has been making increasingly more mistakes in the past two weeks even in the reviewer and advisor roles. Astra is usable for coding but slow and very expensive. In the past two days I’ve used up over 70% on simple copy editing, with dedicated short specs and short sessions. Really little one can do to make it more efficient. Similar work took 20% at most just a month ago. I’m looking to use Astra for milestone reviews/advisory. Perhaps a $100 Pro downgrade will be enough. But my main work is now on open-weight models. And you don’t need to depend on someone to send you a reset. And it’s cheaper by the end of the month too.
With Claude the limits are not even fun anymore - my weekly $100 Max plan quota is gone in one day on merely review invocations, no coding. And my $200 Pro quota is gone in two with some coding. Sonnet 5 is not usable for coding. And Opus 5 tends to always make a couple avoidable mistakes on every task. Fable 5.1 is ok but tends to ignore skills and to work around explicit instructions. Completely canceled all my Claude subscriptions.
With Qwen 3.8, DeepSeek 4.1 Flash, GLM 5.3 I’ve been getting Opus 5-level performance, with less blah blah and no overengineered churn. Public benchmarks are really not telling the real story. The models are more dependable and more predictable. They have their own failure modes. Sometimes DeepSeek 4.1 Flash is quite stubborn but it fails in a good way. Bad for full autonomy - I need to intervene, but it sticks to the rails and instructions - other than Fable and Opus that try to outsmart you and your harness.
Grok is interesting but has been a bit underwhelming on Grok plans - my SuperGrok allowance is depleted in a single session overnight. SuperGrok+ gives more but it’s still about the same as with OpenAI, Claude is way less now.
Since the allowance volume has been shrinking with the major model providers, to me, open-weight alternatives are really necessary now to at least maintain the momentum and budget.
But at work it’s really an uphill challenge - it’s become impossible to convince the tech leadership once they got hooked on Anthropic. They No facts will help. Some people underestimate how expensive Claude really is after getting used to the subscription plans with allowance resets. OpenAI models are expensive too.
Per Xiaomi, MiMo v2.6 training run cost $3.47m. A far cry from the estimated costs ($100m+) for the Big 5 (MSL, xAI, GDM, OAI, Ant). I wouldn't be surprised if salaries and R&D costs have similar drastic disparities.
For a model that matches Muse Spark 1.3 in benchmarks, MiMo v2.6 Pro is incredibly cheap, given its cache rates will remain $0.0036 per million.
I sorta got the impression that the $3.47 million only covered post-training , given that few of the graphs start at zero. Is a barely-trained model going to score 48 on DeepSWE v1.1 ?
Mimo2.5 is really good, but tended to loop too much for my taste. Locally, Pro2.5 wasn't much better. I would reach for it for one shots, hopefully they sorted it out with v2.6, it's a model that's slept on by many. I found that most people that used it did so because it was free. It's a top model worth exploring if you have never given it a go.
jampekka | 15 hours ago
throwa356262 | 15 hours ago
Human error means this wasn't just stopped together by some bot.
jampekka | 15 hours ago
[OP] theanonymousone | 15 hours ago
MichaelNolan | 8 hours ago
egeres | 14 hours ago
GodelNumbering | 12 hours ago
Why?
big-chungus4 | 9 hours ago
W what?
SyneRyder | 12 hours ago
The AA benchmark is a weighted average of other benchmarks and some internal ones. I think the difficult part is finding benchmarks that reflect your own use of the models.
seahorseemoji | 11 hours ago
I’ll grant that maybe world knowledge isn’t that important for these models. But writing ability is important for human understanding, and I think the weird turns of phrase and word choices reflect the labs’ underweighting of the importance of human understanding.
sipjca | 9 hours ago
conception | 9 hours ago
Shekelphile | 6 hours ago
There isn't much compelling reason to use these unless you are just averse to giving money to openai/altman. A $20 codex sub gives you ~$150 of luna use per weekly limit, while there isn't any good subsidized options for chinese models at all (and the few who were subsidizing, like opencode, rugpulled by reducing monthly limit to $60 to $15 with no notice to users).
staticman2 | 5 hours ago
Shekelphile | 5 hours ago
The answer is just a second $20 subscription.
nchmy | 5 hours ago
Shekelphile | 5 hours ago
dom96 | 14 hours ago
KillSwitch-Bench 1.0
1 - https://bench.killswitch-lang.org/ricardobeat | 11 hours ago
drittich | 9 hours ago
dom96 | 5 hours ago
conception | 8 hours ago
dom96 | 5 hours ago
gandreani | 8 hours ago
kosolam | 14 hours ago
guelo | 12 hours ago
kosolam | 11 hours ago
guelo | 11 hours ago
Gareth321 | 13 hours ago
unsupp0rted | 13 hours ago
I've stopped using Astra entirely and remain on Sol orchestrating Luna Xhigh, but it's still not nearly a week's usage for a week's allotment.
And even then, whenever a new model is about to come out, it feels like the model I'm using is being dumbed down substantially.
I have no evidence for this and can have no evidence for this, but I can vote with my wallet regardless.
Gareth321 | 12 hours ago
Even when I try to stick with Sol X/High, my limits are at best half of what they were before Astra launched, and the intelligence has declined markedly.
I cancelled my $100 plan. This is absolutely absurd and frankly unusable now.
Muromec | 12 hours ago
Gareth321 | 11 hours ago
f6v | 10 hours ago
sjbzbeiks | 9 hours ago
When I’m doing work on a repo where I’m implementing a standard and the agents have to read the standard to keep from hallucinating my usage skyrockets.
Hell this changes depending on which language I’m working with.
christophilus | 10 hours ago
dangoodmanUT | 9 hours ago
sauwan | 8 hours ago
gandreani | 8 hours ago
I wonder what dangoodmanUT is using! This is the time to compare!
viccis | 6 hours ago
It's also the case that working on massive codebases is just a different beast. If they've been slopmining a monorepo for months with 200x, then their codebase is probably Lovecraftian at that point and requiring extensive effort to iterate on.
ralusek | an hour ago
jmaker | 6 hours ago
Agreed on the “being dumbed down” observation. It appears they’re most powerful at release time and then are gradually “optimized” so every new model feels more powerful. But there’s no evidence on routing to a deployment with other weights. It would be plausible to do so though at least at peak times.
jmaker | 6 hours ago
seviu | 5 hours ago
According to API usage, they cut you off at around the equivalent of 900$ of API usage, whatever that means. It's very difficult to track all this, and very subjective. What validates me is that of all my friends I am not the only one.
I can only imagine the 100$ users must be feeling the rug being pulled even harder.
Anyway, this has led me to get a Spark, and a second is on the way.
phoghed | 11 hours ago
Serious question: does anyone have evidence of this?
It’s something that’s constantly asserted, and has been since 2023. Every time someone posts a site that tries to track this though, I look at it and it’s just a flat line.
conception | 9 hours ago
By and large they don’t. I have seen this drop a few times, eg before fable came out opus dropped a lot probably due to less compute available.
My guess is it’s a combination of getting used to the new cliff models fall off on and forgetting that model performance drops significantly when context fills up.
So new model comes out, people try it and it’s amazing on a task or two. Then they start using it, context window fills up and it gets a lot worse.
loloisi | 11 hours ago
But this week they seem to have tweaked the system to a point at which all models (Astra, Sol, Luna) hit rate limits all_the_time without me being anywhere close to the weekly limit.
Early results with MiMo 2.6pro are quite encouraging for anything that's non-UI work so likely switching spend for the time being
jmaker | 6 hours ago
I get way more usage for way less money without any quality or performance degradation. My $200 Codex Pro plan allowance is depleted in 2-3 days. Sometimes Tibo announces a usage reset. But GPT-5.6 models are really not good for coding. Sol has been making increasingly more mistakes in the past two weeks even in the reviewer and advisor roles. Astra is usable for coding but slow and very expensive. In the past two days I’ve used up over 70% on simple copy editing, with dedicated short specs and short sessions. Really little one can do to make it more efficient. Similar work took 20% at most just a month ago. I’m looking to use Astra for milestone reviews/advisory. Perhaps a $100 Pro downgrade will be enough. But my main work is now on open-weight models. And you don’t need to depend on someone to send you a reset. And it’s cheaper by the end of the month too.
With Claude the limits are not even fun anymore - my weekly $100 Max plan quota is gone in one day on merely review invocations, no coding. And my $200 Pro quota is gone in two with some coding. Sonnet 5 is not usable for coding. And Opus 5 tends to always make a couple avoidable mistakes on every task. Fable 5.1 is ok but tends to ignore skills and to work around explicit instructions. Completely canceled all my Claude subscriptions.
With Qwen 3.8, DeepSeek 4.1 Flash, GLM 5.3 I’ve been getting Opus 5-level performance, with less blah blah and no overengineered churn. Public benchmarks are really not telling the real story. The models are more dependable and more predictable. They have their own failure modes. Sometimes DeepSeek 4.1 Flash is quite stubborn but it fails in a good way. Bad for full autonomy - I need to intervene, but it sticks to the rails and instructions - other than Fable and Opus that try to outsmart you and your harness.
Grok is interesting but has been a bit underwhelming on Grok plans - my SuperGrok allowance is depleted in a single session overnight. SuperGrok+ gives more but it’s still about the same as with OpenAI, Claude is way less now.
Since the allowance volume has been shrinking with the major model providers, to me, open-weight alternatives are really necessary now to at least maintain the momentum and budget.
But at work it’s really an uphill challenge - it’s become impossible to convince the tech leadership once they got hooked on Anthropic. They No facts will help. Some people underestimate how expensive Claude really is after getting used to the subscription plans with allowance resets. OpenAI models are expensive too.
cmrdporcupine | 5 hours ago
The open model launches and competition generally definitely seem to light a fire under their asses.
ignoramous | 13 hours ago
For a model that matches Muse Spark 1.3 in benchmarks, MiMo v2.6 Pro is incredibly cheap, given its cache rates will remain $0.0036 per million.
NortySpock | 12 hours ago
https://mimo.xiaomi.com/rl/
imjonse | 12 hours ago
drbscl | 12 hours ago
Mostly due to lower cost of living; Shenzhen is way cheaper than SV
f6v | 10 hours ago
drbscl | 7 hours ago
tensegrist | 10 hours ago
iwhalen | 7 hours ago
Interested to see if it also beats DeepSeek V4.1 Flash.
[1]: https://mimo.xiaomi.com/mimo-v2-6
segmondy | 9 hours ago
deanc | 9 hours ago
MallocVoidstar | 9 hours ago
deanc | 9 hours ago