Rendered at 16:06:11 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
joelwallis 19 hours ago [-]
I been using MiMo-V2.5 to do most of my work as software engineer, on a variety of projects I'm working on, and I been VERY happy with ROI. The model is very powerful! Not perfect – I've run in hallucination loops once or twice, but nothing a stop-then-continue wouldn't solve.
The cost is unbelievably low, and the quality of intelligence I get is equivalent to when I was working mostly with Anthropic models (late last year/early this year). I'm fully invested in MiMo and I'm very happy with it.
--
PS: I also check almost daily to see if other models are capable of doing such great work. And they do – DS4F is powerful and DS41 is impressive, GLM 5.3 Flash gets a job done well, etc. – but when I add cost of M-token in the ROI math, Jeez! MiMo is an order of magnitude better.
ehsankia 7 hours ago [-]
> late last year/early this year
That's an eternity when it comes to coding models.
In my personal experience, we've had almost a step change every ~3 months this year, at least for bigger one-shot tasks. For example looking at Gemini Flash 3.0 vs 3.5 vs 3.8, it went 5% -> 30% -> 75% on DeepSWE, all since the start of the year.
pelagicAustral 6 hours ago [-]
tbf, I the happiest I've been working with claude is late last year/early this year (before March)...
epolanski 4 hours ago [-]
That's because Opus 4.6 was the last good assistant model.
Everything after it might be more "intelligent" but is super tuned around end-to-end task (and related benchmarks), not to act as an assistant.
Now it's *you* being the assistant, reviewer, etc.
epolanski 4 hours ago [-]
In one sense you're right.
In another one, Opus 4.6 level already solved 90% of my work-day tasks, so while better models have been instrumental into handling a higher % that does not mean that defaulting on cheaper models can't be good.
I run DS 4.1 flash daily, and then cross check with gpt-6-astra and I've nuked 90% of my AI monthly bill while having higher limits and better performance/intelligence than I did just at the beginning of this summer.
sinuhe69 56 minutes ago [-]
You use the low cost Mimo-V2.5 and not its big brother Mimo-V2.5 Pro? I also made good experience with Mimi-V2.5 when used in conjunction with prewalk mode. But then other models got so cheap and perform better, so I only use Mimo for background tasks.
The availability of Mimo over Openrouter got however, much worse recently.
miyuru 10 hours ago [-]
Same here. It’s the first AI provider I actually gave money to, since they offered the model for free with a Mimo code for the first month or so, and it was great.
These days, there are more intelligent models like DS4.1, but Mimo is very obedient, so I plan things with another model and give the implementation to Mimo.
abustamam 27 minutes ago [-]
Which models do you like for planning? IME I haven't seen much success with using cheap/open models for planning, so I use Claude for planning still.
rapind 15 hours ago [-]
I’ve been very pleased with DS 4.1 flash. Not so much the 4.0 models, but for coding (Rust) it’s been great so far (3 solid days of work).
I’ll give Mimo a try.
trollbridge 14 hours ago [-]
MiMo is my backup whenever DeepSeek is down, had the price bump, is slow, etc.
UltraSpeed was absolutely awesome. I miss it.
DS 4.1 Flash is amazing. Well worth the extra cost.
walrus01 19 hours ago [-]
I've found that mimo v2.5 works for very basic things like a python script to do one thing, but it also is very 'dumb' compared to qwen 3.8-flash-next (I think the benchmark scores for terminal and coding specific benches back this up). And definitely not in the same class as like a GLM5.2 or 5.3. It's fast but makes basic mistakes that only get caught later.
girvo 14 hours ago [-]
The fact I can run Qwen 3.8 Flash Next locally, forever (on my DGX Spark-alike) is genuinely shocking to me. It’s crazy good for how small it is. Fast, too.
And I am not a web developer! It's an extraordinary model.
(Mouse and keyboard required)
walrus01 13 hours ago [-]
Yeah, I'm guessing you have a variant that fits in <128GB with 262k context? I have the unsloth Q8 GGUF of it here in a setup that with full context and ton of extra llama-server "--cache-ram" sits around 200GB RAM usage on a 256GB system, it's probably the best thing I've found for a 256GB class machine. Enough headroom for a rope/yarn extension to 524288 context if I need it.
gmerc 13 minutes ago [-]
RTX6000 Blackwell with 96GB is enough to run it with NV4, 256k context, KVcache, multimodal at 130t/s (SGLang). It's toasty, you're using up 94GB of those 96, but it works and the results are great
girvo 12 hours ago [-]
Yep, the engrams are on NVMe (the speed penalty was lower than I expected) and it is quantised to fit.
It’s good enough that I’m considering a second spark, or selling this and buying an M5 Ultra with 256GB for it
17 hours ago [-]
alwinaugustin 17 hours ago [-]
I am also using 2.5 and it is giving me solid results. Its available free on Openrouter
rahmatawaludin 5 hours ago [-]
Could you elaborate on how to get free mimo access on openrouter?
farlight 3 hours ago [-]
What was meant is probably opencode. It's free there, along with several other models. You have to use their harness to access it. It's alright.
May I ask why you ended up there instead of just using the heavy subsidized subscription. I’m actually curious.
eli 17 hours ago [-]
Mimo has subsidized subscriptions too
james2doyle 19 hours ago [-]
2.5 Pro or the regular 2.5?
I always found that those Mimo models to be really good at tool calling and following instructions
flexagoon 16 hours ago [-]
How does it compare with DS 4.1 Flash in your experience, if you ignore the cost?
baxtr 10 hours ago [-]
Could you elaborate on how you check daily? Do you swap models for certain tasks?
ignoramous 9 hours ago [-]
> GLM 5.3 Flash gets a job done well, etc. – but when I add cost of M-token in the ROI math, Jeez! MiMo is an order of magnitude better.
API may be expensive, but I do 900m tokens (95% cached, ~0.4% output) on Z.ai's $18/mo coding plan with GLM 5.3 Flash.
miroljub 9 hours ago [-]
I wouldn't call that inexpensive.
For comparison, I am currently at 6.6B tokens, 95% of monthly quota on a 10$ command code plan, mostly using DeepSeek flash 4.1, or some of the free models for easier tasks.
How fast is it compared with the other Chinese models?
ricardobeat 17 hours ago [-]
They both are in the 50-100 tok/s range. The Mimo v2.5 Pro Ultraspeed beta could reach 1000 tok/s, hoping they can do something similar for the new model, it was amazing.
wangxili1997 8 hours ago [-]
[flagged]
electroglyph 16 hours ago [-]
[flagged]
NuclearPM 16 hours ago [-]
Real?
electroglyph 13 hours ago [-]
mimo 2.5 has been a big underperformer since shortly after it's release imo. i cancelled my sub after the first month. purposefully using 2.5 right now is just handicapping yourself for no reason.
NuclearPM 13 hours ago [-]
I understand now. You used the wrong word.
yeeeloit 18 hours ago [-]
[flagged]
senordevnyc 18 hours ago [-]
Yeah, this Brazilian dude who has been a contributor here on HN longer than your anonymous account is shilling for a Chinese model company. Makes sense.
platinumrad 18 hours ago [-]
Are you accusing them of astroturfing? Why is it strange for someone to say something topical?
dr_dshiv 17 hours ago [-]
Well, if open source AI is dangerous (for OpenAI/Anthropic IPOs?), this is like watching a time bomb.
dzonga 16 hours ago [-]
the open burial started when zAI served their latest model on all Chinese chips.
now we r just noticing the grave getting dug deeper.
skybrian 15 hours ago [-]
For my own usage, Luna is cheap enough that I don't care if other models are cheaper. I'm interested if another model is in some way better and not too expensive.
rapind 15 hours ago [-]
Luna is great but makes a lot of mistakes at high and lower in my experience (large rust codebase). I use Luna Max for asynchronous subagent reviews and am very happy with its work, but it’s slow af.
epolanski 4 hours ago [-]
From my experience DS 4.1 flash is a much more capable model than luna.
ijidak 13 hours ago [-]
What plan are you on?
Trying to understand why users are using Luna when Sol seems essentially unlimited on the pro plan. Unless you have jobs running 24/7.
skybrian 4 hours ago [-]
Plus plan. For professional use, $100/month would be ok but it’s rather steep for hobbyist use.
teki_one 13 hours ago [-]
Sol is useless atm on the Plus plan, 1-2 questions 5-10m to get through the 5h allowance. (used to be good, can change any day)
Also worth taking a look at is the mimo harness. It's a fork of opencode with some new modes added for long horizon tasks. One of the better open harnesses out there at the moment.
Cookingboy 16 hours ago [-]
2.6-pro just reached 63.7% by step 10, it's on step 11 right now.
Even flash reached 60.7% by step 12, and it's on step 16 now.
This is so exciting lmao.
arcanemachiner 6 hours ago [-]
DeepSWE is saturated now IMO, and is basically worthless. Lots of new models get around 74%. Shame too, because it was a pretty decent benchmark for a few months there.
brookst 4 hours ago [-]
It is saturated, but that doesn’t mean worthless. Seeing 72% is low-signal, but 30% is still meaningful.
markasoftware 7 hours ago [-]
gemini 3.8 flash is also 74% and google just started letting all their engineers use claude...go figure
ehsankia 7 hours ago [-]
> and google just started letting all their engineers use claude
That's misleading.
1. Having different models available is useful for A/B testing and helping improve Gemini itself.
2. They have an enterprise offering for Antigravity (their agentic coding platform), and they need to test that it works well with non-Gemini models too.
passive 17 hours ago [-]
Neat! I've been trying out their next model for the last week, which I assume is a version of this, and it's been a good experience so far.
I had used 2.5-pro for a hefty chunk of development, and found it to work like a somewhat forgetful senior engineer who was new to my project. Very capable, would almost always choose a reasonable option, if not always the best one for the project, and not great at multi-tasking. Generally, made me comfortable not scrutinizing the code line-by-line, but still needed a bit of steering once projects got to a reasonable size.
The next model is a clear step up in the multi-tasking capability at least, with me very rarely having to steer the implementation of a well-defined issue. In terms of code, I found MiMo-V.2.5-pro to be extremely conservative, implementing minimal solutions. The next model seems a little bit more ambitious, in positive ways, making good guesses about gaps/next steps. It also seems to be a fair bit better at design, at least for the little bit I've done, it was good at translating my concepts to practical elements on screen, and cleaned things up nicely as I made suggestions.
krm01 19 hours ago [-]
This is pretty neat. What would be a good reason for the other Model providers to not do this?
kibae 19 hours ago [-]
Speculating here, but I assume researchers can make a reasonable estimate of the size of closed models based on factors like training time, training speed, and the number of tokens processed.
Also, Anthropic and OpenAI probably want to keep each other on their toes so they don’t end up on the wrong side of another Opus 4.6 / GPT-5.3-Codex situation, where one lab releases a model only for the other to drop a better one hours later.
jwpapi 19 hours ago [-]
I think first of all it’s not an obvious idea, also the marketing surplus for other providers is not as big for openai/anthropic as for xiaomi and last but not least I’m pretty sure you can withdraw methodology from here.
I’m saying who has a million dollars for me, so I can make my own model?
Bolwin 10 hours ago [-]
I don't really remember a situation, which of those models supposedly beat the other?
I still opus 4.6 though not for code
nikcub 12 hours ago [-]
this is remarkable transparency in an otherwise hyper competitive and secretive industry
liuliu 19 hours ago [-]
When you run benchmarks while training, isn't that the definition of contamination? Asking because I am not sure if this is normal in big labs now.
jampekka 19 hours ago [-]
Kinda yes. The benchmarks become part of the validation set, which means the models get slightly overfit to them if they are used as criteria for stopping the training. But a lot less compared to using them in the training data.
I'd guess everybody uses at least some benchmarks as stopping criteria, which is kinda sensible, but it also does induce some benchmaxxing, and explains partly why the newest models always tend to eke out in benchmarks.
Correct. If just stopping criteria, that is less contaminated. The question gets muddier once you also use it to determine hyperparameters during small-scale runs.
lucrbvi 19 hours ago [-]
They are using it to evaluate checkpoints during the training, they are probably not using the benchmarks for training the models. It's a common practice for big reinforcement learning runs.
sspiff 4 hours ago [-]
They run one step/iteration on an additional chunk of training data, then use the snapshot of the weights after that iteration in a separate validation benchmark while continuing to train on another chunk of data for the next iteration.
They result of the benchmark does not feed back into the training, it simply serves to provide a measurement of progression over time.
nodja 18 hours ago [-]
They exist to detect degradation. Datasets are not perfect and if a batch contains too much bad data it can ruin a run, also an opportunity to find bad data and improve the dataset filtering.
SwellJoe 19 hours ago [-]
You gotta have something to aim at. And, presumably, the benchmark is not part of the training data, it is the test against which the model is tested at each stage; is behavior moving in the right direction?
esafak 18 hours ago [-]
Not if you don't train against them.
kingstnap 18 hours ago [-]
It's implicitly trained against. There is like information leakage with researchers messing with the training parameters and checkpoints used.
It's not the direct feedback loop of RL but its not far.
brookst 4 hours ago [-]
It’s pretty far.
It’s the difference between “study law until you can pass any random bar exam” and “here are 200 legal questions and we’ll drill them, with me correcting and explaining when you get one wrong, until you can pass exactly these 200”.
Your right that tuning can aim for a benchmark, but it does not leak any information about the answers.
kingstnap 37 minutes ago [-]
The first implies generalization. It's not a test of generalization.
It's actually close to the second. "Here are 200 software questions, will drill you on *other stuff* until you can pass exactly these 200. If the other stuff isn't improving your scores we will change ratios of it till it does."
The reason it benchmaxes is that *other stuff* ends up looking more and more like SWE Bench without you realizing it.
ttul 15 hours ago [-]
$5 per second if my eyes don’t fool me. That’s ~$432K per day. Enough to rent 3,000 B300 nodes on Modal.
stymaar 7 hours ago [-]
Which isn't that much when you compare to the kind of DC that US actors are using.
lostmsu 21 minutes ago [-]
That's posttraining. Pretraining is the expensive part.
brookst 4 hours ago [-]
Do we know what kind of DC US actors are using specifically for training, versus inference and delivery?
fzysingularity 18 hours ago [-]
Very cool to see the openness here, and likely more like this will come from smaller startups where they win users on transparency.
ProfessorLayton 19 hours ago [-]
2.6 Pro: >started 2026-09-15 10:32 UTC
For some reason I thought training took much, much longer than what the progress bar suggests.
This is really neat, I'm currently using mimo 2.5 pro, and it's decent (or great given the price). Hopefully their next one is multimodal.
GaggiX 19 hours ago [-]
These are post-training reinforcement learning steps.
krackers 19 hours ago [-]
Yes, updated the submission title to say "post-training" to hopefully prevent further confusion
thehamkercat 19 hours ago [-]
This is crazy, but sadly anthropic/openai will never do this, what has happened to this world, where chinese companies are more open than US or even EU companies
b3lvedere 7 hours ago [-]
Is that a bad thing?
medlazik 19 hours ago [-]
Neoliberalism, that famously open and transparent economic ideology
atemerev 11 hours ago [-]
Ah, one Donald Trump, a famous neoliberal.
stymaar 7 hours ago [-]
Were Sam Altman and Dario Amodei different men before Trump was in charge?
brookst 4 hours ago [-]
To some degree, sure. Remember “open” AI?
I don’t think Trump changed them, but Trump is absolutely a symptom of larger social collapse in the US, and that collapse has affected Altman and Amodei. We’re not even pretending that truth matters or that the wealthy can ever suffer consequences, and those two seem quite liberated by that.
stymaar 3 hours ago [-]
> To some degree, sure. Remember “open” AI?
Was OpenAI open in any way under Biden Two years ago?
brookst 3 hours ago [-]
Did you miss the part where I literally said, and I quote “I don’t think Trump changed them”?
EMIRELADERO 10 minutes ago [-]
The closedness was always the plan.
"As we get closer to building AI, it will make sense to start being less open. The Open in OpenAI means that everyone should benefit from the fruits of AI after its built, but it's totally OK to not share the science (even though sharing everything is definitely the right strategy in the short and possibly medium term for recruitment purposes)."
-Ilya Sutskever (email to Elon musk and Sam Altman, 2016)
speedgoose 19 hours ago [-]
I didn't know 2 thirds of the training data would be source code.
jerrygenser 19 hours ago [-]
that is the the "data used to improve the model" when signing up for the subscription plans
leothetechguy 19 hours ago [-]
this is the rl run, not the pretraining run
ahmadyan 19 hours ago [-]
even in pre-training, usually 30%-50% is code these days.
leothetechguy 7 hours ago [-]
That would be far too high in my opinion. But happy if anybody can give insights from their own experience with pretraining runs.
ssn2000 12 hours ago [-]
Total run cost is $1.2M until now, what resources are they using to train their model? Wish they shared more details on that and what the MFU metrics are.
rao-v 14 hours ago [-]
I absolutely love that someone is doing this! Why isn’t IBM for Granite or Google for Gemini?
If you are going to develop a near frontier model, and you don’t think you have special sauce up your sleeve, why not making training runs and RL environment scores etc. visible to the world?
I’m genuinely learning quite a bit just from the dashboard
gtirloni 13 hours ago [-]
They think they have the special sauce. Even if they do, what would they get in return for doing that?
kkotak 13 hours ago [-]
Wouldn't us observing this break down the model superposition and make it dumber? :)
brcmthrowaway 12 hours ago [-]
Found the Dark Matter (2024) watcher
ernsheong 17 hours ago [-]
Mino 2.5 has been my workhorse for coder and tester agents (the ones planner agents delegate tasks to)
jstummbillig 11 hours ago [-]
Wow, spending money on training an almost-frontier-model is much more time intensive than I thought it was.
wolttam 19 hours ago [-]
Hah, it would be great to see more labs pick this up.
wg0 10 hours ago [-]
"Slow down this much openness in AI or we won't get our trillion dollars valuations!"
Google had this GPT long go and a wise man within Google noted:
"We don't have any maot neither does anyone else."
The AI bubble burst is guaranteed and is only delayed by IPOs.
user43928 7 hours ago [-]
Nothing is guaranteed.
Open models have not yet caught up with February's Mythos checkpoint.
Meanwhile OpenAI is solving millennium problems, and their compute is still fully utilized.
wg0 6 hours ago [-]
"Stealing millennium problems" would be more complete if not accurate description. And that 99.99% of the market is not interested in solving millennium problems is the other fact.
b3lvedere 7 hours ago [-]
I wonder what we will do with the discarded data centers and its hardware..
Ylpertnodi 5 hours ago [-]
Copper can be stole, but that's already in progress.
rozab 19 hours ago [-]
Why are they doing this? To try head off accusations about distillation?
bayindirh 19 hours ago [-]
Sometimes you're confident about what you're doing and show how you work to the world.
Keeping the garage door open, or at least making the door translucent. It's always cool.
jampekka 19 hours ago [-]
That China's official policy is now to prefer open models and open model development may be a part of it.
culi 18 hours ago [-]
BRICS just had a New Delhi meeting where Xi pushed a 5-point plan on AI cooperation that centered on open source models
Aboutplants 19 hours ago [-]
With that policy in place, labs might be incentivized to be creative in their openness. This being fun/free PR
brookst 4 hours ago [-]
I don’t see how it would head off such accusations. This is post-training, and even it’s data could be pulled from other models or run against other models in realtime. Not saying that’s the case, just that the dashboard does not disprove.
anemic 18 hours ago [-]
Bottom of the page says "Open is what we value."
hsbalanxvxjsmab 5 hours ago [-]
lol that’s rich
dr_kiszonka 15 hours ago [-]
Very curious that everyone here (so far) seems to assume this dashboard presents real data.
hsbalanxvxjsmab 13 hours ago [-]
Haha yeah pretty wild how easily you can see the data is fake by the repeating numbers (refresh the page the progress goes back in time constantly) + watch for restarts. They say they happen but 0 data correlates the log messages. Just a replay of old data or being fed by an llm so they convince people they are open
Bolwin 10 hours ago [-]
The intermediate tickers are fake but real data comes in and resets it. Its like a progress bar essentially. We don't call progress and bars fake
hsbalanxvxjsmab 5 hours ago [-]
I do when their fake like this site is. Insane people blindly believe this stuff
thenews 14 hours ago [-]
been using the 2.5 mimo for side projects, works amazing
monneyboi 8 hours ago [-]
Refreshing, now let's make this a default feature. I imagine a "Upcoming models" list with links to these kind of dashboards.
The Chinese labs are just making fun of the US labs at this point.
Where is the cool shit from the US labs?
culi 18 hours ago [-]
With other software, devs convince their managers of the importance of using open source stuff in their stack. With AI, it's usually managers choosing what models to use for the devs. The US labs don't need to give a damn how much devs like open source
reddec 3 hours ago [-]
I was dev, and now I am Senior level manager.
Open Weight models are current main focus for many companies with full alignment with top management for very simple reasons:
- stable and predictable performance (no pre-launch models degradation)
- ability to tune them for specific business cases (though still rare tbh)
- better (at least 60% Opus vs Kimi (real,3rd party)) and more competitive pricing
- flat pricing if tokenusage is big enough to justify renting GPU
- decent quality
- much higher guarantees that data will not be sent somewhere (assuming 3rd party inference providers)
- and cherry on top: flat and minimal pricing with absolute confidentiality using Alibaba Apsara stack of recently released AMD Instinct Coder box[1]
This isn't about liking open source. This is about the labs just being cool and doing cool shit instead of the opposite which is Anthropic where all they talking about is killing everyone and taking everyone's job.
dlisboa 15 hours ago [-]
These labs are still (for the time being) made of people, who reflect their lives onto the work.
The US population is much more pessimistic and doomsday driven these days, whereas the Chinese are more optimistic and future driven.
noir_lord 18 hours ago [-]
> The US labs don't need to give a damn how much devs like open source
In the short term, true.
In the long term, unknown but typically when you hold progress that way while other countries don't you at best end up becoming siloed while the rest of the world continues on without you.
hsbalanxvxjsmab 13 hours ago [-]
You mean all of the frontier models that the Chinese distillation clones are copying? Yeah kinda cool imo. If a dashboard showing training for a model that doesn't even come close to anything us labs have released in 6 months is "cool", then you're a loser
bicepjai 13 hours ago [-]
Hahaha. Is that Sam or Dario with throwaway account. This sounds like calling social security, a free handout. Who distills the distillaters? Get it?
hsbalanxvxjsmab 5 hours ago [-]
lol omg hahahahahah lol that’s so funny
impulser_ 11 hours ago [-]
Why the fuck would you or I care about that?
Anthropic and OpenAI literally stole from every human in history and youre out here complaining that the Chinese are distilling models and releasing them to the public?
Why do you care?
hsbalanxvxjsmab 5 hours ago [-]
Because without those labs to distill from the pathetic Chinese labs wouldn't have anything. Im not impressed by them copying US labs not sure why you are. But go off ccp bot
atemerev 11 hours ago [-]
No crying in the copyright casino.
dude250711 17 hours ago [-]
Distillation in real-time? Very interesting!
Cookingboy 16 hours ago [-]
That "training cost" is just live revenue count for Anthropic/OpenAI API calls!
/s
hsbalanxvxjsmab 13 hours ago [-]
This is so very clearly fake? See the message stating the flash 2.6 flash run was restarted and 0 graphs correlate that restart
Retro_Dev 12 hours ago [-]
A restart of the process does not necessarily mean reverting the model state. I don't know why you would even do that, because you'd lose all the progress you made.
hsbalanxvxjsmab 5 hours ago [-]
It said restarted step 15 5 mins ago and the progress showed they were working on step 16 for a day
levocardia 19 hours ago [-]
You'd think they would make it less obvious that they are running their whole operation with Claude
ricardobeat 17 hours ago [-]
If you're thinking of the UI style, definitely not Claude. It is incapable of writing a clear sentence like "what each step's samples are made of", would have used all-caps for everything, more padding and gradients.
conception 14 hours ago [-]
I hope this is /s because it’s very easy to get Claude to write sensibly. That’s why AI slop writing is so annoying because it’s so easy to avoid with any amount of effort at all.
ricardobeat 6 hours ago [-]
In my experience Opus and Sonnet 5 subtly ignore most instructions related to writing style, and continue to sound the same half of the time. Do you have a successful skill/prompt to share?
My use case is generally easy to read instructions for lay people of an international/ESL audience. Have it write it's whatever and then run that on it and it comes out... actually pretty good. Use it for emails, etc, when it doesn't need a personal touch and just needs to be clear.
These skills won't get you a snazzy blog post, but I imagine could be augmented to produce something significantly better than the incomprehensible non-sense that it spews out by default.
SwellJoe 19 hours ago [-]
It's not obvious to me. What's the tell?
jambutters 16 hours ago [-]
They'd be running in the red then cause they charge way less than Claude. Sorry but it just doesn't make logical sense. They have open source, papers, and self hosting too
iammrpayments 10 hours ago [-]
Did you come here to astroturf or are you a big fan of Claude
The cost is unbelievably low, and the quality of intelligence I get is equivalent to when I was working mostly with Anthropic models (late last year/early this year). I'm fully invested in MiMo and I'm very happy with it.
-- PS: I also check almost daily to see if other models are capable of doing such great work. And they do – DS4F is powerful and DS41 is impressive, GLM 5.3 Flash gets a job done well, etc. – but when I add cost of M-token in the ROI math, Jeez! MiMo is an order of magnitude better.
That's an eternity when it comes to coding models.
In my personal experience, we've had almost a step change every ~3 months this year, at least for bigger one-shot tasks. For example looking at Gemini Flash 3.0 vs 3.5 vs 3.8, it went 5% -> 30% -> 75% on DeepSWE, all since the start of the year.
Everything after it might be more "intelligent" but is super tuned around end-to-end task (and related benchmarks), not to act as an assistant.
Now it's *you* being the assistant, reviewer, etc.
In another one, Opus 4.6 level already solved 90% of my work-day tasks, so while better models have been instrumental into handling a higher % that does not mean that defaulting on cheaper models can't be good.
I run DS 4.1 flash daily, and then cross check with gpt-6-astra and I've nuked 90% of my AI monthly bill while having higher limits and better performance/intelligence than I did just at the beginning of this summer.
The availability of Mimo over Openrouter got however, much worse recently.
These days, there are more intelligent models like DS4.1, but Mimo is very obedient, so I plan things with another model and give the implementation to Mimo.
I’ll give Mimo a try.
UltraSpeed was absolutely awesome. I miss it.
DS 4.1 Flash is amazing. Well worth the extra cost.
And I am not a web developer! It's an extraordinary model.
(Mouse and keyboard required)
It’s good enough that I’m considering a second spark, or selling this and buying an M5 Ultra with 256GB for it
https://opencode.ai/docs/zen/#pricing
I always found that those Mimo models to be really good at tool calling and following instructions
API may be expensive, but I do 900m tokens (95% cached, ~0.4% output) on Z.ai's $18/mo coding plan with GLM 5.3 Flash.
For comparison, I am currently at 6.6B tokens, 95% of monthly quota on a 10$ command code plan, mostly using DeepSeek flash 4.1, or some of the free models for easier tasks.
now we r just noticing the grave getting dug deeper.
Trying to understand why users are using Luna when Sol seems essentially unlimited on the pro plan. Unless you have jobs running 24/7.
https://www.debtdefaultclock.us/
Fable scores 70%, Kimi K3 69%, Astra 74% (all on max effort).
https://deepswe.datacurve.ai/blog/deepswe-v1-1
Even flash reached 60.7% by step 12, and it's on step 16 now.
This is so exciting lmao.
That's misleading.
1. Having different models available is useful for A/B testing and helping improve Gemini itself.
2. They have an enterprise offering for Antigravity (their agentic coding platform), and they need to test that it works well with non-Gemini models too.
I had used 2.5-pro for a hefty chunk of development, and found it to work like a somewhat forgetful senior engineer who was new to my project. Very capable, would almost always choose a reasonable option, if not always the best one for the project, and not great at multi-tasking. Generally, made me comfortable not scrutinizing the code line-by-line, but still needed a bit of steering once projects got to a reasonable size.
The next model is a clear step up in the multi-tasking capability at least, with me very rarely having to steer the implementation of a well-defined issue. In terms of code, I found MiMo-V.2.5-pro to be extremely conservative, implementing minimal solutions. The next model seems a little bit more ambitious, in positive ways, making good guesses about gaps/next steps. It also seems to be a fair bit better at design, at least for the little bit I've done, it was good at translating my concepts to practical elements on screen, and cleaned things up nicely as I made suggestions.
Also, Anthropic and OpenAI probably want to keep each other on their toes so they don’t end up on the wrong side of another Opus 4.6 / GPT-5.3-Codex situation, where one lab releases a model only for the other to drop a better one hours later.
I’m saying who has a million dollars for me, so I can make my own model?
I still opus 4.6 though not for code
I'd guess everybody uses at least some benchmarks as stopping criteria, which is kinda sensible, but it also does induce some benchmaxxing, and explains partly why the newest models always tend to eke out in benchmarks.
https://en.wikipedia.org/wiki/Training,_validation,_and_test...
They result of the benchmark does not feed back into the training, it simply serves to provide a measurement of progression over time.
It's not the direct feedback loop of RL but its not far.
It’s the difference between “study law until you can pass any random bar exam” and “here are 200 legal questions and we’ll drill them, with me correcting and explaining when you get one wrong, until you can pass exactly these 200”.
Your right that tuning can aim for a benchmark, but it does not leak any information about the answers.
It's actually close to the second. "Here are 200 software questions, will drill you on *other stuff* until you can pass exactly these 200. If the other stuff isn't improving your scores we will change ratios of it till it does."
The reason it benchmaxes is that *other stuff* ends up looking more and more like SWE Bench without you realizing it.
For some reason I thought training took much, much longer than what the progress bar suggests.
This is really neat, I'm currently using mimo 2.5 pro, and it's decent (or great given the price). Hopefully their next one is multimodal.
I don’t think Trump changed them, but Trump is absolutely a symptom of larger social collapse in the US, and that collapse has affected Altman and Amodei. We’re not even pretending that truth matters or that the wealthy can ever suffer consequences, and those two seem quite liberated by that.
Was OpenAI open in any way under Biden Two years ago?
"As we get closer to building AI, it will make sense to start being less open. The Open in OpenAI means that everyone should benefit from the fruits of AI after its built, but it's totally OK to not share the science (even though sharing everything is definitely the right strategy in the short and possibly medium term for recruitment purposes)."
-Ilya Sutskever (email to Elon musk and Sam Altman, 2016)
If you are going to develop a near frontier model, and you don’t think you have special sauce up your sleeve, why not making training runs and RL environment scores etc. visible to the world?
I’m genuinely learning quite a bit just from the dashboard
Google had this GPT long go and a wise man within Google noted:
"We don't have any maot neither does anyone else."
The AI bubble burst is guaranteed and is only delayed by IPOs.
Open models have not yet caught up with February's Mythos checkpoint.
Meanwhile OpenAI is solving millennium problems, and their compute is still fully utilized.
Keeping the garage door open, or at least making the door translucent. It's always cool.
Where is the cool shit from the US labs?
[1] https://www.amd.com/en/ecosystem/oem/supermicro/amd-instinct...
The US population is much more pessimistic and doomsday driven these days, whereas the Chinese are more optimistic and future driven.
In the short term, true.
In the long term, unknown but typically when you hold progress that way while other countries don't you at best end up becoming siloed while the rest of the world continues on without you.
Anthropic and OpenAI literally stole from every human in history and youre out here complaining that the Chinese are distilling models and releasing them to the public?
Why do you care?
/s
My use case is generally easy to read instructions for lay people of an international/ESL audience. Have it write it's whatever and then run that on it and it comes out... actually pretty good. Use it for emails, etc, when it doesn't need a personal touch and just needs to be clear.
These skills won't get you a snazzy blog post, but I imagine could be augmented to produce something significantly better than the incomprehensible non-sense that it spews out by default.