Also, Anthropic and OpenAI probably want to keep each other on their toes so they don’t end up on the wrong side of another Opus 4.6 / GPT-5.3-Codex situation, where one lab releases a model only for the other to drop a better one hours later.
I’m saying who has a million dollars for me, so I can make my own model?
I still opus 4.6 though not for code
I don’t think Trump changed them, but Trump is absolutely a symptom of larger social collapse in the US, and that collapse has affected Altman and Amodei. We’re not even pretending that truth matters or that the wealthy can ever suffer consequences, and those two seem quite liberated by that.
Keeping the garage door open, or at least making the door translucent. It's always cool.
I'd guess everybody uses at least some benchmarks as stopping criteria, which is kinda sensible, but it also does induce some benchmaxxing, and explains partly why the newest models always tend to eke out in benchmarks.
https://en.wikipedia.org/wiki/Training,_validation,_and_test...
It's not the direct feedback loop of RL but its not far.
It’s the difference between “study law until you can pass any random bar exam” and “here are 200 legal questions and we’ll drill them, with me correcting and explaining when you get one wrong, until you can pass exactly these 200”.
Your right that tuning can aim for a benchmark, but it does not leak any information about the answers.
They result of the benchmark does not feed back into the training, it simply serves to provide a measurement of progression over time.
For some reason I thought training took much, much longer than what the progress bar suggests.
This is really neat, I'm currently using mimo 2.5 pro, and it's decent (or great given the price). Hopefully their next one is multimodal.
The cost is unbelievably low, and the quality of intelligence I get is equivalent to when I was working mostly with Anthropic models (late last year/early this year). I'm fully invested in MiMo and I'm very happy with it.
-- PS: I also check almost daily to see if other models are capable of doing such great work. And they do – DS4F is powerful and DS41 is impressive, GLM 5.3 Flash gets a job done well, etc. – but when I add cost of M-token in the ROI math, Jeez! MiMo is an order of magnitude better.
I always found that those Mimo models to be really good at tool calling and following instructions
It’s good enough that I’m considering a second spark, or selling this and buying an M5 Ultra with 256GB for it
And I am not a web developer! It's an extraordinary model.
(Mouse and keyboard required)
I’ll give Mimo a try.
UltraSpeed was absolutely awesome. I miss it.
DS 4.1 Flash is amazing. Well worth the extra cost.
These days, there are more intelligent models like DS4.1, but Mimo is very obedient, so I plan things with another model and give the implementation to Mimo.
API may be expensive, but I do 900m tokens (95% cached, ~0.4% output) on Z.ai's $18/mo coding plan with GLM 5.3 Flash.
For comparison, I am currently at 6.6B tokens, 95% of monthly quota on a 10$ command code plan, mostly using DeepSeek flash 4.1, or some of the free models for easier tasks.
That's an eternity when it comes to coding models.
In my personal experience, we've had almost a step change every ~3 months this year, at least for bigger one-shot tasks. For example looking at Gemini Flash 3.0 vs 3.5 vs 3.8, it went 5% -> 30% -> 75% on DeepSWE, all since the start of the year.
Everything after it might be more "intelligent" but is super tuned around end-to-end task (and related benchmarks), not to act as an assistant.
Now it's *you* being the assistant, reviewer, etc.
In another one, Opus 4.6 level already solved 90% of my work-day tasks, so while better models have been instrumental into handling a higher % that does not mean that defaulting on cheaper models can't be good.
I run DS 4.1 flash daily, and then cross check with gpt-6-astra and I've nuked 90% of my AI monthly bill while having higher limits and better performance/intelligence than I did just at the beginning of this summer.
Where is the cool shit from the US labs?
In the short term, true.
In the long term, unknown but typically when you hold progress that way while other countries don't you at best end up becoming siloed while the rest of the world continues on without you.
The US population is much more pessimistic and doomsday driven these days, whereas the Chinese are more optimistic and future driven.
[1] https://www.amd.com/en/ecosystem/oem/supermicro/amd-instinct...
Anthropic and OpenAI literally stole from every human in history and youre out here complaining that the Chinese are distilling models and releasing them to the public?
Why do you care?
now we r just noticing the grave getting dug deeper.
Trying to understand why users are using Luna when Sol seems essentially unlimited on the pro plan. Unless you have jobs running 24/7.
I had used 2.5-pro for a hefty chunk of development, and found it to work like a somewhat forgetful senior engineer who was new to my project. Very capable, would almost always choose a reasonable option, if not always the best one for the project, and not great at multi-tasking. Generally, made me comfortable not scrutinizing the code line-by-line, but still needed a bit of steering once projects got to a reasonable size.
The next model is a clear step up in the multi-tasking capability at least, with me very rarely having to steer the implementation of a well-defined issue. In terms of code, I found MiMo-V.2.5-pro to be extremely conservative, implementing minimal solutions. The next model seems a little bit more ambitious, in positive ways, making good guesses about gaps/next steps. It also seems to be a fair bit better at design, at least for the little bit I've done, it was good at translating my concepts to practical elements on screen, and cleaned things up nicely as I made suggestions.
Fable scores 70%, Kimi K3 69%, Astra 74% (all on max effort).
Even flash reached 60.7% by step 12, and it's on step 16 now.
This is so exciting lmao.
That's misleading.
1. Having different models available is useful for A/B testing and helping improve Gemini itself.
2. They have an enterprise offering for Antigravity (their agentic coding platform), and they need to test that it works well with non-Gemini models too.
/s
If you are going to develop a near frontier model, and you don’t think you have special sauce up your sleeve, why not making training runs and RL environment scores etc. visible to the world?
I’m genuinely learning quite a bit just from the dashboard
Google had this GPT long go and a wise man within Google noted:
"We don't have any maot neither does anyone else."
The AI bubble burst is guaranteed and is only delayed by IPOs.
Open models have not yet caught up with February's Mythos checkpoint.
Meanwhile OpenAI is solving millennium problems, and their compute is still fully utilized.