The issue is once you solve the hard problems, the lower models start messing things up that were working and reverting all fixes for the hard problems. They'll just go off and do dumb stuff.
A car that feels safe to be driving at 200 mph is going to feel more comfortable at 60 mph, compared to one for which 60 mph is at the very limits of its abilities. Analogies only go so far so I'm not sure there's anything to be learned from that though.
Yes, even for crud apps there's a huge variability in how nicely you can make them, and how many mistakes or footguns models make along the way. I'm very curious to what level these people are building to.
In the Artificial Analysis index, MiMo 2.6 Pro is smarter than GPT-Sol 6.1 Low at the same cost, and only slightly dumber than Medium. MiMo 2.6 Flash is marginally cheaper and smarter than GPT-Luna 6 Max. (There is no GPT 6+ Terra, which would otherwise be in that range.) These are not negligible or trivial results.
It's very unlikely to have any meaningful effect. CXMT DRAM will also be hoovered up the moment it hits the market at the same inflated prices. Even if they had the capacity (they don't) why flood the market when you can just print money?
reply