Why would you accept it when the benchmark's ranking is obviously nonsense.
It literally has muse spark 1.3 above 6 astra, 5.6 sol and fable 5. Anyone who has played with any of these models for any amount of time would immediately realize that this is total bunk.
Its "pure capabilities" are definitely worse than Fable, but I find codex has a much more pleasant style and is in comparison much more generous with its limits.
Ugh, people are still saying the Codex limits are more generous. They're not, Claude's are over 2x higher, have been for months! [1] It's just that Claude uses far more tokens, 2-3x is common. Except sometimes GPT will use just as many or even go into a compact loop and then your quota is gone, little headroom for hard tasks.
That website seems to suggest that Opus 5 spends ~57 cents per task, while GPT 5.6 Sol spends ~49 cents per task? That ratio doesn't feel quite right to me. Artificial Analysis says Opus 5 High costs nearly ~3x as much as GPT 5.6 Sol High for a given task: https://artificialanalysis.ai/models/comparisons/claude-opus...
Yeah I subscribe to both and watch the numbers too, and it drives me nuts
Who wants to actually watch anyways, rather than worry about it my team just created our own harness that prioritize usage + intelligence and assigns work out (and records token usage..)
These exist and are used. The issue is that because they're so much smaller, they're also much worse, so they tend to have lots of false positives while still being easy to circumvent.
The fundamental issue is that for these systems bigger is essentially always better (if affordable). So if we can squash something like Kimi K3 down to run on a 'normal' system, that just incentivizes devs to increase the model size until once again we're at the limit of what can be run.
This is not what the bitter lesson is about. It's not "don't develop better methods, just scale", it's that those methods which scale best win. LeCun's work is fundamentally about devising a method that scales better with data than LLMs. Agree with him or not about the feasibility of it, but this is fundamentally still a bitter lesson-pilled mindset.
However "world models" have been a thing for the entire history of computers. They have changed names over time: "rules engines", "expert systems", "semantic web", and so on and so forth.
And they have failed every single time.
The bitter lesson essay was written precisely to dismiss that approach, which used to dominate conferences and scientific publications of the era. A general learning system, given sufficient computation power, will always outperform specialized crafted systems in the long run.
Think of it like this: if a world model is a useful abstraction, the general learning system will create it by itself during its training, without us needing to implement it by hand. This is the bitter lesson. And it comes for us all.
This is some bizarre victim inversion. The providers of closed models are the ones who are trying to use regulation to stop their open model competition, not the other way around.
Now, what if you order a construction crew for building your house? What if you also hire an architect to design the plan? And an overseer who manages everything?
Surely at some point you would stop saying "I built this house", even if you ordered and financed it?
This is getting a bit off-topic, but there are certainly people who will still say that they built it. That is, in a nutshell, how people try to reconcile desert theory with the existence of the super-wealthy.
>I just did a ~6 month project in ~2 weeks using a frontier model.
Claims like this are hard for me to take seriously because 'good' models have been available since the start of the year. So, if they really 10x one's productivity, then people should be able to have gotten done 5 years worth of work since then, but I've never actually seen anybody show any project like this.
Opus 4.5/4.6 are what many people consider the first 'good' models and it's from last year/start of this year.
But fine, let's say everything before gpt 5.5 was unusable crap. Then there should still be projects that would normally have previously taken ~2 years done in just two months. Where are they?
My guess at what's happening is that people are mostly using the tools on low impact or speculative projects. Notice that he said he wouldn't have attempted it without AI.
That's been my experience too. I had an idea that I didn't need so I hadn't bothered doing it, but AI made it easier to just have a go. I suspect people aren't using AI as much on their main profit-making projects (which also are going to be bigger, more complex and not greenfield - which is all harder for AI).
Also give it a chance - as you said "good" models have only been available very recently and you wouldn't expect everyone to start using them instantly.
Sure, but "10x faster, but only applies on small greenfield, throwaway projects" is a major caveat. In fact, there's a good chance this doesn't disprove the original blog post, you could be way faster on small projects but slower on 'real' projects.
>Also give it a chance - as you said "good" models have only been available very recently and you wouldn't expect everyone to start using them instantly.
But I'm not expecting everyone to have built something like that, but surely among millions of users someone should have, especially the people proclaiming insane productivity gains? There are no super impressive open source projects done using AI and all the companies boasting about how all their code is AI written now don't show much improvement either.
reply