Hacker Newsnew | past | comments | ask | show | jobs | submit | ImprobableTruth's commentslogin

Why would you accept it when the benchmark's ranking is obviously nonsense. It literally has muse spark 1.3 above 6 astra, 5.6 sol and fable 5. Anyone who has played with any of these models for any amount of time would immediately realize that this is total bunk.

That's the ankle. The actual knee is hidden in the feathers of the body.

Its "pure capabilities" are definitely worse than Fable, but I find codex has a much more pleasant style and is in comparison much more generous with its limits.

Ugh, people are still saying the Codex limits are more generous. They're not, Claude's are over 2x higher, have been for months! [1] It's just that Claude uses far more tokens, 2-3x is common. Except sometimes GPT will use just as many or even go into a compact loop and then your quota is gone, little headroom for hard tasks.

[1] https://devforth.io/agents-for-code/?sortby=monthly-value And I can confirm the numbers, I subscribe to both and watch the numbers


That website seems to suggest that Opus 5 spends ~57 cents per task, while GPT 5.6 Sol spends ~49 cents per task? That ratio doesn't feel quite right to me. Artificial Analysis says Opus 5 High costs nearly ~3x as much as GPT 5.6 Sol High for a given task: https://artificialanalysis.ai/models/comparisons/claude-opus...

No no our coffee is not more expensive! The serving sizes are just smaller!

Yeah I subscribe to both and watch the numbers too, and it drives me nuts

Who wants to actually watch anyways, rather than worry about it my team just created our own harness that prioritize usage + intelligence and assigns work out (and records token usage..)

https://go7workhorse.com

Still beta, please try it and give me feedback.


I don't get it. It's the same result.

> people are still saying the Codex limits are more generous. They're not

They are if you follow Tibo on the resets.


Well at the end of the day, I can finish more work with the codex limits.

Codex (+Sol) feels a lot more human for sure. Fable 5 is so, so wordy.

It used to be true up to 2w ago, but with the new/reinstated 5h limits I wouldn't be so sure anymore...

that's news to me, I'm still getting weekly limits, no hourly limits.

Claude has a better 5hr limit?

It's probably GLM 5.3 flash, so weaker but cheaper.

With vision on top

These exist and are used. The issue is that because they're so much smaller, they're also much worse, so they tend to have lots of false positives while still being easy to circumvent.


The fundamental issue is that for these systems bigger is essentially always better (if affordable). So if we can squash something like Kimi K3 down to run on a 'normal' system, that just incentivizes devs to increase the model size until once again we're at the limit of what can be run.


This is not what the bitter lesson is about. It's not "don't develop better methods, just scale", it's that those methods which scale best win. LeCun's work is fundamentally about devising a method that scales better with data than LLMs. Agree with him or not about the feasibility of it, but this is fundamentally still a bitter lesson-pilled mindset.


However "world models" have been a thing for the entire history of computers. They have changed names over time: "rules engines", "expert systems", "semantic web", and so on and so forth.

And they have failed every single time.

The bitter lesson essay was written precisely to dismiss that approach, which used to dominate conferences and scientific publications of the era. A general learning system, given sufficient computation power, will always outperform specialized crafted systems in the long run.

Think of it like this: if a world model is a useful abstraction, the general learning system will create it by itself during its training, without us needing to implement it by hand. This is the bitter lesson. And it comes for us all.


This is some bizarre victim inversion. The providers of closed models are the ones who are trying to use regulation to stop their open model competition, not the other way around.


This is how it is with all of these guys, their only principle is "what's good for me", and will twist all narratives to fit it


That's why competition is good though


Now, what if you order a construction crew for building your house? What if you also hire an architect to design the plan? And an overseer who manages everything?

Surely at some point you would stop saying "I built this house", even if you ordered and financed it?


This is getting a bit off-topic, but there are certainly people who will still say that they built it. That is, in a nutshell, how people try to reconcile desert theory with the existence of the super-wealthy.


>I just did a ~6 month project in ~2 weeks using a frontier model.

Claims like this are hard for me to take seriously because 'good' models have been available since the start of the year. So, if they really 10x one's productivity, then people should be able to have gotten done 5 years worth of work since then, but I've never actually seen anybody show any project like this.


> 'good' models have been available since the start of the year

today: https://www.anthropic.com/news/redeploying-fable-5

35 days ago: https://www.anthropic.com/news/claude-opus-4-8

70 days ago: https://openai.com/index/introducing-gpt-5-5/ <-- first model I've found useful

77 days ago: https://www.anthropic.com/news/claude-opus-4-7

119 days ago: https://openai.com/index/introducing-gpt-5-4/

182 days ago: The start of the year


Opus 4.5/4.6 are what many people consider the first 'good' models and it's from last year/start of this year.

But fine, let's say everything before gpt 5.5 was unusable crap. Then there should still be projects that would normally have previously taken ~2 years done in just two months. Where are they?


I’m starting to see huge solo projects with 100 commits per day turning up on GitHub.

Last year these were unusable slop.

Now they’re “getting there”. Not quite as good as hand crafted code written by humans, but usable.


My guess at what's happening is that people are mostly using the tools on low impact or speculative projects. Notice that he said he wouldn't have attempted it without AI.

That's been my experience too. I had an idea that I didn't need so I hadn't bothered doing it, but AI made it easier to just have a go. I suspect people aren't using AI as much on their main profit-making projects (which also are going to be bigger, more complex and not greenfield - which is all harder for AI).

Also give it a chance - as you said "good" models have only been available very recently and you wouldn't expect everyone to start using them instantly.


Sure, but "10x faster, but only applies on small greenfield, throwaway projects" is a major caveat. In fact, there's a good chance this doesn't disprove the original blog post, you could be way faster on small projects but slower on 'real' projects.

>Also give it a chance - as you said "good" models have only been available very recently and you wouldn't expect everyone to start using them instantly.

But I'm not expecting everyone to have built something like that, but surely among millions of users someone should have, especially the people proclaiming insane productivity gains? There are no super impressive open source projects done using AI and all the companies boasting about how all their code is AI written now don't show much improvement either.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: