Hacker Newsnew | past | comments | ask | show | jobs | submit | nullbyte's commentslogin

The intelligence difference between models like DS4.1 and Sol/Opus is NOT negligible.

The premium price is only worth fo the hardest problems.

For CRUD shoveling, models like DS4.1 are enough.

And the intelligence gap between cheap and premium is closing, as can be seen from the title of this post.


The issue is once you solve the hard problems, the lower models start messing things up that were working and reverting all fixes for the hard problems. They'll just go off and do dumb stuff.

If both DS4.1 and Opus can complete the tasks you throw at it at good enough quality the differences are negligible.

Who cares if your car can go 200mph if all you need is 60. If my requirement is 60mph, I want a faster 0-60, not a higher top speed.


A car that feels safe to be driving at 200 mph is going to feel more comfortable at 60 mph, compared to one for which 60 mph is at the very limits of its abilities. Analogies only go so far so I'm not sure there's anything to be learned from that though.

Yes, even for crud apps there's a huge variability in how nicely you can make them, and how many mistakes or footguns models make along the way. I'm very curious to what level these people are building to.

Opus 5.5: $4/$20

Deepseek 4.1: $0.02/$0.60

Just to illustrate how cheap the Corolla is in your analogy. Also Opus output would be $50 without competition.


I suspect most people in this position would be using the subscriptions, though.

A $20 Anthropic subscription is $500+ equivalent of API credit. You can build quite a lot on the $20 plan and get to use the best model.

Their API pricing has healthy margins built in.


Or just a cheaper way to reach 0-60

Specially if using the 200mph car when you need it is just a /model away.

In the Artificial Analysis index, MiMo 2.6 Pro is smarter than GPT-Sol 6.1 Low at the same cost, and only slightly dumber than Medium. MiMo 2.6 Flash is marginally cheaper and smarter than GPT-Luna 6 Max. (There is no GPT 6+ Terra, which would otherwise be in that range.) These are not negligible or trivial results.

Love this, very cool.

I used to use a service called ngrok for this, but it's nice that Cloudflare is offering one now.


CXMT is a memory fab that just opened in China. They achieved a 90% yield on DDR5. Hopefully that could ease up the supply squeeze.


CXMT has been in a game for a while, at least since 2025, because my SBC is using their LPDDR4 RAM.


It's very unlikely to have any meaningful effect. CXMT DRAM will also be hoovered up the moment it hits the market at the same inflated prices. Even if they had the capacity (they don't) why flood the market when you can just print money?


Even though cost-per-token is low, Deepseek v4 tends to burn an immense number of tokens to accomplish tasks.


It still ends up being one or two orders of magnitude cheaper per task on benchmarks.


I can go for days on end without topping up my deepseek account. When it can’t solve a problem I switch to GPT and have to top up in real time.


Great article


Face search? What do you mean?


sanest emacs user


Since Bitwarden is open source, can't somebody create a community-driven fork? Maybe a self hosted option?



Someone can, and did


"We also support inpainting, enabling targeted audio editing and the continuation of short recordings."

I didn't know there were models for that. Very cool!


PlayDiffusion is a notable one. But the state of the art is quickly evolving.


82.7% on Terminal Bench is crazy


Is it? There are 5 other models near ~80% and it was achieved in March... which in AI-world seems like a century ago.

https://www.tbench.ai/leaderboard/terminal-bench/2.0


those are not verified. I've tried forgecode and I cannot believe they didn't do something to influence the benchmarks


Yup, they were found to be sneaking the answer key using agents.md

https://debugml.github.io/cheating-agents/#sneaking-the-answ...


Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: