The cost-per-task in the charts from the now-remove blog post put it more at Sol-level cost per task, however. It seems like the model is significantly more token efficient in the benchmarks
My assumption is in the long term that efficiency will break interpretability, which will lead to questionable alignment. As efficiency drives capitalism and evolution we'll run headlong at it and try to deal with the risk as a side effect.
My guess is for distillation, they need to forward the prompt to Anthropic to get the real Anthropic model's response so they can train their own models on it
Improvements in model performance aren't always strictly compute-constrained in a way that makes them reliant on Moore's Law. Open weight models-- in particular, from Chinese labs-- are optimizing model intelligence with less compute. They're "behind" frontier models by months, but as others have noted, it's possible to get Sonnet 4.5+ level performance at reduced cost, today, from open weight labs.
This has been a thing for years, and so much so, that there's an entire TV with a dedicated second screen that shows you ads underneath your main screen: https://www.telly.com
Wow. I guess I'm surprised it took this long for banner ads to reach TVs.
I think I have my next startup idea: a physical ad blocker for this thing. We could even have multiple styles: yellow sticky note, duct tape, painters tape.
And if you want cheaper ones, we can print our own ads on your ad blocker!
No one rich enough flying what the average person would consider a "private jet" or private plane would be flying VFR from uncontrolled airport to uncontrolled airport. The "ultra rich" are not puttering around in single-engine Cessnas
This is Open AI's fault (and literally every AI company is guilty of the same horrid naming schemes). Codex was an old model based on GPT-3, but then they reused the same name for both their Codex CLI and this Codex tool...
I mean, just look at the updates to their own blog post, I can see why people are confused.
Google just did it too. "Gemini Ultra" is both a model (https://deepmind.google/models/gemini/ultra/) and their new top-tier subscription plan (a la Open AI's Pro plan). Why is this so difficult?
Confusing people is the best way to get them to throw their hands up, stop thinking critically, and start paying. all businesses do this. Mega corps have resources to enforce clarity, but they dont because theyre stupid? Ill eat my words if thats the case....
Given that there's a dozen agentic coding IDEs, I only use Cursor because of the few features they have like auto-identification of the next cursor location (I find myself hitting tab-tab-tab-tab a lot, it speeds up repetitive edits). Are there any other IDEs that implement these QOL features, including Void (given it touts itself specifically as a Cursor alternative)?
I think QOL will shift away from your keyboard. Give Claude Code a try and you’ll understand what I mean. Developer UX will shift away from traditional IDEs. At this point I could use notepad for the the type of manual work I do vs how I orchestrate Claude Code.
The reason I have never bothered with Claude Code (or even other agentic tools), is that I still code mostly by hand.
When I am using LLMs, I know exactly what the code should be and just am using it as a way to produce it faster (my Cursor rules are extremely extensive and focused on my personal architecture and code style, and I share them across all my personal projects), rather than producing a whole feature. When I try and use just the agent in Cursor, it always needs significant modifications and reorganization to meet my standards, even with the extensive rules I have set up.
Cursor appeals to me because those QOL features don't take away the actual code writing part, but instead augment it and get rid of some of the tedium.