Hacker Newsnew | past | comments | ask | show | jobs | submit | andrewmunsell's commentslogin

The cost-per-task in the charts from the now-remove blog post put it more at Sol-level cost per task, however. It seems like the model is significantly more token efficient in the benchmarks


that's all openai models but i'm very happy openai continues to focus on efficiency rather than reasoningtokenmaxxing


My assumption is in the long term that efficiency will break interpretability, which will lead to questionable alignment. As efficiency drives capitalism and evolution we'll run headlong at it and try to deal with the risk as a side effect.


either way we cannot see those thoughts anyway


It's also now live on Ollama Cloud as of a couple minutes ago


My guess is for distillation, they need to forward the prompt to Anthropic to get the real Anthropic model's response so they can train their own models on it


Improvements in model performance aren't always strictly compute-constrained in a way that makes them reliant on Moore's Law. Open weight models-- in particular, from Chinese labs-- are optimizing model intelligence with less compute. They're "behind" frontier models by months, but as others have noted, it's possible to get Sonnet 4.5+ level performance at reduced cost, today, from open weight labs.


This has been a thing for years, and so much so, that there's an entire TV with a dedicated second screen that shows you ads underneath your main screen: https://www.telly.com


Wow. I guess I'm surprised it took this long for banner ads to reach TVs.

I think I have my next startup idea: a physical ad blocker for this thing. We could even have multiple styles: yellow sticky note, duct tape, painters tape.

And if you want cheaper ones, we can print our own ads on your ad blocker!


"ERROR: PLEASE REMOVE AD BLOCKER TO CONTINUE VIEWING."

"TO RESUME PLAYBACK, PLEASE STAND UP AND SHOUT THE NAME OF THE BRAND ON THE LOWER SCREEN"

"THANK YOU. PURSUANT TO YOUR CONTRACT WITH US, A VIDEO OF YOUR PROMOTIONAL ENDORSEMENT WILL BE POSTED TO OUR SOCIAL MEDIA ACCOUNTS"


No one rich enough flying what the average person would consider a "private jet" or private plane would be flying VFR from uncontrolled airport to uncontrolled airport. The "ultra rich" are not puttering around in single-engine Cessnas


not going to argue with that, there were single engine comments on the thread though.


Yes, but I had to update the Codex CLI manually via NPM to see it. The VS Code extension auto-updated for me


> incorrect, its an o3 finetune.

This is Open AI's fault (and literally every AI company is guilty of the same horrid naming schemes). Codex was an old model based on GPT-3, but then they reused the same name for both their Codex CLI and this Codex tool...

I mean, just look at the updates to their own blog post, I can see why people are confused.

https://openai.com/index/openai-codex/

Edit:

Google just did it too. "Gemini Ultra" is both a model (https://deepmind.google/models/gemini/ultra/) and their new top-tier subscription plan (a la Open AI's Pro plan). Why is this so difficult?


Confusing people is the best way to get them to throw their hands up, stop thinking critically, and start paying. all businesses do this. Mega corps have resources to enforce clarity, but they dont because theyre stupid? Ill eat my words if thats the case....


They should use one of their LLMs to get some better naming schemes - seriously LLMs are pretty good at this set of task


They absolutely cannot be worse than the humans involved here. gpt-4o followed by a series of o models so you have gpt-4o and o4? Wonderful.


Given that there's a dozen agentic coding IDEs, I only use Cursor because of the few features they have like auto-identification of the next cursor location (I find myself hitting tab-tab-tab-tab a lot, it speeds up repetitive edits). Are there any other IDEs that implement these QOL features, including Void (given it touts itself specifically as a Cursor alternative)?


I think QOL will shift away from your keyboard. Give Claude Code a try and you’ll understand what I mean. Developer UX will shift away from traditional IDEs. At this point I could use notepad for the the type of manual work I do vs how I orchestrate Claude Code.


The reason I have never bothered with Claude Code (or even other agentic tools), is that I still code mostly by hand.

When I am using LLMs, I know exactly what the code should be and just am using it as a way to produce it faster (my Cursor rules are extremely extensive and focused on my personal architecture and code style, and I share them across all my personal projects), rather than producing a whole feature. When I try and use just the agent in Cursor, it always needs significant modifications and reorganization to meet my standards, even with the extensive rules I have set up.

Cursor appeals to me because those QOL features don't take away the actual code writing part, but instead augment it and get rid of some of the tedium.



Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: