Hacker Newsnew | past | comments | ask | show | jobs | submit | damsta's commentslogin

> all requests to the Pro model will be routed to V4.1 Flash and billed at Flash's price

While V4.1 Flash performance and cost looks promising this auto re-routing sounds concerning


I think they urgently need to free the compute power currently wasted on the big pro models.

I'd say it's because it's not available yet on subs

Why release it now instead waiting those few days until it is available for everybody?

Because they saw how much hype Glasswing was getting in April

From what I've seen it only made people mad, not hyped, so the person that thought it was a good idea miscalculated a bit. Now waiting for Anthropic's post about their usage promo or something similar to redirect people to them.

I was using Opus 5 and now I'm forced to switch to Sonnet 5 and bust the prompt cache which will eat a large percentage of my 5h usage...

I don't like any of current solutions when it comes to compaction. I'd love to have a way to say what exactly should be summarized, because most of the time I just need to compact some noisy MCP tool calls, test runs and things like that. Just let me pick what should be summarized and keep the rest as is.


You can do that in Pi!

> Extensions can intercept and customize both compaction and branch summarization

https://pi.dev/docs/latest/compaction

Just make an extension (or ask Pi to write an extension for itself) that intercepts compaction and leaves only what you want, or rewrites it in any other way. Should be just a few lines.


Sounds like you might like subagents. Agent > subagent receives agent context (presumably cached)->tool call->compact/summarise->return to main agent


I think /handoff on pi (or at least oh my pi) is what you are looking for


I mean, not to be flippant but can't you just prompt the agent to write a file as you're getting closer to the compaction limit? I tend to just go to roughly 50-70% context utilization and then tell the agent to summarize the conversation and save it to a file, manually /clear, then say let's continue that last conversation. You can inspect the summary first and make any changes.


The way I've done it has been when there's a longer task, I give it a markdown document that has a plan with numbered steps, and then for each step start a fresh context window and tell it to update that doc as it goes with any decisions taken or deviations from the plan, or other context needed for future steps. If it gets close to the end of the context I tell it to summarize the current state and anything a fresh context would need to know.


Solid 2 looks great! The async features are really interesting, I'll give them a try in my side projects. Congrats on the release!


So if Fast mode is 1.5x faster at 2x the price, will Ultrafast cost 20x as much? $100/$900 per 1M tokens?


> 3.7 Flash is available through the end of the year at an introductory price of $0.75/1M input tokens and $3.75/1M output tokens.

> Introductory pricing expires on December 31, 2026. Starting January 1, 2027, $1.50/1M input tokens and $7.50/1M output tokens will apply.


Clearly this model will be irrelevant by Jan. 2027, why would Google even bother to say this?


It's basically a "if we really have to support this for a long time, we want to be compensated for that" pricing strategy. It's about long term maintenance cost being greater _because_ it will be irrelevant.


Its probably just a corporate symptom, weird stuff like this happens in messy large orgs.


Maybe they know something we don't. What if all frontier lab do this? Maybe this is actual cost of running these llm.


They want to maintain the perception that Flash is worth $7.5/mot, so they can charge more for the next one.


Unfortunately with agents Meta can easily run hundreds of automated experiments on real browsers with each commit in the uBlock Origin's repo and each change in various filter lists to find clean workarounds and submit PRs with patches. You can try fighting this on the other side with similar approach, but why waste time and tokens, just stop using Facebook.


There's too much money in it for Meta to not throw a serious budget at it that will outweigh any effort the community can make.

Like you said, we shall all just stop using it.


I wonder what the price is going to be after the "significant increase" they plan to introduce in the near future.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: