I see where you are coming from. But 6.1 Sol seems like a new frontier in pricing, not intelligence. I do think the deceleration stuff was mostly bluster, but I don't think this release in particular contradicts it too much.
I don't want to drag the discourse away from this achievement, but I hate how this is announced with blatant corporate advertising (our internal model, here are the benchmarks, gpt astra TM yours now for the low low price of £200pcm). I just didn't think Navier-Stokes falling would be sponsored by McDonald's.
Still. I am crying right now. Navier-Stokes is solved.
Agreed. Opus 5 is doing just fine, slightly better than 4.8. It's personality is insufferable, but I find myself catching fewer problems at code review. It generally understands my conventions and isn't so eager to accrue tech debt.
Both. Command the author to use the style, then command a reviewer to check it. Write one skill called review-prose with your rules, and another called write prose which tells the author they will be judged by review-prose, so you only write the rules once.
What I have found is that getting a model to rewrite a badly written passage is hard, because it seems to key off what it reads. It might swap some vocabulary around ok, but it doesn't fix structures very well. So getting it close to the preferred style in the first place is better.
To take this further, if you must fix existing bad prose, write a clean-prose skill which extracts the bare structure of the prose with none of the style, hands it to an author subagent who isn't poisoned with the original bad prose, then hands the output to a reviewer subagent.
Opus 5 writing is horrendous, so I have been experimenting with improving the output!
No. If the target of your "Bayesian inference" is whether the chain kik->gmail->ISP is reliable, then it isn't independent evidence. But that isn't the same as the target of inference in court, which is guilt or innocence, and obviously having kik is additional evidence for that. As I mentioned, the chain kik->gmail->ISP would not even be disputed in a run-of-the-mill accusation in the US, any more than DNA evidence gets scrutinized for lab mix-ups. You would need expensive attorneys and experts for that.
What you're describing is very close to what education theory calls "assimilation" and "accommodation". When we assimilate knowledge, it just fits into our schema of understanding. When we accommodate knowledge, we need to change our schema, and that is where the feeling of being challenged (and often the feeling of profundity) comes from.
For example, a child learns that "foreigner" means someone from outside their country. Then, when they're 11, they go on their first holiday abroad and realise "Wait! _I_ am a foreigner here!"
So, maybe one way to frame what you're saying, is that LLM output tends towards being easily assimilable.
The article presents hypothesis as fact with insufficient science. I have no problem discussing speculation, but if the author wants to promote their claims to any more than that I would want to see more journalism.
reply