> seems at least plausible that the Israeli security services did in fact know about Oct 7th but chose not to act
Regardless of the facts in this specific event, you can make much the same argument about many of the major 21st century terror attacks - there were multiple warnings leading up to 9/11, and the FBI and CIA had eyes on several of the perpetrators before the attacks. They didn't "choose not to act" in so much as institutional disfunction prevented the dots from being connected or any preemptive action being taken, but from the outside it looks pretty much the same.
The coding side of this isn't even the problem - that a broken change like this made it to production is an operational failure.
Why didn't testing catch this? Why was there no canary? How come alarms didn't wake up an on call as soon as the call-volume on the backend dropped? Why was this update rolled out to every customer at once, instead of gradually?
I think verification is still a grey area in the agentic software dev.
QA and ppl testing it often resort to coding agents for the QA work, and we all know how good agents are at convincing themselves.
> How is it a "multiplier" rather than an "equalizer"?
Because without the responsible human engineer in the loop, it'll all gradually decay in a cascade of edge-cases. This happens with human written code as well (every "we'll replace this prototype before we ship" you've ever worked on), but with LLMs it happens at 10-100x the rate.
The other good(?) news on this front is that Windows 11 only supports chips with at least AVX2, so if you are willing to cut off everyone on win10 and earlier, you can stick to that
Pretty sure it only actually currently needs SSE4, since they started compiling it with POPCNT for x86. The official system requirement is Intel 8th gen CPU, but Intel made Kaby Lake and Coffee Lake CPUs that lack AVX.
These are always fun little what ifs. In 2012 I sold $50k in Amazon stock to pay off my student loans - a hilariously bad financial decision in retrospect, since the loans had capped interest rates, and that $50k of 2012 amazon stock would be worth over a million today...
I would really love if we brought back some colloquialisms in this field. Not that long ago most folks in tech would have had pretty blank looks on their faces when someone started talking about the "Pareto frontier"
> Is It better a model that takes me to 90% in 1 dollar or one that takes me to 95% in 2 dollars?
It's pretty important to understand if your own work domain is one where the last 5% matters. In a lot of day-to-day software engineering tasks, it doesn't, and one can get crazy mileage out of the cheaper models. OTOH, if you are performing novel research, that last 5% may be worth whatever it costs...
The 90% and 95% are against some blend of tasks meant to be broadly representative. A pricey model seldom fails a problem that cheap models do well, so there's stratification of tasks by difficulty. Someone doing novel research may be in the "hard" 15% of the blend, where P(solution) goes from one third to two thirds.
On the other hand, if it's cheap to tell whether you got a good solution, and you think the 90 and 95% apply to your task blend, then it's almost always worth trying the cheap model first.
My Kia Niro EV, which I'm otherwise quite happy with, has an occasional failure mode where it locks out the ignition for an arbitrary amount of time (generally 15 minutes). It's sometimes, but not exclusively, triggered by scheduled departure turning on the AC. Kia's response is a big shrug...
Think they're referring to the following, when Amodei was still working for OpenAI:
'OpenAI Feared “Optics,” Not the Law – “Dario Amodei, OpenAI’s then-Research Director, responded that ‘as a training set [LibGen is] a bit sketchier.’ [OpenAI researcher Sam] McCandlish explained: ‘I was just worried about optics – i.e. ‘openai uses copyrighted data from sketchy russian website’ showing up on [Hacker News] would be unfortunate.”'
I was curious what he was responding to. Per https://authorsguild.org/app/uploads/2026/09/Class-Plaintiff... it was "On July 19, 2019, McCandlish wrote in an OpenAI Slack channel: “We’re not sure if we’re going to release the Foresight LM Scaling paper publicly, but if we do we were thinking about removing all mentions of LibGen, since it's a bit of a sketchy data source." The paper may or may not be https://arxiv.org/abs/2001.08361 where they write "we also test on similarly-prepared samples of Books Corpus [ZKZ+15], Common Crawl [Fou], English Wikipedia, and a collection of publicly-available Internet Books."
Dario, if you’re reading this, I want you to know that Claude sucks now. Its output isn’t even English anymore, it’s just claudeslop.
You’ve actually ruined it so much that it has now polluted the Chinese models which are copying your work. So now all the models output claudeslop.
You’ve polluted all of the training data in the world. Now we’re never going to be able to train proper models, because everything has slop in it and all the sites have locked down their data to prevent future startups training.
I’m not the person you should say that to, I already don’t use them. My previous comment was specifically about “report that to their customer support”. Anthropic doesn’t have actual customer support
I would to order two soft taco supremes, a beef chalupa, and some cinnamon crispas. I'm sorry, what? Cinnamon crispas are discontinued? Since when? 1988! Unbelievable!
Regardless of the facts in this specific event, you can make much the same argument about many of the major 21st century terror attacks - there were multiple warnings leading up to 9/11, and the FBI and CIA had eyes on several of the perpetrators before the attacks. They didn't "choose not to act" in so much as institutional disfunction prevented the dots from being connected or any preemptive action being taken, but from the outside it looks pretty much the same.
reply