Or rather it's being.. Unsubstantiated. The Nerf conspiracy isn't that there have been a few harness and platform bugs leading to performance regressions, but that OpenAI/Anthropic have maliciously and unethically degraded their model performance post release to shed load and save money.
Yeah it's just inconceivable that companies whose entire business model started by engaging in wholesale for-profit theft and abuse of intellectual property would ever be so unethical as to try to lower their costs, especially just prior to an IPO.
It's at least highly implausible. Why would they engage in pretty uncontroversially illegal deception/fraud if they have so many other legal ways of gaming benchmarks, selling more tokens etc. available to them?
It's like arguing that your bank is scalping you by rounding down interest math on odd days of the month when they can just introduce a perfectly legal bullshit fee or otherwise change their terms to your disadvantage instead.
They wouldn't be breaking any laws at all. It'd be simply tuning their output as they see fit. And that's if you could actually prove it was even intentional, when in reality they have endless plausible deniability of 'oh it was just a technical glitch we've since corrected.'
Banks have far greater transparency and legal requirements. LLM companies are just delivering a black box that they have complete control over. And given the regular 'How's Claude doing this session?' stuff, it's almost certain that they're A-B testing various tweaks on a per session basis.
Again, these things are constantly measured. They sell HEAPS through their API access to enterprise consumers that expect a model to not be nerfed after it's released. And you bet many of those enterprises, some spending many millions each month, are measuring this shit.
So this is a case of extraordinary claims requiring extraordinary evidence.
And even though it's super straight forward to collect the evidence, there seems to be zero substantiating the vibe bro conspiracy theories.
My impression was that, at least with Anthropic, the point of my subsidized subscription is that I'll convince my employer to get an API account. If anything, the motive would instead be to ensh*ttify on the corporations already committed. Those of us inducted as boosters would get the fluffed up product.
Where people pointed out issues early and en masse, and Anthropic denied it was happening, gaslighted anyone claiming this was an issue, then begrudgingly admitted it was an issue, and then spent another two weeks "fixing it".
Anthropic is in a perpetual state of "oops, these 'bugs' degraded our model quality" and only admit the issues when it's immediately obvious and visibly affects a large number of customers.
Otherwise all open benchmarks can be (and are) gamed. And it's quite hard to judge the output of a non-determenistic black box that Anthropic (or OpenAI) constantly tweak.
If you're going to try to argue that companies doing things, completely legal mind you, to increase their profit margins is a conspiracy theory then you're not debating in good faith. Let alone when we're speaking of a subset of companies that were fundamentally built on wholesale unethical behavior carried out for profit. Let alone when we're speaking of companies who are all racing to IPO where short term results matter more than just about anything.
Another issue is also that the risk here is probably literally zero. Any evidence in support of such could easily be dismissed, with completely plausible deniability, as a short-lived technical glitch as opposed to intentional behavior.
> an explanation for an event or situation that claims a secret, powerful group is responsible for a hidden plot, rejecting the standard or official account
I'm sorry, but yeah. The official account is a harness regression and some platform bugs.
Where is the evidence they are underhandedly and unethically regressing their models to shed load and reduce costs? This is the conspiracy theory running rampant through the vibe boroughs; that they are bait-and-switching on model capabilities then "nerfing" them to save money and shed load. Where is the evidence?!
I'm not saying it's illegal, per say, so don't come at me with that straw man bull cock. This bro science conspiracy has been circulating for at least 2 years(I don't even know) and enterprises would certainly be pissed off if they were paying premium API prices for advertised and previously tested model capabilities that are suddenly under performing due to "nerfing" shenanigans.
I think the real story is just how easily people believe that purported conspiracy theory. It speaks to how little trust there is in these AI companies, and in Big Tech in general, that this "conspiracy" theory is perfectly plausible to lots of people
I think this is mostly a thing in the west. I was very surprised when I moved to Asia, and ended up emailing with lawyers back and forth for weeks, getting the right documents; the right advice, and then charged money the moment there actually was "work" to do, and I needed to hire them. It was shocking to me, because by that time I would have already spend $1000+ on a lawyer in Europe or the US.
I am no longer going to be impressed by new model releases unless they introduce an entirely new paradigm of interacting with them, that is going to make the benchmarks look like everything else isn't even 5% as capable.
The org I am at we have build an internal harness for coding, and we only hire devs that do work "AI-first". We did start hiring juniors as well as seniors; the process for either of them is the same. We drop them into any of our repos without any prior knowledge of what the project is about; and ask them to interact with the harness to build a larger feature that has been specified in project management. For the task, they can only use AI; there is no room for any manual coding at this particular org. We grade them more or less on their interaction and problem solving patterns, as well as asking us questions; and how they write prompts. The coding interview last 2 hours, and there are only 3 rounds of interviews. Before the coding interview, we do send out a brief of what it will entail.
Maybe this sounds a little strange, but this has worked exceptionally well, as we have now several juniors working with us, that do still get checked by seniors. The harness itself which holds several levels of standards, rules as well as quality gates, allows these juniors to ship out a huge amount of high quality code. We still reject about 80% of the people who interview with us, because even though they know how to code with AI, their AI interaction patterns just aren't on par with what we need. I really started to like this process. Also to clarify, the harness isn't some Coding agent setup, but rather a whole setup that works well with Codex, Claude or Cursor; its well maintained versioned and constantly improved upon, with a huge amount of automations, rules and hooks. The moment someone starts working on anything at all, this is immediately connected and tracked with project management.
That sounds like a really great idea. Would you be open to sharing more about how the tool is designed? Are there any open-source or commercial tools like that out there right now?
For some people it clearly works, for others it does not. I feel quite hopeless with a Claude subscription, but the $100 ChatGPT subscription has been a lifesaver for a lot of my work. I have very little complaints and couldn't imagine switching.
I am still not convinced there isn’t some secret basement in which each frontier lab is just orchestrating all of these agents to make their products appear much more intelligent than they are with all guard rails turned of and continuous human input.
Well let’s look at facts - provided enough compute and a goal, these system will be in a sort of loop trying out every single thing that’s in their system - they have encyclopedic knowledge and so it’s not unbelievable that a prompt which usually has a lot of implicit human rules in it can be misunderstood by AI and it just tries everything in its arsenal and we hear about the things which actually resulted in damage. I bet most of the time, they just spin in loops without achieving much if my experience with these LLMs is anything to go by. They have an important advantage in one area though, they know a lot and they can spin forget trying all sorts of combinations of things. The danger right now is probably cybersecurity, which is most likely because most orgs have historically underinvested in that area
Not the right comparison to make when comparing gold vs paper currency that can be printed at a whim. The comparison would hold if they pulled out physical fiat currency reserves.
reply