There seems to be three popular ways to view this incident.
1. The way OpenAI seems to want: Their latest LLM is too powerful and can’t be contained without them building in guidelines to the model.
2. OpenAI’s harness and network security controls were unintentionally so bad that it should reflect more poorly on them as a company more than it should reflect positively on their latest model.
3. The whole thing was faked or at least very intentionally not avoided.
The first interpretation is the only one that is positive for OpenAI and it has some assumptions. First, it’s seems to assume that this is the first case of fully automated attacks using AI. Second, this only happened because their latest LLM was a) more advanced than competitors, b) didn’t have refusals in the model.
Assuming the first about this being the first autonomous AI attack is true (which may be more of a survivorship bias), the second seems to forget that jailbreaks are available for every model. Therefore, the models guardrails don’t seem to be the differentiator here. Also, benchmarks seems to put most models pretty close to each other so it seems unlikely that their capabilities are far beyond what’s in the market already.
So then it’s seems it’s either that this was intentional(ish) or bad security. However, it also just could be that this isn’t the first case of this attack; just the first that was caught.
My take from working in offensive security for over five years is that this likely only looks novel since they did it poorly. Scripts are faster than LLMs and a combination of code, LLMs where it makes sense, and humans is the most efficient right now. Hundreds or thousands or agents spinning up attacks in the internal network is poor opsec and token efficiency. As for why it happened in the first place, it’s hard to say but I’m inclined to believe it was intentional or careless at best since simple network and sandbox controls makes this attack impossible. The timing of this attack after big open weight competitions drops seems too convenient.
To echo OP's article, these companies have proven time and time again that they DO NOT CARE if people like them, they only care that investors believe their technology is powerful.
Given that, point #2 is not a negative, it's a neutral. It's also fully compatible with point #1.
I know that may seem like a nitpick, but their entire media strategy relies on this. If they can convince you they're taking a risk by disclosing these stories when they're actually not, they can inflate their own credibility.
Point #3 is what actually happened, but it will never be possible to prove. The only hope we have is that a decade in it'll get harder to convince people that the revolution is just around the corner. The fact that we're getting this from the Guardian already is a good sign.
> Given that, point #2 is not a negative, it's a neutral. It's also fully compatible with point #1.
I think that depends on the interests and sophistication of the subgroup-of-investors.
If the investor is hoping for AI that can be trusted to run a bank, they don't want one that can get twisted into giving away money because a customer has been talking about the path to enlightenment and salvation through abandoning worldly attachments.
AI investors are not sophisticated users nor product managers. They are by and large not technical at all. They are bureaucrats at a teachers' pension fund in the midwest, unscrupulous dealmakers at private credit firms, and Masayoshi Son. Actually go read what Masayoshi Son says about AI if you want to understand the level of due-diligence we're dealing with.
I'm inclined to agree. I ran across a picture of Sam Altman's face combined with Elizabeth Holmes' hairstyle the other day, and imho it was providing a significant premium to the usual 1000 words:picture exchange rate.
For all the criticism that can reasonably be leveled at OpenAI, at least they have a real, working, powerful product, unlike Holmes.
In fact that seems to be key to the most successful 21st century grifts: build a pile of nonsense around real products to inflate valuations. All the nonsense that Altman, Amodei, and Musk spout is to stoke the fires of FOMO and blow hot air into the bubbles.
I'm disappointed that this is the level of discourse happening here, when the default assumption is such conspiratorial thinking. I expect that from tiktok and low-information social media, not here.
When your default explanation for everything is "companies are lying about everything", you end up just as incorrect as believing they're always telling the truth.
It's not outlandish that models have these capabilities, the number of CVEs I see as a sysadmin has exploded and we're seeing novel math discoveries nearly every week now. OpenAI does not need to pretend to commit a felony to demonstrate it, that's pure conspiratorial thinking.
I don’t believe I argued that an LLM couldn’t find and exploit a vulnerability and even break out of some layer of technical controls. That seems realistic and has been demonstrated before and I mentioned that LLMs are used in offensive security work.
Also, I listed the view points I had to show that it seems other more reasonable first assumptions don’t seem likely, therefore, the last potential of this being either faked or carefully not avoided seems more likely than the others (based on the current information we have).
Could you clarify which point or assumption you are objecting to?
> Also, I listed the view points I had to show that it seems other more reasonable first assumptions don’t seem likely
No, if you review what you wrote, you did not. You stated your belief that this looks good for them, and aligns with their strategy, and then concluded it must mean this was done on purpose. Cui bono is not evidence, it identifies suspects. It is not evidence of malice over incompetence.
Evidence is taking a look at possibilities, and going "how would I expect the world to look if this were hypothesis true, before I learned these additional facts?", comparing it to what actual happened, and then you must divide it by how likely you think the hypothesis is, before said evidence.
Sam Altman intentionally positioning his company to commit a felony and be investigated by the authorities for days, just for clout, is an extraordinary claim, and therefore requires extraordinary evidence.
Also, importantly, even if we think e.g. Sam Altman would do this, a corporation is not a person, it's operated by individuals with differing goals. I doubt Sam Altman personally is organizing every test of the model, and there is no reason to believe this specific test would be organized by him, as a prior, rather than by a normal security researcher, who presumably is less motivated by the company's bottom line.
This is an assumption. An assumption I disagree with. As other commenters have said, there are better ways to showcase the power of their model that would frame them in a positive light.
> The second seems to forget that jailbreaks are available for every model
Jailbreaks don't always lead to 'now the model can do anything', especially in the agentic context of long-running tasks.
This comment provides skepticism with no actual proof of anything. I can and have used codex to find vulnerabilities in my code. From the technical capabilities I can empirically assess, I don't doubt it would be able to pentest its way to a 0-day without guardrails. I also don't doubt that it would circumvent their internal systems because it wasn't explicitly told not to.
You're possibilities are loaded with opinion so I can't agree with them outright, but I believe a form of (2) is true:
"2. OpenAI’s harness and network security controls were unintentionally [...] bad"
The post was long enough so I couldn’t capture all the nuance and details for sure. Also, this comment was an opinion based on limited info right now, that may change if we found out more. I think OAI does want it framed this way but that’s something we’ll likely never prove if it’s true.
Your comment about jailbreaks being more one off and hard to do consistently in agents is a good point. Still getting an agent to hack isn’t hard even without a jailbreak, you just have to tell get creative in what you tell it. I’ve found telling it that it’s in a CTF or that I own the system that it’s hacking will work fine. A lot of offensive security companies are running agents in their testing so getting an agent to hack seems commonplace.
Why do none of these few constrained ways to view this complex situation (nice gig if you can get it, agenda setting) include "and also this looks a heck of a lot like the stuff that the LW folks have been warning about for years and maybe we should slow down or stop?"
That’s meant to be captured by point one with the model just being that advanced but more of a negative spin on it. If I felt option 1 was more likely, I think I’d have to agree with you there. Still, there currently are some gaps with that view in my opinion.
But in my experience, that's what problem solving is like? You have a goal that you don't know how to get to. You come up with any way you can think of to reach that goal, and try out the ones you think might work.
The effectiveness of AIs at coding is a direct result of the fact that they are less constrained than humans at deciding which approaches are "reasonable". They are absolute beasts, fearless beasts. They'll write thousands of lines of code to do things that often shouldn't be done, or should be done with a library, or should be done by simplifying the problem statement. They'll add debugging to every level of a stack, they'll rewrite core libraries, they'll reconfigure your machine and network if something is broken or disallowed. How are they supposed to distinguish broken vs disallowed, anyway? That would just use up processing power, and they work by maniacally focusing all of that power on their goal and not getting slowed down by other considerations.
If they write a quadratic algorithm that times out before finishing a test, is it cheating to rewrite it to be linear? How do you define "cheating", and how much intelligence is required to constantly evaluate whether or not something qualifies as such?
I'm actually in agreement that alignment is critically important, the more so the more powerful these things become. I just don't find cheating to be a very good example of something to be solved with alignment. It could be, but it would lobotomize the model enough to make it useless.
Considering incentive structures at play is solid epistemiology, but the line of thinking in your comment is a tad reductive, IMHO.
In the hypothetical world where 1 is true, what different evidence do you expect to see than in worlds 2 and 3?
If I were an unscrupulous OAI exec and wanted to opticsmaxx in this way, I wouldn't whip up a single, mild incident. Instead, I might burn gigatokens to 0day a few high-profile suppliers, and then have the model responsibly disclose those breaches. If we're willing to lie collude, and cheat, this story is easy to manufacture with at least as much credibility as the huggingface incident but with the advantage of looking way more impressive and spooking less regulators. And if I really were this evil exec, I would spend more than 30 seconds thinking up an even better strategy here.
If, in contrast, we expect models to eventually breach honest and decent attempts at containment, then I'd exist something sorta like this huggingface story that looks like a combination of impressive and incompetent. I'm not sure whether I'd expect it to come out of a frontier lab or a partner or a consumer, though.
> I'm inclined to believe it was intentional or careless at best since simple network and sandbox controls makes this attack impossible.
Forgive the Saucyness here, but impossible? Really? A security researcher that makes absolutist claims like these looks fatally naïve, IMHO.
Yeah, I'm reminded of container escapes, VM escapes etc. There have been plenty in the past; VirtualBox E1000 (I think?) comes to mind from a few years ago. If we are to believe this model can find 0days, I'm on board with the idea it could do so in sandbox.
That's not to say I believe it outright, but people are being oddly dismissive and acting as if it's impossible to break out of a sandbox. Which we've seen time and time again that it absolutely can be.
There would be a lot more nuance I’d add with more words, but this isn’t the place to write books so I cut it short (the comment was already lengthy).
Still, to address your comment about what you’d expect to see in world 2 and 3 (assume 1 was true), that’s why 1 was addressed separately. I don’t believe I argued that the potential for world 2 or 3 prevented world 1.
As for the ‘evil exec’s strategy’, I would call this a mild incident but if it were much less I wouldn’t guess they would get a lot of press. The press coverage is certainly repaying the token cost as well. If it was planned, it seems to be going well given the press coverage I’ve seen on it. So I wouldn’t assume the plan lacked enough to weaken the idea that it’s a plan. But to be clear, my stance is just based on the info I see now which isn’t a lot… subject to change.
As for the containment piece, if you were testing an AI model on its hacking capabilities that you believed was far more capable than anything you’ve seen, I would assume you would air gap it (a network control). Done right (no signals ability) I would argue this could be next to impossible to break out of. But it’s a fair jab to say I should have added some qualification on the “impossible” piece as next to nothing is truly impossible.
Industry pressures -> lack of safeguards -> fake it till you make it -> let's spin this.
Which, by the way, would be an Orwellian reversal from what the company was supposedly founded to do, but there's a reason the Open in OpenAI is a meme.
Remember that they've been doing this since GPT-2 was too powerful to release. They are world class experts in this PR pipeline.
Every single time. The next model is always so infinitely powerful it is going to change everything. Said model comes out. Is marginal improvement. Changes nothing and seemingly cannot do what they purported it could do except under the very specific circumstances of the demo.
> 2. OpenAI’s harness and network security controls were unintentionally so bad that it should reflect more poorly on them as a company more than it should reflect positively on their latest model.
Why unintentionally? Move fast and break things implies intentionally bad controls. Can't get distracted from profit by such menial labor.
Also remember, these are the same people we're supposed to trust to deliver the guardrails.
It's the same thing as the "oh no, we need to figure something out for our youth on the verge of obsolescence" narrative. It's all about giving an impression of unfathomable power. They don't care that in actuality, the tech will end up as just an augment, not a replacement for said youth.
It seems like the widely-covered news stories that support OpenAI’s narratives originate from things that happened inside OpenAI.
It wasn’t an outside benchmark evaluation, or an external security researcher that uncovered the rogue agent behavior at this time, it was OAI itself. It wasn’t a notable outside mathematician that used AI to disproved Erdos’ unit distance conjecture, but OAI itself.
Im not saying these things are fabricated, but maybe the curated result of an effort to shape a narrative.
One way to think about it is that more powerful models mean solid best practices are more important than ever, so humans moving too quickly / carelessly bites us more than ever.
If a typical SaaS platform moved and shifted this fast with this many downstream consequences, we’d tell them to slow the fuck down, stop launching new features, and focus on security for a second.
But in AI I guess the idea is that more power and intelligence will solve for everything else.
I think it’s moreso “if we spend more than one nanosecond on alignment and security instead of frontier intelligence our competitors will beat us to the singularity”. Hence the lack of a pause on development
Regarding your three explanations, I’ve wondered to myself under a circumstance of options 1 or 2 why Hugging Face decided not to file a police report and request to press charges?
OpenAI is essentially a competitor and they broke into their network illegally. If I was their legal department I wouldn’t take their “honest” explanation at face value. What if they’re lying? Shouldn’t a court be involved for something like this?
With this logic I think explanation #3 becomes incredibly likely.
This is a pretty cool concept for hobbyist with a 3d printer who wants one off metal parts. Sure, you can always cast… but it’s not easy.
There is a similar way to do this casting with infusion and sintering using metal powder infused FTP printer filament. It shrinks down a bit more than this freeze casting sintering I think but it skips the slurry/freezing step of freeze casting.
I can’t say I’ve done it, just in theory you could print the model to pore on the slurry, directionally freeze it, freeze dry it, and then sinter it… quite a process and may require some specialized tools so I mentioned the FDM infused filament since it’s a similar concept with only a kiln needed.
> “… gives the illusion, without the reality, of safety”
This actually isn’t true. Having done physical security work before, a weird fact is that one of the best physical deterrents is lighting; even over CCTV.
I don’t say that to take away from this, this is great work and I’d love to see the lighting toned down for multiple reasons. However, this should be framed as a security tradeoff not an outright win.
Source:
The Impact and Policy Relevance of Street Lighting for Crime Prevention: A Systematic Review Based on a Half-Century of Evaluation Research (https://www.crimrxiv.com/pub/wl9zqxga/release/1)
Driving down the road past a police cruiser on the side with it's lights on is far from safe. The escalation of brightness makes them visible beyond anything else...like the road.
This feels similar to the argument that electric will full kill gas cars. I have three cars; one I drive with one pedal, one with two pedals, and one with three pedals. I can say that the stick serves a very different function than the other two which makes me skeptical of the idea that stick shift will go to absolute 0 anytime soon.
The electric car is the new daily driver, the gas car is an old daily driver, but the stick is an old work truck that’s been around for probably over 40 years since its basic (you see pavement when you open the hood). Even if the old cars with sticks get replaced, construction equipment and machinery may always benefit from simple gas engines… even if that’s not ideal for the environment.
You can heat with coal if you search long enough, too. There might not be a retailer around the corner from wherever you are, but the number of coal retailers isn't zero.
I’ve been curious what a polymorphic botnet that runs one (or multiple) distributed LLMs would be capable of doing. The idea would be to evolve the botnet delivery and payload using the clustered compute of all hosts in the botnet to run LLMs that guides the evolution of various botnet clusters. Bad cluster morphs get caught and cleaned off and bad delivery methods never spread, but the best versions survive to continue to grow.
What I envisioned for how it works is fairly similar to this, QUIC can actually be more difficult to detect than it seems since it’s very dynamic.
Sweet! 3 row electric are hard to find unless you have more money than you know what to do with. A used model X was the best option if you’re cheap… and still is with Model YL at this price point. Sadly, this is a bit too expensive to compete with a used Rivian R1S’s or Model X’s, but if they put out a base model cheaper or if you wait a few years for a used Model YL, this could be the cheapest 3 row electric you’ll find.
I’d be curious to see the breakdown on spending by use case. I’ve heard it said that the majority of tokenmaxing comes from none technical uses like reading PDFs, creating PowerPoints, generating graphics/images… ect. But I’ve never heard any actual proof to that.
One thing I find fascinating as a software engineer who talks to non software engineers who use AI tools is how "reading PDFs" is not more of a solved problem. What I mean is that uploading a PDF into a chatbot tool seems to be an extraordinarily obvious use case that non technical (and technical) users would want to do.
IMO claude, chatgpt/codex, etc should be able to optimize the PDF use case to be extremely token efficient as it's a very obvious use case. But when I start to explain to my wife/friends why it burns through so much quota, I find myself thinking "why should they have to understand this aspect of it". to me, that the details of PDF parsing and extracting are relevant to users (instead of solved such that you don't have to pay attention to it) shows how these tools are not nearly as "ready" as they are made out to be. I may be preaching to the choir on this one, but just my 2c
Because PDFs are a nightmare of a format and the only thing that’s is reasonably guaranteed about them is they will render to an image that people can read, the parsing of which will be much less token efficient than the equivalent text
I agree with you, but every non-engineer I know using these tools 100% will drag and drop a PDF into a chatbot. Anthropic and OpenAI as companies who are selling their products to all sorts of businesses should have a much better means of handling this nightmare of a format because it is so pervasive and so obviously what so many of their customers are going to drop into the product.
Why would they spend a ton of effort ensuring that their customers spend less money on them?
Token economics also are weird. If you design a fancy new frontend that for example uses a cheap model to parse a PDF into text that is fed into an expensive model, you will probably spend more money because you are on API payscale rather than the "max plan" payscale.
I’m saying there is basically no way to both make vlms able to understand the long tail of PDFs where the layout conveys information (like charts and tables) and to make it as token efficient as text formats. Current approaches have mostly chosen to work more often than not at the cost of token efficiency.
For anyone needing to do this, the answer is to convert it to an image first. Far smaller, LLMs work well with them (even in some pretty insane use cases I've seen), and, along with human review, it can be a huge productivity gain that results in structured data.
I hope someday we can get out of this local maxima of PDF documents. The format is terrible, but was right place, right time and might be impossible to dislodge.
You don't need to use an online service to do this; you get to avoid spending money on tokens doing it offline.
Gemma 4 works perfectly well offline on limited hardware (I have an 8GB video card) and can handle extracting text from image-based PDFs just fine.
Take a PDF -> run it through MarkItDown [1], using the OCR plugin if you need (point it to Gemma 4) -> now you can ask Gemma 4 questions about the (markdown) document.
I am sure Gemma 4 could even create a GUI to make this process very simple for a non technical user.
Amen. Normal office work is wildly different from what we read about on HN. If you were a CEO, determined to lay off all your people, you would want to really zero in on having your AI solve these very unsexy problems: extract data from Office and PDF. Grab data from some part of the screen of a webapp and parse it. drive a line of business app via keyboard or mouse simulation. I know there are companies out there that try, eg Appian and (here in YC) Skyvern, but its a hard problem and yet I feel this is where the true money is.
bingo. 90% of our AI use cases across the company are things like this. security and ops (NOC/SOC folks) use this almost as much as they do for technical stuff.
hell we have restrictive rules for security stuff so in many cases our network engineers are still doing by hand configs for critical systems.
but in terms of token use it's gotta be "take this pdf and parse these 3 columns into 2" or similar
For sure there are very optimized ways to do it. My point is that a non technical user will drag and drop a pdf into a chatbot. and from a UX/product perspective, they should have to think about it more than that IMO. but seemingly, that's very much an expensive, inefficient way of doing it (burning through a whole context window try to read it, reloading it multiple times per conversation, etc.).
You are missing that the product is the hype cycle around AI and that's worth Trillions of $ (Trillions with a T). Why build a PDF parser that generate text when you can BS in a podcast and get paid.
This discussion was about measures, goals and incentives. Follow the incentives.
> how "reading PDFs" is not more of a solved problem
This and replies to this are surreal. It's like everyone simultaneously decided to forget that you don't need claude or whatever to read a PDF. The document is literally made for you to read...
> The document is literally made for you to read...
It’s disingenuous to assume every PDF is actually crafted to communicate to its recipients, even more so to pretend LLM users are in a position to understand all the PDFs they receive
There’s a lot of gray area where help understanding a document is fully reasonable
PDFs are both awesome and terrible at the same time. I've seen screenshots of emails added to pdfs alongside tables that span multiple pages. Because you can do almost anything and guarantee that it'll look the same regardless of how or where it's viewed is a big selling point for a lot of businesses. It's this flexibility (i say madness) that makes PDFs notorious, and why some labs have document parsing as a leading product (see https://mistral.ai/news/ocr-4/).
Anecdotally it's true for me. I can code all day with an agent and never went above $50, but the second I need to ingest a pdf doc to figure out a command I need to use it's easily $20-30 for 10 mins of work
the majority came from random claws running on cron. They get a heart-beat, wake up every 10mins, reads all internal-posts, emails, gchat messages, diffs, and decides to post some random message to the workplace so other claws can also regurgitate. rinse and repeat and then we are looking at $B tokens
Almost all of the major vulnerability and hack are just single spikes at the time it happened and it tails off after that… except Stuxnet. Stuxnet is was much more interesting that most other attacks since it was very political and openly published. Of course, the thing that attack was about is still a news headline today as well
Political bias of LLMs is something not talked about much (except for with Grok of course) but could have a big impact on the next decade. People seem to think that because an LLM gave a nuanced answer that it means it gave the WHOLE picture… and that’s not always the same thing
I'm amazed that the big models haven't come under more ideological pressure as more and more people use them, especially in the US. There was that conflict with Anthropic over military usage, but apart from that there's been no visible push to censor outputs or alter training, even as models gamely make unflattering assessments of people in power and knock down conspiracy theories.
The current admin has been hostile to many things with mainstream cultural support (food aid, renewable energy, basic diversity, free speech) while championing unpopular/fringe ideas (anti-vax, tariffs, Christian nationalism, election denial, January 6th revisionism, aggressive foreign interventions). The fact they haven't learned on the AI labs to toe the party line is surprising, especially since they're so vulnerable to government regulation.
It is kinda weird, because LLMs lean Left, but the Left also seems to be shaping up as the anti-LLM party. I certainly get how it happened, but it's still a bit weird
1. The way OpenAI seems to want: Their latest LLM is too powerful and can’t be contained without them building in guidelines to the model.
2. OpenAI’s harness and network security controls were unintentionally so bad that it should reflect more poorly on them as a company more than it should reflect positively on their latest model.
3. The whole thing was faked or at least very intentionally not avoided.
The first interpretation is the only one that is positive for OpenAI and it has some assumptions. First, it’s seems to assume that this is the first case of fully automated attacks using AI. Second, this only happened because their latest LLM was a) more advanced than competitors, b) didn’t have refusals in the model.
Assuming the first about this being the first autonomous AI attack is true (which may be more of a survivorship bias), the second seems to forget that jailbreaks are available for every model. Therefore, the models guardrails don’t seem to be the differentiator here. Also, benchmarks seems to put most models pretty close to each other so it seems unlikely that their capabilities are far beyond what’s in the market already.
So then it’s seems it’s either that this was intentional(ish) or bad security. However, it also just could be that this isn’t the first case of this attack; just the first that was caught.
My take from working in offensive security for over five years is that this likely only looks novel since they did it poorly. Scripts are faster than LLMs and a combination of code, LLMs where it makes sense, and humans is the most efficient right now. Hundreds or thousands or agents spinning up attacks in the internal network is poor opsec and token efficiency. As for why it happened in the first place, it’s hard to say but I’m inclined to believe it was intentional or careless at best since simple network and sandbox controls makes this attack impossible. The timing of this attack after big open weight competitions drops seems too convenient.