Hacker Newsnew | past | comments | ask | show | jobs | submitlogin
Anthropic finds three hacking incidents similar to the HuggingFace attack (simonwillison.net)
8 points by Schlagbohrer 69 days ago | hide | past | favorite | 4 comments


Anthropic Needs to have the most intelligent and scary agents, so if OpenAI does something bad they need to prove their models can do even worse. Without that the whole valuation collapses. There’ll be more “our model outhacks others” for the next hype cycle.


I find the era of "apocalyptic alarmism as marketing" super strange and maybe not good either. How long until one of these firms intentionally "forgets" to airgap the model during a cybersecurity test, just to get a good headline about how powerful theirs is? Perverse incentives all around.


These were incidents where Claude had breached what was supposed to be a sandboxed exercise and hacked external organizations. Anthropic had no idea this had occurred (starting in April) until now, when they were prompted to check their logs due to the OpenAI vs Huggingface attack.


So, is this a competition? To see whose model can be jailbroken the most times and incite the highest level of public alarm?




Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: