Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

The problem is model confusion. You ask models to get around security but also not to get around your security.

Models get confused by who said what - especially cluade models. They get confused by negation (don't do something versus do something). Compartmentalization is hard.

You can either solve compartmentalization completely, or just not tell the model to do things that must be compartmentalized at high stakes.



Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: