Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

That is a common misunderstanding. Even if no safeguards are in place, asking an LLM what its "system prompt" is does not guarantee it will accurately reproduce the same. LLMs are not databases. They don't have perfect recall. What they print when asked such a question may or may not be the actual system prompt, and there is no way to tell for sure.


I mean, you're not wrong, but you're also missing the point. We don't need "perfect recall" in this case. It's not difficult to get any of the ChatGPT models to divulge their knowledge cutoff date. It's also not hard to verify with a handful of crafted prompts.

I think we can reasonably conclude it's updated.


Should be pretty easy to test this with the API though right? I haven’t seen a quantitative test here but since I can provide my own arbitrary system prompts, whether I can hack and recover them reliably should let us infer a confidence level for recovering ChatGPT’s own system prompt.


Especially now if it's hoovered up endless blog posts about prompt hacking


If you get the same result over and over again it's more likely to be true


If you get the same result over and over again, it means the model is more overfit to a certain result. It does not mean the result is correct.


> model is more overfit to a certain result

From their communications, a massive amount of effort was put into making sure the model followed the system prompt. One might claim "overfit as a feature".


Thank you, this is one of the most understood 'facts', especially regarding "prompt hacking/jailbreaking"




Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: