That is a common misunderstanding. Even if no safeguards are in place, asking an LLM what its "system prompt" is does not guarantee it will accurately reproduce the same. LLMs are not databases. They don't have perfect recall. What they print when asked such a question may or may not be the actual system prompt, and there is no way to tell for sure.
I mean, you're not wrong, but you're also missing the point. We don't need "perfect recall" in this case. It's not difficult to get any of the ChatGPT models to divulge their knowledge cutoff date. It's also not hard to verify with a handful of crafted prompts.
Should be pretty easy to test this with the API though right? I haven’t seen a quantitative test here but since I can provide my own arbitrary system prompts, whether I can hack and recover them reliably should let us infer a confidence level for recovering ChatGPT’s own system prompt.
From their communications, a massive amount of effort was put into making sure the model followed the system prompt. One might claim "overfit as a feature".