Personally I haven’t had that problem. I can readily distinguish spam from non-spam by manual review ~99% of the time. If I could train or instruct an LLM to do the same as I do now, I would be happy. My current false-negative rate with Bayesian spamassassin is more like 50%.