We improved that by creating a skill that is precise about what we want our tests to look like, and telling the AI to always use the skill when writing tests.
It still writes a bit too many and uses a bit too much mocking, but the alternative of humans missing obvious use cases and just being lazy to test important cases is much, much worse.
Of course. LLMs only produce the average. Most of the time if you think you need to mock you should probably just do the real thing. It’s mostly an antipattern.