Without a good way of getting it into the dataset via simulation, or a more speculative approach (world models?), edge cases and unusual circumstances could cause failures.
I think my broader point is people don't do so well in unusual circumstances either: blizzards, heavy rain, dust storms, etc. They'll hydroplane, drive into stopped traffic, etc. We need to decide if we'll hold self-driving cars to some unreasonable standard of perfection or accept them once they are X safer than a human benchmark, even if they still have Y rate of failure per million miles.