People drive into deep water all the time - some states specifically have laws making them financially liable for the cost of rescue because it’s such a stupid thing to do. But still, they do it.
I think the fact that there are laws about it is not good evidence that it should be in the training data. Laws often cover weird edge cases, and if an edge case happens 50 times in 100 years, there's likely a law covering it. At the same time, that's probably not enough occurrences for it to naturally end up in a dataset -- the edge case would probably need to be intentionally sought out. I'm not saying that driving-in-deep water only happens 50 times in 100 years, it's certainly more common than that, I'm just saying that despite laws on the topic, it may still be too rare to be well represented in training data. For example, in real life I've only seen a car drive into deep water once. Even if we include recordings that I've seen, that would maybe bring it up to 20?
Without a good way of getting it into the dataset via simulation, or a more speculative approach (world models?), edge cases and unusual circumstances could cause failures.
I think my broader point is people don't do so well in unusual circumstances either: blizzards, heavy rain, dust storms, etc. They'll hydroplane, drive into stopped traffic, etc. We need to decide if we'll hold self-driving cars to some unreasonable standard of perfection or accept them once they are X safer than a human benchmark, even if they still have Y rate of failure per million miles.