Training the cars based on data collected from roads is heavily biased towards incident free conditions. This does not give any training or feedback on the rare occasions such as these. If there was a learning algorithm deciding what to do (assuming hand coded rules are brittle and hence one would want to learn handling these scenarios) then it perhaps has no training data.
Evaluating the cars based on incidents per million is fine but doesn't tell anything about how the incidents would have been handled if they had happened. There is no incentive for the learning algorithm to slow down the car to prevent fatality, if all it cares about is an incident happening and not the severity.
One possible solution, autonomous cars are trained in real life simulations (using realistic lighting conditions, dummies and what not) to be able to handle the rare incidents and they are also required to pass regulatory testing in similarly realistic conditions to test for their behavior in rare incidents, before they are allowed to drive on actual roads.