That is what the comment is saying. Of course the vision stuff is done with machine learning - that is after all the state of the art. But that is a tiny part of the self-driving problem. So you can recognize pedestrians, other cars, lanes, signs, maybe even infer velocity and direction from samples over time. But then the high-level planning phase isn't typically a machine learning model, and so if you record all the state (Uber better do or that's a billion dollar lawsuit right there) you can go back and determine if the high-level logic was faulty, the environment was incomplete etc.