You still need training data.
Even if you had a simulation set up for training, and enough compute on the vehicle to run something like MuZero real time inference for self driving, you would still run into the problem of humans setting up all the possible driving scenarios, which will likely leave some out and the model will never learn the right actions for those - meanwhile, to any human, the action would be very obvious. Over time you could probably get close enough to a very very small error rate, but you would still be hesitant to trust the system.
A true superhuman driving agent would most likely be able to take a picture of a scene, and then give a prediction of evolution of object position in 3 dimensions for a given time window, and give a confidence score on that. For example, if the single image is from few on a highway, it should be able to predict that cars are in fact moving, because the chance of cars standing still on a highway is very low.
And to train a model like that, you most likely need a base model that can "understand" physics.
Additional images from the past or the future would generate prediction that is more accurate, and also improve the model self assessed confidence. And then you would have some heuristic algorithm like MCTS based on confidence levels on the best course of action.