Well, if all you have is machine learning, you're going to need a lot of trials. Hence the need for repeatable simulations.
From the article:
"Alternatively, one could follow the Tesla Autopilot approach and deploy their research code in “shadow mode” across a fleet of robots in the real world, where the model only makes predictions but does not make control decisions."
Does Tesla really do that? Has anyone decoded what they're uploading? How much upload bandwidth does each car use? Or is this just hype?