1. When the driver has full control of the vehicle
2. When the 'autopilot' is engaged and the driver is ready to intervene if necessary
So if the AI passes the safety test on Type 1 data, Tesla can promote it to being tested on Type 2. And if it passes that safety test it can be promoted to full autonomous control.
The 'autopilot' mode effectively does for Tesla what Google's test drivers do, but for free and on a much larger scale. Seems to me Tesla have a very strong hand here.
So if there's an accident, Tesla can check to see if the autopilot would/could have avoided it. If they can turn round to lawmakers and say that "X% of accidents could be avoided if hands-off autopilot was legal" it should help speed up the regulatory side of things.
"It would have avoided this accident" (by braking, steering). It can say nothing about the future, "... but it would still have been in a collision 0.42 seconds later".
For all the collisions it would have avoided, there's another subset where it would only have "delayed" the collision. But that won't be mentioned. Because it doesn't fit the narrative.
This is why "partial automation" for the initial data collection still produces valid results. You can replay data against updated models and do "what if" testing without actually sending the car back out on the road again.
Eventually, (and with a lot of handwaving) the difference converges towards 0 and the network gives the right answer.
In this case, what you need is mostly "what would most humans do?"
There would be things to refine about that (e.g. prevent speeding; analyse how humans reacted right before crashes etc. and improve on responses), but as a starting point it is immensely useful.
Which is to say, if Google wants self-driving Google Maps vans, then collecting data using Google Maps vans makes sense. But if they want general-purpose self-driving cars, then collecting data using Google Maps vans will only give them a very narrow set of data.
In the OP where I mention outfitting an example 2019 car, that car might be a current-model-year car with roughly similar vehicle dynamics to the proposed 2019 model. The only thing that they are paying particular attention to is the specific sensors in use and their placement on the vehicle. This apparently is critical to development of the driving model, the camera on the test vehicle must be in the precise position and direction and must react in the same way as the production model or the data is bordering on useless.
Now, with a full 3D pointcloud and enough sensor data you might be able to translate one recording into a lower-resolution resampling to model the production version, but I've not heard of anyone doing that. Test data is collected on the production sensor suite, no changes allowed.
Definitely two data sets that should be put together.
I'm assuming that once a Google Maps van has covered an area, it intentionally avoids that area until a significant time later.