The standard for success is not zero deaths, but it should at least be a system no more dangerous than human drivers.
Generally speaking, the average human driver causes a “ding” every 100k miles, a insurance worthy event every 250k miles, a police worthy event every 500k miles, and a fatality every 60M miles. Assuming an average driving speed of 60 mph, which is obviously an overestimation, those are 1.6k, 4k, 8k, and 1M hours respectively.
Evaluating a L2 autonomy system is somewhat challenging since in truth you are evaluating a human+autonomous system, so you should expect the overall safety to increase unless the autonomy system is actively making things more dangerous. However, for a L4/L5 autonomy system, which is Tesla’s goal, we can reasonably approximate the quality by miles per intervention as an intervention is a failure of the autonomy.
As a reasonable baseline based on video evidence of FSD, we can assume that each intervention prevents at least a “ding”. Therefore, we can evaluate the quality of FSD relative to a minimal L4 system as the ratio of time per disengagement versus 1.6k hours. Using those same videos as evidence, we see the FSD system requiring intervention every few minutes, i.e. <10 minutes, on average. That is 1/10,000 of the minimal quality expected from a fully functional L4 system.
So, for Tesla to create a minimal acceptable L4 autonomy solution that would actually usher in a age of safe self driving systems, they need to improve the core functionality of their beta product by 10,000x. It is that gap between what they are delivering and what they still need to achieve that is outrageous.