Yes, it's called statistics and probability theory.
Edit: Parent was edited, was previously (paraphrased)
> I'm guessing you have no technical understanding of how this works
My understanding of statistics is:
- I can halve the % insanity by adding another 100% of good labels.
- If I want to reduce the insanity of labels to 1/33th of ~33% I need to add another 3200% of good labels.
- If I want to reduce the insanity to 0% I need to balance the bad labels with an infinite amount of good labels.
Is there anything I'm missing entirely except probability theory? Is probability theory the answer or is there something else?
People who talk about the danger of humans driving cars always seem to talk about the raw numbers, because humans drive cars a lot and the raw numbers are rather large.
But when we talk about automated driving, it's in percentages, because it's not being done on the same scale.
So to compare apples to apples, you'd have to convert the number of fatalities to an accuracy percentage. Have you considered trying? There is certainly more than one way to do it, but it would greatly contribute to the discussion if you made some attempt.
Telsa's early results for their very limited "self-driving" technology has shown a huge reduction in accidents for any given period of time the vehicles are on the road.
Humans are much safer than people on average, when driving in conditions suitable for Autopilot.
This makes zero sense and isn't how "average" works. For the same 1000 hours on the road, a Tesla car with Autopilot will have fewer accidents than a car driven for 1000 hours by humans. This changes as driving conditions get worse, and humans outperform Autopilot.
The way "average" works is that you average over something - a population or set. It is very important to be clear about what that something is and whether it's appropriate.
Why do you believe that Autopilot outperforms humans in comparable conditions? If this is based on Tesla marketing, I'm extremely prejudiced against them, and assume out of hand that they simply aren't making the right comparison and don't care. However, if you think that is incorrect, you could elaborate on why you have the opinion you do.
> Why do you believe that Autopilot outperforms humans in comparable conditions?
Because they have the data that proves it?
> I'm extremely prejudiced against them
And I've chosen to take them at face value with a grain of salt, and to believe that for the data they've collected from the hundreds of thousands of Tesla's with millions of hours of data using Autopilot, it's fair to say they have a large enough sample to draw conclusions about the safety of their cars vs. any incident rates from pretty much any other distribution.
Does it make more sense as "Humans, when driving in conditions suitable for Autopilot, are much safer than people on average"?
"I've chosen to take them at face value with a grain of salt, and to believe that for the data they've collected from the hundreds of thousands of Tesla's with millions of hours of data using Autopilot, it's fair to say they have a large enough sample to draw conclusions about the safety of their cars vs. any incident rates from pretty much any other distribution"
You seem to be saying that if you have a lot of data it doesn't matter what you compare it to. That seems wrong to me. Also, I don't have this data, and you are not bothering to help me find it.
The immaterial distinction between "humans" and "people" still makes that sentence confusing. I take it that you mean "driving a mile (either as human or autopilot) is safer in conditions that are good for autopilot than driving a mile in average conditions"? Or more directly, isn't your question really "are the conditions the same for the averaged human drivers and the averaged autopilots"?.
My implication is that I severely doubt the conditions are the same, when somebody touts a comparison, and I would need clear and convincing evidence otherwise to change my mind. As well as strong evidence of good intent and trustworthiness by the source of the information.
It's not just about being intentionally deceitful, but about the fact that it's hard to do the right comparison, so people feel justified in giving up on it.
Would a statement like this be a surprise for you?
1. You can't have an infinite amount of good labels 2. Humans are in charge of labeling too.
The question is if you can reliably overcome the number of bad labels in your training set, so that 33% of bad labels equates to <33% "insanity" in the system.
How would a system reliably discredit missing labels while still learning from good labels? The simplest solution would be that system is able to spot the bad/missing labels itself with some certainty, but that seems like a catch 22.