Without an error bound, ML can’t be in charge of anything that could put human lives at risk.
This is also why I don’t understand all the hype about FSD / L5 autonomous driving. We don’t even know yet if such error bounds even exist, so we don’t even know if machine learning is even the right tool for FSD yet. All certification entities for control systems that put human lives at risk in aviation, automotive, etc. require those error bounds. So it actually doesn’t really matter if Tesla comes up with a “maybe L5” system, without right error bounds, their cars won’t be certified as L5 and drivers will need to keep hands on the steering wheel.
I know it sounds like a flippant question, but for certain applications, if we can get a model that's better than human, then it doesn't need to be perfect.
And they way we currently do this in all sorts of ways is to pair a human with a computer so that they each do what they're best at. It doesn't have to be about full automation.
(the difference from an actual software is that humans are based on some crazy nanotech from the future that nobody can completely control)
Excellent summary :D
Sure, the "legacy" intelligence / climate / food could also have extreme tail risks, it's just that it's been tested for 100s of millenia... whereas new technology might be better in the average case or even 99th percentile, but the 1% (or 0.0001%) is unknown and potentially much worse.
However, it seems to me that people resolve this more along ideological / political lines than with any kind of rational reasoning.
Climate change isn't a "tail risk". It is a hard wall our civilization is approaching fast. If we do not solve it, it will undo the conditions we depend on to live.
Edit to add: a better metric would be something like "billions of decisions made where human life was at stake".
Makes sense doesn't it?
Well, actually no. The Google cars drive the same route every day in Mountain View with no deviation, so those millions of miles are really the same 10 miles over and over. Even seen 3 in a row behind each other.
Fools the regulators, and apparently you, though.
I don't think its possible to fool the regulators, but even if it was possible there are dramatic consequences for lying. Volkswagon tricked regulators into thinking their car didn't emit as much as it really did in 2007-2015, and ended up having to pay 33.3 billion dollars in fines and returns[2]. Volkswagon stock went from a high of around 27 to 12. 5 years later, it's only recovered to around 16. The massive lawsuit that would be formed if Tesla's (or other automated car maker) car was found to be unsafe would likely be even larger. Furthermore, it would crush the self-driving car business semi-permanently. VW at least still gets to make environmentally safe cars, Tesla and Waymo's main value proposition is dead in the water.
Basically, even though there are incentives to lie and cheat, I think the incentives to try to make something that is safe is far, far greater. And while I'm sure there are Google cars driving the same route every day, I'm sure there are also tests being done in all sorts of conditions all over the planet. I think that practically, socially, and politically speaking, we can use billions of miles driven. The stakes are way too high to lie.
[1] https://en.wikipedia.org/wiki/List_of_self-driving_car_fatal... [2] https://en.wikipedia.org/wiki/Volkswagen_emissions_scandal
Humans don't have "error bounds" either, and you trust them just fine.
However, we do have inertia with humans running such situations, and until there is something provably/demonstrably better, I don't see the current situation changing.
This includes myself of course.
I don't think it's an obvious conclusion that error bounds aren't important for automation because they aren't calculable for a human. They are just very different beasts.
Let's do a thought experiment: let's say, we had a self-driving car that's verifiably 10x better than human on average, yet does not provide "error bars". I know we don't have one now, and unlikely to have one in the foreseeable future, but bear with me here, for the sake of argument.
Would you trust it, rather than a random human Uber driver?
FWIW, I'm amazed that driving cars manually is perceived as normal every time I drive one. I can easily accelerate 2+ metric tons of metal to 100+ MPH, and get distracted, launching this deadly projectile with me inside into oncoming traffic. Most roads have _no dividers_. This does happen from time to time, lots of people die. Nobody gives a shit.
Humans suck so bad at so many things that robots will be better than them at a lot of fairly unconstrained tasks in the next decade or two. And I'm pretty certain they won't have error bars while doing what they do. Humans don't.
Note that X doesn't need to be high, or higher than a human, for example a self-driving car with a 60% failure rate is fine as long as we know when and how it fails. It's this "when and how" that's completely missing from self driving cars.
And in machine learning in general we just don't have any guarantees that a model trained on some dataset will perform the same way in the real world.
I've dipped my toes into robotics a few times, and every time I end up being reminded of just how painful it is. Even when you're working in an idealized simulator, it's extremely easy to find a little edge case that causes completely bizarre behavior. And moving out of the simulator only makes things far, far worse.
It's really easy to forget about those sorts of details and brush it away as just things to be solved while developing the automation. But they don't just go away so easily. Error bounds are there to help manage these issues and ensure we know how to best use the automation.
In many instances I think people would tend to prefer the Uber driver in a moderate reading of your scenario. A human driver is likely to perform somewhat consistently and predictably. If they drift around the corner and leave a long skid mark in front of your house you can make some assumptions about how they are going to drive. If you get in and see them struggling to keep their eyes open, you can again make some assumptions about their performance. A well-rested and safe driver is extremely unlikely to suddenly throw themselves into on-coming traffic with no warning. You can refuse or stop the ride if you judge you are not safe.
Automation is a different beast. It's liable to fail in ways that a human driver would not. It may be performing wonderfully until something a human would not even notice happens, and which point it may indeed throw you into on-going traffic. For example, look at adversarial examples in deep learning. As a passenger in this case you don't have a way to judge your own moment-to-moment safety. Even if it is safer on average, the sheer unpredictability and the resulting stress is likely going to eat significantly into any gains.
1) Humans can estimate their own uncertainty. Ask a person to show how long a meter is, and they'll give you an estimate. Then ask them to show you the "error bounds", i.e. what they're "quite certain" the meter is longer than and shorter than. You are likely to get sensible bounds. Now, humans aren't amazing at this, but the brain does have capacity for estimating how uncertain it is.
2) No, you don't really trust humans. This very fact that humans are often imperfect in estimating their own uncertainty makes us very stupid sometimes. How many times you were sure you know something for a fact, for it to turn out to be completely false. This is why society tries to not put too much responsibility into the hands of a single person, or at least to provide help and/or safety mechanisms if that is the case.
We do have error bars on simple measurements already. They're right there in the data sheet for the sensor. What you're asking for are error bars on things several levels distant in the layers of abstraction. Humans suck at that. We only cope the same way machines do: through constant negative feedback.
The problem is that management doesn't want to hear about the worst case, because it looks bad for them politically.
Yes, but humans are already here and doing all those dangerous and difficult things. But if we're going to replace them with something, it makes sense to want to replace them with something that's better than them. I mean us.
Otherwise, we might as well stick with the humans. Especially since we already know how to make humans (whereas self-driving cars, not so much).
A great example of this is motion planning, where papers both on sample-based methods (such as SST), and on search based (descendants of the A* family) argue at length the theoretical optimality and convergence properties.
On another note, I think requiring more theoretical analysis as a guarantee of safety could partially be an AI-winter meme rather than practical solution. Point in case: do people run a quick check of aerodynamics maths before boarding a flight? No - they rely mostly on the engineering and regulatory process that gradually made passenger flights safer.
So I to explain my manager that we just cannot do better and know for sure that the robot is really in position X, especially with the limited sensing the project would afford. Sure, you can do the classic AGV thing and add magnetic markers everywhere. Or use more sensors to get higher accuracy, but none of those were popular options.
I work in the perception/localization domain and I am not aware of any large developments in that direction. I do know that there are certain ML based perception systems that got some levels of ASIL certification.
Usually mathematics and (proper) algorithms is the topic where everything is 100% (good quality actually finished work, not an early prototype released as final is assumed). It just may or may not to be fully relevant to our life.
You are off by a few nines.
Transportation machines are expected to have 5 or 6 of them. And those are the most dangerous kind we keep around. Everything else is more reliable.
Anyway, machine reliability is calculated over usage, not lifetime. It gets much larger numbers.