You really want a sensor fusion strategy for devices making life-or-death decisions on your behalf.
You really want a sensor fusion strategy for devices making life-or-death decisions on your behalf.
If you somehow do sensor fusion perfectly and the models never disagree, then why even have the LiDAR? At that point you’ve solved vision using cameras.
It did get me thinking about different senses and how I prioritize them. Given that humans strongest sense is visual, it's interesting to me that the priorities of what to trust seem opposite of my expectations. It seems to me that visual signals are seemingly least "trusted" compared to the other senses. As if the logic goes, "my sense of smell is so bad that if it detects danger, it must be very dangerous".
Similarly, I'm sure dogs can smell rotting flesh long before meat is unsafe.
In the context of vehicular autonomy it's more complicated than that because for example a radar sensor picking up an overhead sign or a truck parked on the shoulder as if it's an obstruction on the road is something you want to ignore when you're going 80 MPH on the highway rather than slamming on the brakes.
When you're the only vehicle on the road, stopping if anything goes wrong is always the safest idea. When you're one of hundreds of vehicles in a high speed flow of traffic stopping would put you and everyone else on the road at significantly greater risk.
I can almost see the objection that multiple cameras are as good as one camera+lidar, but I think it's a mistake to trust any system that can't check itself for consistency across multiple bands. It doesn't take a very good radar to keep you from ramming a fire truck. In fact, whatever runs the cruise control's distance sensor should have been enough to prevent a bunch of the Tesla oopsies reported in the press. When tackling one of the hardest engineering problems faced by humankind, it seems stupid not to take advantage of all the data you can get.
Each sensor is given an uncertainty model, that is for example a stochastic model by adding e.g. Gaussian noise to the "true" value it is measuring. Further, you have another model that describes the dynamics, e.g. equations from physics can tell you where the car will go if you know the current speed, position, etc..
1. The kalman filter computes a probabilistic prediction of what it thinks is happening by using the dynamics model. That is, based on what it knows so far, where will the car (probably) be when the next measurement comes up?
2. When measurements from various sensors come in, the Kalman filter uses Bayes's theorem to compute a mean (posterior), in which each measurement is weighted by the probability that the measured value is correct (using the uncertainty models; "correct" here means "in agreement with the prediction"). In other words, sensor that are inaccurate (large variance) will be considered less in the computation of the mean, while more accurate sensors are given more importance.
Once the mean of the measured quantities are computed they are used again in step 1 and the whole thing is repeated. As you can see, that disagreements are justified by the inaccuracies of the sensors, and the process of performing a probabilistic weighted average solves the problem. For the Kalman filter in particular, it can be be show (mathematically proved) that this process minimizes the variance (uncertainty) of the measured quantities (which btw. is an amazing result if you think about it).
A simple example of why this gets complicated: if I have a point in one camera and a point in another camera and I know they correspond to the same real-world spatial point, I can calculate some distance, and the statistics of that calculation can be captured by a halfway-reasonable error model. But how did I know they corresponded to the same point in the first place? Well, because they look the same according to some image feature... or because some deep neural network told me so.. etc. There just aren't very good ways to model just how haywire ^that^ process can go. So at the end of the day, once you let this evil into your perception system, using statistics to blend your sensors together is undermined, and all of your precious covariances just turn into tuning knobs you can twiddle.
The dirty secret is that almost all robotic perception systems are hiding unprincipled, un-modelled heuristics in the data association process. This is kicked under the rug because it doesn't really fit into traditional estimation theoretic frameworks. In a lot of papers you'll see academics push it aside by just calling that the "front-end", which they brush aside as a little widget you put on the front. If you're lucky they'll do ablations across a couple different options.
Of course, this is just one level deeper down the iceberg. It goes far deeper. Even if you could model the statistics of a depth camera well, the statistics of "what are all the objects in your scene about to do" is another couple of orders of magnitude more un-modellable. Often engineers will do something like attach a "constant-velocity" model to the agents in a scene. Imagine trying to bin all of the reasons you might stop walking in a straight line into a bubble that describes how "noisy" that picture of the world is! Now you can begin to appreciate just how hopeless it is to explicitly model uncertainty in the world around us.
You can start to add unjustified assumptions and that'll make the world model weaker, yes. But starting with pretty basic assumptions like "you can segment an object from an series of images because each solid object will move on its own trajectory", or even more basic like "objects have edges" and then a few dozen samples per second, and suddenly you have a fairly robust way to detect things.
Same for predicting where something will go. If you can estimate an objects current velocity, acceleration, and jerk with reasonable precision, you don't really need a highly predictive heuristic for the world model.
For decision making you need more robust heuristics, like "the car to my left has right of way at the stop sign", but you don't need that level of heuristic to identify that there is a car and that it isn't part of the pavement and that it is currently sitting still.
However, you should also consider that working with heuristics is pretty much all engineers do. We ought to solve problems, even if the theory is not there yet. A great example is how we got air travel long before we had any real understanding of the fluid dynamics happening around the fuselage (as it was generally computationally intractable). So, sometimes a simple epipolar camera model with noisy clouds around the subjects is sufficiently accurate for the task. The real problem IMHO is that the degree to which these rudimentary approximations are tested for safety is not nearly enough with respect to how critical they are in the whole system.
A while ago I stumbled upon this presentation on system safety [1] which had an interesting perspective coming from the aerospace industry. In aerospace they test the shit out of every component to make sure that a failure does not cause an airplane to crash. In comparison waymo, uber and everyone else has done almost nothing in terms of testing for safety before putting out their products.
[1]: Richard Murray: "Can We Really Use Machine Learning in Safety Critical Systems?" https://youtu.be/Wi8Y---ce28?si=HsqgiLngdHojpYO9
On technical level two sensors are clearly better than one even if you just pick one in case of disagreement, but as others have said Kalman filters and other more advanced techniques exist. There is a reason airplanes or spacecraft have multiple redundant sensors like this for decades.
The argument only makes sense if you want to save money, but then say you are being cheap up front.
It is indeed Tesla marketing that posits otherwise, which is wrong, and Tesla fans eat it up.
Somehow everyone else has figured it out, and even Tesla knows how to do it and have done it for years.
It’s pure bunk that’s used to cover for other decisions and now gets parroted around
I own a Tesla but I also acknowledge that Elon will claw every last dollar he can to increase margins even by a few cents.
Can you explain this? If you always pick the same one in case of disagreement, what is the purpose of the other sensor? You're not getting any additional information when they agree.
Imagine a stop light with a lidar sensor that broadcasts that information. Car ahead broadcasting that it's stopped in the fast lane.
They're hardly worth discussing as serious solutions today.
We are perfectly capable combining both inputs and act accordingly. An AI system can easily learn how to combining visual and LIDAR input to make the right decision given the circumstances.In the end it is just a decision tree.
Same as if I'm driving in winter and hit a patch of ice - Visually, it looks identical to the rest of the snowy, icy roads here, but I will drive if I feel and here that I'm spinning my wheels or sliding sideways.
If I smell coolant, oil, or a belt, even if my eyes tell me my gauges disagree I'll be pulling over for those issues as well.
Conversely, if I've replaced a tire and the TPMS light says the removed tire in the cargo area has low pressure (duh, that's why I changed it) but I know the spare (without a TPMS transmitter installed) is good, or otherwise know that the idiot light is a false positive, I'll trust my other senses over my eyes.
* Screeching breaks from the right
* Sounds of a rattling bicycle trying to undertake you
* An approaching emergency vehicle
* Other cars beeping their horn at you
Of course you will confirm such situations visually, but you definitely using hearing in addition to sight.
Like one sensor says you have a bus stopped in front of you and the other says it's all clear? And your choices are full steam ahead or prepare to not ram the apparent bus?
There are two cameras in this computer vision system. If one of them goes offline, is obscured, becomes dirty, or malfunctions you lose stereo vision and depth perception. So you'd obviously need 3+ cameras. And now we're right back at sensor fusion challenges again.
If you hear a sound but see nothing, you do become much more aware of the generic area afterwards, something like that should also be possible.
That being said, even automobiles make safety tradeoffs for cheapness or feasibility. However, we really shouldn't allow any tradeoffs for a completely unnecessary feature like "self driving". Imagine if wanting your car to have android auto or similar meant it couldn't use the lights, because a tradeoff was made.
In a life or death situation, you should opt for the system which will keep you alive more, not the one that costs less.
That’s probably also the cheapest viable strategy.
If a sensor-fusion car cuts accidents per mile (vs human) by 100x, but can only be deployed on 100,000 cars a year, and a camera-only car kills 10x more than that per mile, but can be put on 10,000,000 cars a year for the same cost, the camera-only car will end up saving 10x more people than the “better” system.
(I exaggerated both the improvement ratio and cost ratio because I like multiplying by powers of ten)
If the outcome of your safety discussion ends up suggesting a “safer”, “more expensive” solution that will definitely leave more people dead and injured then that analysis, then something is seriously wrong.
Going for absurd safety standards or expectations is absurd and self defeating. Again, as the other anon said, a practical solution that helps improve safety without handwaving material realities (cost, feasibility, adoption rates) is always better than a "safer" option that won't actually be used.
Obviously corporations try to make more money, but people also dont like buying more expensive cars.
A calculation that leads you to underdesign a product's safety and leaves no room for this product's safety improvement, in terms of mechanical or electronic update, is clearly not thought as being safe in that regard, regardless of the economies of scale or even low-term utilitarian goals (that would be expressed as: people spending money on a tesla would be safer in the short run, rather than using no automatic driving at all while waiting for a better product).
This is an important difference, and there is a societal choice to make here: do we (as society) want to buy now, and potentially have regrets later (when the safety of the product degrades with time, causing it to also have a record of people's deaths), or do we want to proactively force a notion of safety onto cars that is more than just being good enough at an arbitrary point in time, so that we hav more confidence over the long-term viability of that (societal) investment? As you can guess I gravitate towards the later, but of course it's a gradient, with several choices in-between, because pushing that thinking to an extreme would lead to stagnation, which would not do anything in terms of improving safety, as you noted.
There is no robust proof that any self-driving system outperformes a well-trained driver.
We could take that money and invest it into advanced driving lessons
As far as LIDAR itself: sure, yeah, you get depth info out of it. But depth info is only part of the problem, and frankly it's clear at this point that it's one of the easiest. The hard parts are in the recognition side: not "is that pedestrian in your path" (easy), but "is that pedestrian going to step into the street or not" (hard). And that's a computer vision problem. You can use LIDAR output as vision input, sure, but it's has no advantages.
Tesla was right, basically.
That's proven false by the cars continuing to drive into stationary objects. This failure mode is not ambiguous.
Camera input is garbage for interpreting geometry, especially from very smooth or very discontinuous surfaces, and especially with the shitty low resolution and low dynamic range cameras they use, and especially with non-stereoscopic cameras with no motion freedom relative to the vehicle body. Lidar is a necessary crutch for working around the fact that, while hypothetical cameras that don't exist might work well, all available cameras are unsuitable for the purpose, and calculating multi-view geometry accurately costs time.
> The hard parts are in the recognition side: not "is that pedestrian in your path" (easy), but "is that pedestrian going to step into the street or not" (hard). And that's a computer vision problem.
Pedestrian motion is not strictly a vision challenge but a general category of environment understanding (mass, momentum, motion mechanics). Vision is only one possible input mode preliminary to modeling.
It has? This again gets to "are they safer than human drivers?", because the competition hits stationary objects all the time. If you have data let's discuss data, but "proven false" is, again, just spin.
> Pedestrian motion is not strictly a vision challenge but a general category of environment understanding
Semantic evasion. You agree that it's "not a problem solved by LIDAR", right? It needs a camera. You can use a LIDAR output as a (somewhat inferior) camera, but it's not providing any advantages.
It looks like you're jumping from "depth is easy with cameras" (demonstrated false) to "safer than humans anyway without it" (speculative and not demonstrated by anyone), so who here is evading? That they're safer is not demonstrated. That depth is easy with just cameras is demonstrated to be false by the continuing failures in the presence of extreme financial incentive to not have those failures.
> You agree that it's "not a problem solved by LIDAR", right? It needs a camera. You can use a LIDAR output as a (somewhat inferior) camera, but it's not providing any advantages.
The LIDAR addresses the part where all current cameras are unsuited to mapping physical world geometry under driving conditions. It's not one or the other, but you appear to be assuming an imaginary not-the-one-we-live-in reality where only one is needed because you assume that all available cameras aren't actually very bad. But they are all actually very bad. So we continue to need both for the indeterminate future until someone invents mechanically robust extreme fidelity stereoptic cameras with motion freedom independent from the vehicle body, which is what humans use.
Humans are unsafe predominantly because of inattention, not ability. Camera-only vehicles are unsafe because of camera ability before you even get to the attention part.
Tesla's repeated failures over the years (and your conviction toward what Tesla is doing regardless) demonstrate a dangerously erroneous belief that object identification is the first and most important step for path planning. But that's not how humans drive, and it's not how to drive safely. The vehicle should avoid driving into any space that isn't going to be open smooth road, period, so the most important step is mapping geometry. There are no cameras currently suited for that. This is not a theoretical limitation. Just a practical one. Becoming suitable with current cameras would require many more cameras with much more processing per frame, so if you're trying to save costs vs lidar, you won't.
humans are absolutely doing sensor fusion, brains are bayesian inference machines. do not underestimate the power of the visual system.
and no, that the brain does is not an argument in favor of LIDAR-less cars. the eyeball + visual cortex system is alien technology compared to our feeble models. beware the hubris of a man who has learned to classify golden retrievers.
Not in the sense in the upthread comment they aren't, no. We have two cameras and two microphones. The latter is limited to weak detection of horns and tire screeches and not much else, and the former are too close together to give stereoscopic depth information at traffic distances.
We have a camera, basically. We do lots of stuff with the camera, sure. But that's not sensor fusion.
We sure as hell don't have anything like LIDAR.
They're only too close together to give very precise depth information, but they do still provide useful depth information. They also double the incoming light and SNR, which is why the average person performs better on visual acuity tests with both eyes than with only one or the other.
> We have a camera, basically. We do lots of stuff with the camera, sure. But that's not sensor fusion.
They're varifocal cameras with very good dynamic range that receive double the light input and that also have full freedom to move around, both rotationally and translationally, which provides, among other things, more depth information and better object boundary segmentation from controlled parallax and focus, and which involves the continuously varied activation of many different muscles and sensory nerves because they're attached to the extremely complex and sensitive proprioceptive structure called the rest of your body, which your brain fully uses as input when processing visual information. And we know that your brain uses this other information, because not having this other information causes reduced perception and motion sickness.
So, no, they're not just cameras. And, yes, we do sensor fusion.
Sonar sensors are most accurate at medium ranges, but they are notorious for detecting ghost objects that do not really exist. Infrared range sensors are more reliable but are only accurate at very short range. So when a sonar sensor detects an object 8.4 meters away, you use the infrared sensor to double check. If the infrared sensor says there's an object 9 meters away in the same direction, you assume the object is real but is actually 8.4 meters away. If the infrared sensor says the nearest object in that direction is 20 meters away, you assume the sonar sensor made something up.
If you have enough types of sensors, you can also use a "majority rule". If two of 3 sensor types agree, you assume the 3rd is an anomaly. Lidar is excellent for this because it is accurate across a very large range, so it tends to overlap with most of your other sensors. This increases that odds that when there is a disagreement, one of the agreeing sensors will be capable of accurately measuring the distance to the object.
Do AI systems have the potential to weight or inform those transactions based on historical historical data then? The “experienced” aspect of learning all the things that turned out to be true or false in previous comparisons or data decision points would seem to be the obvious missing piece, but I have never really understood the specifics.