If you somehow do sensor fusion perfectly and the models never disagree, then why even have the LiDAR? At that point you’ve solved vision using cameras.
If you somehow do sensor fusion perfectly and the models never disagree, then why even have the LiDAR? At that point you’ve solved vision using cameras.
Each sensor is given an uncertainty model, that is for example a stochastic model by adding e.g. Gaussian noise to the "true" value it is measuring. Further, you have another model that describes the dynamics, e.g. equations from physics can tell you where the car will go if you know the current speed, position, etc..
1. The kalman filter computes a probabilistic prediction of what it thinks is happening by using the dynamics model. That is, based on what it knows so far, where will the car (probably) be when the next measurement comes up?
2. When measurements from various sensors come in, the Kalman filter uses Bayes's theorem to compute a mean (posterior), in which each measurement is weighted by the probability that the measured value is correct (using the uncertainty models; "correct" here means "in agreement with the prediction"). In other words, sensor that are inaccurate (large variance) will be considered less in the computation of the mean, while more accurate sensors are given more importance.
Once the mean of the measured quantities are computed they are used again in step 1 and the whole thing is repeated. As you can see, that disagreements are justified by the inaccuracies of the sensors, and the process of performing a probabilistic weighted average solves the problem. For the Kalman filter in particular, it can be be show (mathematically proved) that this process minimizes the variance (uncertainty) of the measured quantities (which btw. is an amazing result if you think about it).
A simple example of why this gets complicated: if I have a point in one camera and a point in another camera and I know they correspond to the same real-world spatial point, I can calculate some distance, and the statistics of that calculation can be captured by a halfway-reasonable error model. But how did I know they corresponded to the same point in the first place? Well, because they look the same according to some image feature... or because some deep neural network told me so.. etc. There just aren't very good ways to model just how haywire ^that^ process can go. So at the end of the day, once you let this evil into your perception system, using statistics to blend your sensors together is undermined, and all of your precious covariances just turn into tuning knobs you can twiddle.
The dirty secret is that almost all robotic perception systems are hiding unprincipled, un-modelled heuristics in the data association process. This is kicked under the rug because it doesn't really fit into traditional estimation theoretic frameworks. In a lot of papers you'll see academics push it aside by just calling that the "front-end", which they brush aside as a little widget you put on the front. If you're lucky they'll do ablations across a couple different options.
Of course, this is just one level deeper down the iceberg. It goes far deeper. Even if you could model the statistics of a depth camera well, the statistics of "what are all the objects in your scene about to do" is another couple of orders of magnitude more un-modellable. Often engineers will do something like attach a "constant-velocity" model to the agents in a scene. Imagine trying to bin all of the reasons you might stop walking in a straight line into a bubble that describes how "noisy" that picture of the world is! Now you can begin to appreciate just how hopeless it is to explicitly model uncertainty in the world around us.
You can start to add unjustified assumptions and that'll make the world model weaker, yes. But starting with pretty basic assumptions like "you can segment an object from an series of images because each solid object will move on its own trajectory", or even more basic like "objects have edges" and then a few dozen samples per second, and suddenly you have a fairly robust way to detect things.
Same for predicting where something will go. If you can estimate an objects current velocity, acceleration, and jerk with reasonable precision, you don't really need a highly predictive heuristic for the world model.
For decision making you need more robust heuristics, like "the car to my left has right of way at the stop sign", but you don't need that level of heuristic to identify that there is a car and that it isn't part of the pavement and that it is currently sitting still.
However, you should also consider that working with heuristics is pretty much all engineers do. We ought to solve problems, even if the theory is not there yet. A great example is how we got air travel long before we had any real understanding of the fluid dynamics happening around the fuselage (as it was generally computationally intractable). So, sometimes a simple epipolar camera model with noisy clouds around the subjects is sufficiently accurate for the task. The real problem IMHO is that the degree to which these rudimentary approximations are tested for safety is not nearly enough with respect to how critical they are in the whole system.
A while ago I stumbled upon this presentation on system safety [1] which had an interesting perspective coming from the aerospace industry. In aerospace they test the shit out of every component to make sure that a failure does not cause an airplane to crash. In comparison waymo, uber and everyone else has done almost nothing in terms of testing for safety before putting out their products.
[1]: Richard Murray: "Can We Really Use Machine Learning in Safety Critical Systems?" https://youtu.be/Wi8Y---ce28?si=HsqgiLngdHojpYO9
On technical level two sensors are clearly better than one even if you just pick one in case of disagreement, but as others have said Kalman filters and other more advanced techniques exist. There is a reason airplanes or spacecraft have multiple redundant sensors like this for decades.
The argument only makes sense if you want to save money, but then say you are being cheap up front.
It is indeed Tesla marketing that posits otherwise, which is wrong, and Tesla fans eat it up.
Somehow everyone else has figured it out, and even Tesla knows how to do it and have done it for years.
It’s pure bunk that’s used to cover for other decisions and now gets parroted around
I own a Tesla but I also acknowledge that Elon will claw every last dollar he can to increase margins even by a few cents.
Can you explain this? If you always pick the same one in case of disagreement, what is the purpose of the other sensor? You're not getting any additional information when they agree.
We are perfectly capable combining both inputs and act accordingly. An AI system can easily learn how to combining visual and LIDAR input to make the right decision given the circumstances.In the end it is just a decision tree.
Same as if I'm driving in winter and hit a patch of ice - Visually, it looks identical to the rest of the snowy, icy roads here, but I will drive if I feel and here that I'm spinning my wheels or sliding sideways.
If I smell coolant, oil, or a belt, even if my eyes tell me my gauges disagree I'll be pulling over for those issues as well.
Conversely, if I've replaced a tire and the TPMS light says the removed tire in the cargo area has low pressure (duh, that's why I changed it) but I know the spare (without a TPMS transmitter installed) is good, or otherwise know that the idiot light is a false positive, I'll trust my other senses over my eyes.
* Screeching breaks from the right
* Sounds of a rattling bicycle trying to undertake you
* An approaching emergency vehicle
* Other cars beeping their horn at you
Of course you will confirm such situations visually, but you definitely using hearing in addition to sight.
It did get me thinking about different senses and how I prioritize them. Given that humans strongest sense is visual, it's interesting to me that the priorities of what to trust seem opposite of my expectations. It seems to me that visual signals are seemingly least "trusted" compared to the other senses. As if the logic goes, "my sense of smell is so bad that if it detects danger, it must be very dangerous".
Similarly, I'm sure dogs can smell rotting flesh long before meat is unsafe.
In the context of vehicular autonomy it's more complicated than that because for example a radar sensor picking up an overhead sign or a truck parked on the shoulder as if it's an obstruction on the road is something you want to ignore when you're going 80 MPH on the highway rather than slamming on the brakes.
When you're the only vehicle on the road, stopping if anything goes wrong is always the safest idea. When you're one of hundreds of vehicles in a high speed flow of traffic stopping would put you and everyone else on the road at significantly greater risk.
Imagine a stop light with a lidar sensor that broadcasts that information. Car ahead broadcasting that it's stopped in the fast lane.
They're hardly worth discussing as serious solutions today.
There are two cameras in this computer vision system. If one of them goes offline, is obscured, becomes dirty, or malfunctions you lose stereo vision and depth perception. So you'd obviously need 3+ cameras. And now we're right back at sensor fusion challenges again.
If you hear a sound but see nothing, you do become much more aware of the generic area afterwards, something like that should also be possible.
I can almost see the objection that multiple cameras are as good as one camera+lidar, but I think it's a mistake to trust any system that can't check itself for consistency across multiple bands. It doesn't take a very good radar to keep you from ramming a fire truck. In fact, whatever runs the cruise control's distance sensor should have been enough to prevent a bunch of the Tesla oopsies reported in the press. When tackling one of the hardest engineering problems faced by humankind, it seems stupid not to take advantage of all the data you can get.
Like one sensor says you have a bus stopped in front of you and the other says it's all clear? And your choices are full steam ahead or prepare to not ram the apparent bus?