Since they updated the neural net to also recognize the vehicle orientations the dancing has stopped.
I have seen a lane change cancel recently, a couple weeks ago though.
You can tell the car is a 'nervous' driver. It plays it way too safe, but I guess that's a good thing at this point.
To be correct but one second late is to be completely inaccurate. The system is trying to estimate the current position of the car, but also predict future positions.
So a little bit of imprecision is fine since it improves accuracy related to predicting the future positions of the cars. A slight move in one direction may indicate a lane change, so it is always useful to be aware of that so as not to accelerate past a car whose measurement appears to be more inaccurate, since they actually might be moving. If you did the same thing with a human's "sixth sense" perception of the positions of the cars, you'd definitely find that they move a lot compared to their actual positions when the head is turned since our ability to merge our vision and our inertial sense is not very good for the most part.
The same issues arises with AR/VR, it's useless to know a more accurate position of the user if it's not the present position, because then that will definitely lead to motion sickness.
You could probably model inertia with n prior frames of probability fields.
What if they are hit by a truck? Maybe not 100,000 m/s^2 but if you assume that cars can't accelerate in directions they aren't pointed, you will be wrong at the worst possible time.
A better approach would be to include temporal data in the inputs to the neural net so it can learn how to do the prediction and filtering itself using all the context available in the input imagery, instead of processing each frame completely independently and feeding low-dimensional symbolic results into some other system. But you'd need a very large dataset and a very large neural net.
On the latest software and with HW 2.5, this is not true. It's still very much there.
What would be worrying is the model misclassifying an object, not detecting it at all, or having the bounding box consistently off.
Including human vision. The raw sensory data is pretty messy, and with some ingenious experiments some researchers can get a glimpse of exactly how messy.
The research on GANs also shows how computers can be fooled by things that wouldn't confuse humans.
If true, it makes sense the video would also have been captured using this emulation layer, explaining why it's not latest-and-greatest-fast.
And, of course, often real-time systems have much worse throughput performance than a non-real-time system on the same hardware. After all, latency guarantees are not free.
And yes, in practice you will probably combine a small real-time core with most of the code (by volume) running non-real time. Lots of one-time setup at the beginning of the right, or longer-range route replanning doesn't need to finish within tight deadlines. The human analogue is your co-pilot reading the map; or when you are fiddling with the aircon or radio dial.
It doesn't matter if you have a highly performant pipeline for detecting other cars if you have a random 200 ms VM pause as another car blows a stop sign in front of you.
This is the thinking behind many types of speed control street layouts. You should also know where to look to anticipate where danger is most likely to come from, and be ready with some kind of action. This is why we do hazard identification tests as part of the driving test - looking in the right directions at the right times is crucial for operating a vehicle safely.
250ms is a fairly average reaction time for something visual that you are ready for - but you should really be giving yourself as much time as possible - if somebody bombs past a traffic light at 70mph, even if it is green, most people would agree that it was an unsafe move. This goes doubly for an autonomous car, that is unable to play the positioning negotiation game that humans are masters of as a result of being social creatures.
And if you're not hovering your foot over the brake pedal, you're not getting 250ms.
The bigger problem is that what you might actually be getting is (almost) arbitrarily long pauses with a long tailed distribution. So sometimes 200ms, rarely a second, and every once in a while perhaps two seconds, etc; and no guarantees on the longest pause.
in motion the car drives just fine with the caveat they have not enabled signal recognition. I use TACC and at times full AP on my daily commute which includes road speeds from 35 to 55. I particularly like it on rainy days. I treat it like having a high school kid being chauffeur... I am a back seat driver who just happens to be in the driver's seat.
as for visual representation like in the video or waymo's demo videos, like many other things in life when you see how the sausage is made it is a wonder how we all survive it. The key difference between Tesla and Waymo is Tesla is not geo fenced, same with Cadillac's supercruise which is not available except on interstate.
who has the best solution, I am not willing to place a bet on that yet
The usual priors such as the Jeffreys prior [1] often do not work, because the posterior distribution will not be normalizable and estimates made by minimizing the expected loss will be inadmissible.
[1] In Bayesian probability, the Jeffreys prior is a non-informative (objective) prior distribution for a parameter space; it is proportional to the square root of the determinant of the Fisher information matrix.
Why is this of relevance?
It has the key feature that it is invariant under a change of coordinates for the parameter vector. That is, the relative probability assigned to a volume of a probability space using a Jeffreys prior will be the same regardless of the parameterization used to define the Jeffreys prior. This makes it of special interest for use with scale parameters.
Why is this an issue?
Accordingly, the Jeffreys prior, and hence the inferences made using it, may be different for two experiments involving the same theta parameter even when the likelihood functions for the two experiments are the same—a violation of the strong likelihood principle.
Ie objects don't spring into existence or disappear or fly around at 400 mph or change direction at 2000 gees.
Basically, no, you don't just throw it at some Bayesian math or a Kalman filter to make it look prettier. Yikes.
But yeah, the scales are heavily tipped.
However the main difference is that, when we are consciously looking directly at something, we can almost always tell with 100% certainty what we're looking at, up to a considerable distance. I can see a car pulled over to the side of the highway a solid half mile ahead sometimes, and have plenty of time to respond. Computer vision doesn't have this additional strength.
As always though, the strength that computer vision has over us is it never gets tired or distracted, and it never operates in "default mode" where sensory inputs don't get full (or even much at all) conscious attention.
Even when some new information forcefully comes into play, your brain is often able to adjust your memory so you believe you knew it along, so long as the initial percept is fresh enough and had enough uncertainty.
All of this feels to you like a perfect unbroken stream of direct seeing but it is an illusion. You don’t see anything directly, you get fuzzy spurts of probability and turn it into your world in your mind. A world that’s likely to be unrecognizable to the next person.
That's just your brain again. You might mistake a bike for a lamp post, and switch between beliefs several times, before you figure it out, then convince yourself you knew it the whole time.