Since they updated the neural net to also recognize the vehicle orientations the dancing has stopped.
I have seen a lane change cancel recently, a couple weeks ago though.
You can tell the car is a 'nervous' driver. It plays it way too safe, but I guess that's a good thing at this point.
To be correct but one second late is to be completely inaccurate. The system is trying to estimate the current position of the car, but also predict future positions.
So a little bit of imprecision is fine since it improves accuracy related to predicting the future positions of the cars. A slight move in one direction may indicate a lane change, so it is always useful to be aware of that so as not to accelerate past a car whose measurement appears to be more inaccurate, since they actually might be moving. If you did the same thing with a human's "sixth sense" perception of the positions of the cars, you'd definitely find that they move a lot compared to their actual positions when the head is turned since our ability to merge our vision and our inertial sense is not very good for the most part.
The same issues arises with AR/VR, it's useless to know a more accurate position of the user if it's not the present position, because then that will definitely lead to motion sickness.
You could probably model inertia with n prior frames of probability fields.
What if they are hit by a truck? Maybe not 100,000 m/s^2 but if you assume that cars can't accelerate in directions they aren't pointed, you will be wrong at the worst possible time.
A better approach would be to include temporal data in the inputs to the neural net so it can learn how to do the prediction and filtering itself using all the context available in the input imagery, instead of processing each frame completely independently and feeding low-dimensional symbolic results into some other system. But you'd need a very large dataset and a very large neural net.
On the latest software and with HW 2.5, this is not true. It's still very much there.