Sure, but I bet Andrej Karpathy's team is using SOTA and mitigates most of the issues like the ones you mentioned. OK it's an "argument from authority" one might believe or not and I doubt they'd ever disclose things that make it work in their specific case. But realistically, if you e.g. observe results of real-time semantic segmentation, you see that surfaces flip a bit in each frame but mostly stay correct; you can e.g. use average IoU/coverage from past n frames to estimate what is going on. They have a fleet of cars driving around receiving more training data in all kinds of environments to handle various perceptual conditions as well.