The idea is that if you can take an image and predict the distances to the objects in it, you can then create an image that is a transpose of the original based on your predictions, but at a different distance. In case of Tesla moving forward, that would be a split second later, closer to the original image. As the car moves forward, you actually take that second image and compare the virtual to the original.
At this point you probably get it wrong, but you adjust your predictions and recalculate until your virtual transposed image looks exactly like the real image, at which point you've got the distances pretty darn close.
I remember watching a talk about their computer vision, and it was pretty impressive. The challenge is that despite knowing where the billboards are, they still look that way because sometimes, in obscure situations, in little old towns with fuzzy rules, road signs actually do appear in weird places, and that's where the car often trips up because without sufficient data it just can't build a prediction model to know when to respond, and when to suppress the new input.
I have a side project where I manage to get a helluva lot of depth info using a single photo from a single camera (monocular photogrammetry).