I think this is a critical part of cracking self-driving cars - "scene understanding", or more accurately, identification AND MODELLING of all traffic participants. A tiger chasing a deer will identify and model what the deer will do, what the terrain is, and take appropriate decisions on how to best pursue; a good driver will do the same - e.g. understanding that a round object bouncing in the street of a residential neighborhood is a ball, and where a ball comes, a child will follow.
I also think that this feedback loop of "Model the world, take action, observe, update model" is a key component of consciousness (whereas we model ourselves!) and key to AGI. I would not be surprised that AGI comes from a self-driving car development model.