For example, let's say that the car is approaching an intersection, and suddenly sees a puddle on the road to the left getting brighter - a visual world model like this might extrapolate a scenario that the brightness is the result of a car moving towards the intersection assigning this some probability, and signing another probably to a scenario that it's just a flickering headlight, and the car would then decide whether and how much to slow down.
In this example there is a sensor, but it definitely doesn't tell the robot "exactly what is there", and while we could try to write rules about what it should do, the Bitter Lesson tells us it's better to just let it create its own model.
Self driving vars have cameras as part of their sensor suite, and have models to make sense of sensor data. Video will help with perception and classification (understanding the world) with no agency needed. Game-playing will help with planning, execution, and evaluation. Both functions are necessary, and those that come after rely on earlier capabilities