The Vision transformer is capable of predicting that a running child is likely to follow that rolling soccer ball. It is capable of deducting that a particular situation looks dangerous or unusual and it should slow down, or stay away from danger in ways that previous crop of AI could not.
Imo, the only thing currently preventing transformers to change everything is the large amount of compute power required to run them. It's not currently possible to imagine GPT4-V running on a embedded computer inside a car. Maybe AI asic type of chip will solve that issue, maybe Edge computing and 5G will find it's use-case... Let's wait and see, but I would bet that transformer will find it's way in many places and change the world in many more ways than bringing us chatbots.
It all needs to be onboard. That’s where money should be going.
More seriously, for safety critical applications, LLM have some serious limitations (most obviously hallucinations). Still, I beleive they could work in automotive application assuming: high quality of the output (better than current SoA) and very high token count (hundreds or even thousand of token/s and more), allowing to bruteforce the problem and run many inferences per seconds.
Clearly we are not there yet.
I wasn't intending to say it would be useful today, but pushing back against what I understood to be an argument that, once we do have a model we'd trust, it won't be possible to run it in-car. I think it absolutely would be. The massive GPU compute requirements apply to training, not inference -- especially as we discover that quantization is surprisingly effective.
That's also where I would see transformers or another AI architecture with reasoning capabilities shine: the fact that it can reason about what is about to happen would allow it to handle edge cases much better than relying on dumb sensors.
As a human, it would be very difficult to drive a car just looking at sensor data. The only vehicule I can think of where we do that is submarines. Sensors data is good for classical AI but I don't think it will handle edge case well.
To be a reasonable self-driving system, it should be able to decide to slow down and maintain a reasonable safety space because it is judging the car in front to be driving erratically (ex: due to driver impairement). Only an AI that can reason about what is going on can do that.
What is vision if not sensor data?? Our brains have evolved to efficiently process and interpret image data. I don't see why from-scratch neural network architectures should ever be limited to the same highly specific input type.
I also think dumb sensors is unfair, there are Neural Network solutions for processing LIDAR data so we are talking about a similar level of intelligence applied over both sensors.