Yeah cause every car needs 8xH200 pulling 10kW to run a VLM at realtime speeds. Would be unfortunate if 4G dropped out under some trees while using the API after all.
And the disinclination of these companies to push the weights of their cutting edge models into people’s cars where they can be dumped.
When the models stop improving, we will get model-specific ASICs that are much more power-efficient.
Soo, never? Granted Cerebras is a thing, if the process can be commoditized.
At the moment the area of edge inference at speed seems pretty bleak though.