Not an ML guy so maybe missing something important here. It seems like many use cases could be covered by even dumber, simpler models created with the aide of general purpose LLMs.
Like, can't I just tell Claude or GPT to help me create a BERT or <insert old boring ML framework model> that's even cheaper and faster?
For instance, Frigate runs 4.6mb yolov9 image object detector on a $50 USB Coral stick. It takes like 1-2 watts of power, responds in 15ms, and is completely local. A pay-per-token cloud hosted not-really-an-LLM-but-probably-based-on-one that runs in 150+ ms still seems terribly inefficient.