There’s lots of work on distillation, smaller models, approximations, etc. People already have simpler forms of these running on smartphones. Models seem to be growing faster than we can make them small though :D
Yes, these things keep us up at night as well :-).
I think this is largely unnecessary, can't things like TPUs handle the inference?
Using TPUs is expensive. Also some applications may need small latency
Putting all your speech/text onto cloud machines runs counter to e2e encrypted messaging.
I think you can get hardware like a TPU for consumer products?