Inference is much less energy intensive and could be done on small chips.
Regardless, I'm not as certain as the author about the future of ML on small devices. Some ML models are huge and needs to be updated frequently, therefore there is little sense in downloading those to small devices. In such cases, it makes much more sense to send feature data to a remote server that can generate a prediction within milliseconds, and then transmit that prediction back to the device.