Is the energy usage so different between local and cloud inference? Both require electricity, the local option even more than the better optimized cloud variant even perhaps. How either is powered makes the crucial difference I suppose. Both can potentially run on solar as well as gas or nuclear.
It's the training that takes the most energy, and that needs to happen for locally running or cloud models regardless.
What am I missing here? EDIT:
Proposal H mentions "LLM usage accelerates the destruction of our ecosystem"
I was thinking solely about energy usage, but there is of course also water usage. A local setup is not water-evaporator cooled most likely.