> Individual LLM requests are vanishingly small in terms of environmental impact;
The article uses open source models to infer cost, because those are the only models you can measure since the organizations that manage them don't share that info. Here's what the article says:
> The largest of our text-generation cohort, Llama 3.1 405B, [...] needed 3,353 joules, or an estimated 6,706 joules total, for each response. That’s enough to carry a person about 400 feet on an e-bike or run the microwave for eight seconds.
I just looked at the last chat conversation I had with an LLM. I got nine responses, about the equivalent of melting the cheese on my burrito if I'm in a rush (ignoring that I'd be turning the microwave on and off over the course of a few hours, making an awful burrito).
How many burritos is that if you multiply it by the number of people who have a similar chat with an LLM every day?
Now that I'm hungry, I just want to agree that LLMs and other client-facing models aren't the only ML workload and aren't even the most relevant ones. As you say adtech has been using classifiers, vector engines, etc. since (anecdotally) as early as 2007. Investing algorithms are another huge one.
Regarding your USENET point, yeah. I remember in 2000 some famous Linux guy freaking out that members of Linuxcare's sales team had a 5 line signature in their emails instead of the RFC-recommended 3 lines because it was wasting the internet or something. It's hard for me to imagine what things were like back then.