Those in AI data centers never stop running and completely utilize their capacity. The difference in power usage is astronomical.
Those in AI data centers never stop running and completely utilize their capacity. The difference in power usage is astronomical.
I don’t claim to know, but we ought to be able to have a rational debate on this.
There's nothing irrational about suggesting AI GPUs are consuming far more power
Apparently a single gaming GPU can be used to run an LLM that serves hundreds of concurrent requests.
> Benchmarking Llama 3.1 8B (fp16) on our 1x RTX 3090 instance suggests that it can support apps with thousands of users by achieving reasonable tokens per second at 100+ concurrent requests.
You're essentially arguing that shipping naval diesel aggregates must be trivial because you can fit a dozen moped motors on the bed of your pickup truck just fine.
I have no insight into how many GPT-4 users are served per GPU, but I would assume OpenAI heavily optimizes for that, considering the cost to run that thing. It's probably in the same ballpark: hundreds-thousands of concurrent user requests per GPU. Still better than one GPU per gamer, even if it requires 10x the energy.