I don’t claim to know, but we ought to be able to have a rational debate on this.
I don’t claim to know, but we ought to be able to have a rational debate on this.
There's nothing irrational about suggesting AI GPUs are consuming far more power
Apparently a single gaming GPU can be used to run an LLM that serves hundreds of concurrent requests.
> Benchmarking Llama 3.1 8B (fp16) on our 1x RTX 3090 instance suggests that it can support apps with thousands of users by achieving reasonable tokens per second at 100+ concurrent requests.
You're essentially arguing that shipping naval diesel aggregates must be trivial because you can fit a dozen moped motors on the bed of your pickup truck just fine.
I have no insight into how many GPT-4 users are served per GPU, but I would assume OpenAI heavily optimizes for that, considering the cost to run that thing. It's probably in the same ballpark: hundreds-thousands of concurrent user requests per GPU. Still better than one GPU per gamer, even if it requires 10x the energy.