Apparently a single gaming GPU can be used to run an LLM that serves hundreds of concurrent requests.
> Benchmarking Llama 3.1 8B (fp16) on our 1x RTX 3090 instance suggests that it can support apps with thousands of users by achieving reasonable tokens per second at 100+ concurrent requests.