I don’t know much about gaming GPU workloads and how they differ from LLM inference workloads, but I’m definitely rich enough to pay for an LLM subscription and definitely too poor to afford the GPUs it would take to get anywhere near that level of inference.