Sort of works for inference, but not for anything else. For now, I bought a 4x4090 workstation… but I am crying how power-inefficient it is.
You might like this article, which looks at the arithmetic intensity of LLM processing: https://www.baseten.co/blog/llm-transformer-inference-guide/