The memory bandwidth and size seems to be there, but what is the tokens per sec on like a qwen model? And you can basically do 3x opus 4.5 on the $100 a month claude plan. Your payback will be near infinity years after electricity.
how about measuring returns in the sense that I no longer need to share my flagship idea with some random 3rd party just because it hosts the LLM I am using?