However, if you don't trust cloud providers or inference providers for whatever reason then you probably aren't going to be excited to enter a co-op model where you're still effectively renting access to hardware that you don't directly own. There are already reasonably priced options to rent bare metal from a cloud provider.
The only way I see it working is if it's a bunch of medium to large sized businesses getting together to be able to rent out the spare capacity on hardware that they physically control. So an AWS equivalent where each rack is owned by a different company and retail VMs migrate between them transparently. But I question the overall economics of such an arrangement.
If we're imagining a co-op, then the participants should all be equal owners in an organisation that owns the hardware itself, otherwise it's not much of a co-op really.
The advantage is having a voice in setting direction, investment and ultimate shared incremental profit recovery over the rent capture from hyperscalers.
It's an interesting idea really. But it also comes with cooperation, which means trust, and a lot of for-profit enterprises are bad at operating cooperatively outside of 2- or 3- way engagements.
I don't see it happening. A current gen GPU with a huge and fast block of memory isn't a perfect fit for these algorithms but it's relatively close. With cryptocurrency, mass small sha256 hashing was a totally different kind of computation.
I don't think that's true. The best fit out of what's presently available perhaps. Inference is almost entirely memory bandwidth bound at present, to the extent that GPUs with HBM have a massive advantage over those with GDDR. TPUs appear to be a much better overall design.
I expect that a hypothetical advance in fabrication enabling processing elements to be placed directly adjacent to dense RAM on the same silicon (not merely in the same package) would be superior in all regards.
Processing scales better than DRAM does. I think an HBM-like stack where the bottom layer has the math units is probably the ultimate form of that.
And it's possible that flash instead of DRAM is actually the better play, as long as you can hook up enough in parallel. RIP Optane.
Isn't it already happening with Cerebras? It's mentioned at the end of OpenAI's GPT 5.6 announcement:
"We're also launching GPT‑5.6 Sol on Cerebras at up to 750 tokens per second in July"