To be cost effective with inference providers, you have to find some way to be using it 24/7.
If they decided to collude, they could absolutely say "from now on you no longer have access to model X because you're an asshole"
The commercial inference offering are also downstream of one of those 3 projects (or trt-LLM if they're nvidia). It would impact Ollama, and fireworks, together, and everyone else.
Don't tempt fate.
if it doesn't apply to you then just come back in a couple years and see what the situation is then. 1 million context window, 1 million tiny layers to fit in 4gb RAM at a time, with 256gb of fast unified RAM in every consumer device? Or a different concept entirely
in the meantime, z.ai probably doesn't reply to US subpoenas so you can shift all your incriminating conversations over to that and use GLM anyway. who cares if the Party trains on your data and steals your IP and ignores you for legal matters, when the alternative in the US is just a thin corporate layer and party who steals your IP and will snitch on you for legal matters.
You're better off setting a budget and buying the best machine you can afford in that range, or picking a VRAM target and accepting the class of models you can run on it. Those models will almost certainly improve over time and your skills will adapt to the limitations. Hardware is so valuable right now that it's not even likely to be a significant loss if you had to sell.
Right now I think 24 GB is probably the best bang for your buck (used 3090), because you also get a high end gaming/gpgpu device which is nice anyway. 32 GB you can do with AMD or Intel, but NVIDIA is megabucks and at this point you're really paying for RAM. Unfortunately the ship has sailed on "reasonably" priced RTX 6000s, which at one point were about $7k and are being listed at $10k++.