But that's absolutely false about the Nvidia moat being only training. Llama.cpp makes it far more practical run inference on a variety of devices. Including ones with or without Nvidia hardware.
This is a leading-edge software library that provides a huge boost for non-Nvidia hardware in terms of inference capability with quantized models.
If you don't understand that, then you have missed an important development in the space of machine learning.
At length:
- yes, local inference is good. I can't say this strongly enough: llama.cpp is a fraction of a fraction of local inference.
- avoid talking down to people and histrionics. It's a hot field, you're in it, but like all of us always, you're still learning. When faced with a contradiction, check your premises, then share them.