Edit for clarity: You’re comparing a platform (Bard, GPT) to a model (llama, etc). The majority of folks playing with local models are missing the platform.
In order to close the gap, you need to hook up the local models to LangChain and build up different workflows for different use cases.
Consequently, this is also when you start hitting the limits of consumer hardware. It’s easy to download a torrent, double click the binary and pass some simple prompts into the basic model.
Once you add memory, agents, text splitters, loaders, vector db, etc, is when the value of a high end GPU paired with a capable CPU + tons of memory becomes evident.
This still requires a lot of technical experience to put together a solution beyond running the examples in their docs.