There's a huge difference between training and running models.
The multimillion (or billion) dollar collections of hardware assemble the datasets that we (people who run LLMs locally) run. The non-open datasets that we-host-LLMs-for-money companies do the same, and their data isn't all that much fancier. Open LLMs are catching up to closed ones, and the competition means everyone wins except for the victims of Nvidia's price gouging.
This is a bit like confusing the resources that go in to making a game with the resources needed to run and play that game. They're worlds different.