https://www.semianalysis.com/p/google-we-have-no-moat-and-ne...
Original author is an ML researcher, and the crux of his argument is that most weights in a LLM are significantly overdetermined. Once you have ingested several terabytes of natural language, you know how to generate natural language.
The remaining misses are facts that it has never seen, usually because they are so obvious that nobody thinks to write them down explicitly. And so more training data doesn't necessarily help LLM performance, unless you're either ingesting either extremely basic facts that are so obvious most adult discourse overlooks them, or you're ingesting expert knowledge that's highly specific and only discussed in a few forums. Reddit data could perhaps help with the latter, but a.) Reddit is usually not the place to go for expert discourse and b.) there are other better sources of data for it. You'd usually be better off training on trade publications, scientific journals, or fandom than Reddit.
Also the LLaMa/RedPajama approach makes this stupidly easy, because you can pass around patch sets to the model trained on specific mini-corpora, and then update the weights appropriately. Hence why the author of the Google memo believes neither Google nor OpenAI has a viable moat.