A good way to harvest new training material is to eavesdrop real human conversations from non polluted sources (such as microphones listening to people talk in public places or texts), transcribe them, and feed them to LLMs.
But our normal convos are plagued of mistakes, bad grammar, etc
It doesn't take much to clean up say 95% of mistakes I reckon, as it tends to be pretty repetitive, and unless there's a bunch of wordplay happening, intention can be discerned.