"Automating curation of the datasets with the latest and greatest GPT sounds very possible however."
It might also lead to recursive garbage generation. If the AI has flaws and those flaws are from the initial training, then I see no way, how the flawed data can ever generate clean data.
"The datasets are massive, and curating them into buckets of fact/fiction by hand would be next to impossible. "
And it is not impossible. It is just a lot of work. And maybe work we just have to do, if we want to create reliable AIs, that do not fail at random times.
Now this alone might not completely remove hallucinations, but solid training data is just the base of it all.
(and since we are living in the area of fake news, I welcome all efforts towards established facts and data out of general principle)