Right now if you look at GTP-3's output it seems like it's approaching a convincing approximation of a bluffing college student writing a bad paper, correct sentences and stuff but very 'cocky'. It cannot tell right from wrong, and it will just make up convincing rubbish, 'hoping' to fool the reader. (I know I'm anthropomorphizing but bear with me).
Current models are being trained on a huge amount of internet text. As smarmy denizens of hackernews we know that people are very often wrong (or 'not even wrong') on the internet. It seems to me that anything trained on internet data is kinda doomed to poison itself on the high ratio of garbage floating around here?
We've seen with a lot of machine-learning stuff that biased data will create biased models, so you have to be really careful what you train it on. The dataset on which GTP-n has to be trained has to be pretty huge(?); and moderation is hard(?) and doesn't scale; it's easier to generate falsehood than truth; and the further we go along the more of internet data will be (weaponized?) output of GTP-(n-1); So won't the arrival of AGI just be sabotaged by the arrival of AGI?
Has anyone written something about the process of building AGI that deals with this?