I just can't stop thinking though about the vulnerability of training data
You say good enough. Great, but what if I as a malicious person were to just make a bunch of internet pages containing things that are blatantly wrong, to trick LLMs?
You say good enough. Great, but what if I as a malicious person were to just make a bunch of internet pages containing things that are blatantly wrong, to trick LLMs?
So Reddit?
I’d imagine the AI companies have all the “pre AI internet” data they scraped very carefully catalogued.