AI has poisoned its own well
tracydurnell.com
tracydurnell.com
Majority of online users are lurkers, so all of these models are extremely biased on whom they got their information from.
I suppose if you are only vaguely familiar with Reddit you might it assume it's like a vBulletin setup where only a handful of subforums are set up by admins rather than users.
Same thing could be said of people that aren't online too much or barely never on tech forums, like John Carmack.
The thing is, there's no objectively unbiased sample of information. You could weight towards tweets or books, NYT or the Economist, arXiv or journals. The books could be bestsellers or classics. There's no right answer, just preferences.
On the net, search and content quality is already pretty low for certain key words. If a word is part of the news cycle, expect hundreds of badly researched news paper articles, some of which might be generated as well. Or if they weren't, you wouldn't notice a difference if they become that.
But I don't believe the companies made a mistake. They could even protect their position with the data they already acquired and classified. Maybe a quality label would say genuine human®.
If all that fails large companies would also be able to employ thousands of low wage workers to classify new content. The increasing memory problem persists, I think that is a race where the model that can extract data as efficiently as possible will win. But without the data sets, there is no way to verify performance.
I think the simple idea using chatGPT to filter it's input set with result in drastically higher quality dataset. Add in the ability for chatgpt to use plugins and such like wolfram alpha to fact check it's answers then chatgpt can actually start generating input data from know good sources in areas it knows it's having quality issues.
I mean, it literally looks like the beginning of self-learning AI.
And it doesn't seem clear to me. It may be true, but far from obvious.
For example, DALL-E, which was better than its predecessors. Was it better because of mode input data, that is, was more input data a necessary condition for its improvement? Or even the biggest reason? Reading the OpenAI blog makes it sound as if they had new kinds of AI models, new ideas about models, and that the use of data from a web crawler was little more than a cost optimisation. If that's true, then then it should be possible to build another generation of image AI by combining more new insights with, say, the picture archives of Reuters and other companies with archives of known provenance.
Maybe I'm an elitist snob, but the idea that you can generate amazing pictures using the Reuters archive sounds more plausible than that you could do the same using a picture archive from all the world's SEOspam pages. SEOspam just doesn't look intelligent or amazing.
I was surprised to learn that it didn’t take an enormous amount of data to train Llama.
Meta used 1.4 trillion tokens training Llama, but that’s only about a million times the Harry Potter collection [1]. Given that the kindle store has 12 million books [2], it’s credible to get 1.4T tokens just from that. Twitter has 200 billion new tweets per year, so it’d only take 7 tokens per Tweet for that to produce a Llama sized dataset every year.
[1] https://en.wikipedia.org/wiki/LLaMA [2] https://blog.fostergrant.co.uk/2017/08/03/word-counts-popula... (1 million words) [3] https://www.omnicoreagency.com/twitter-statistics
I don't know about the text models, but e.g. Stable Diffusion (and most of its derived checkpoints) has very recognizable looks.
By the way, does anyone know if such generative models could be used as classifiers, answering the "what's the probability that this input was generated by this model" question? That'd help solving the "obtaining quality training data" problem: use the data that has low probability of being generated by any of the most popular models. It's not like people started to produce less hand-made content anyhow!
About 7 years ago when I managed a deep learning team at Capital One, I did a simple experiment of training a GAN to generate synthetic spreadsheet data. The generated data maintained feature statistics and correlations between features. Classification models trained on synthetic data had high accuracy when tested on real data. A few people who worked for me took this idea and built an awesome system out of it.
Since the poisoned well is a known thing now, it seems like a solvable problem/
I think the truth is we’ve written all that ever needs to be written, and even if they universe becomes populated by on AI LLM chat bots communicating, they will be fine to feast off of what we’ve left them as a legacy.
I highly doubt this statement.
How useful would ChatGPT 4 be if it was trained only on data up to 2013? (Assuming the total amount of data it was trained on was the same) Would it be like talking to a human who was sitting in their basement for the past decade? I am not sure how useful that would be to me.
These days, it's a rare piece of text that surprises the reader with creative or ingenious or outside-the-box thinking. If we want higher level cognition from future LLMs, where will such deep thought come from? Surely not the training data used now: email, tweets, reddit, mass media, etc. GIGO indeed.
Much has been written, yes. But not much of that is worth reading.
I don't think it's any more rare now than it ever was. It's always been rare.
> Much has been written, yes. But not much of that is worth reading.
Sturgeon's Law is immutable and timeless.
Are we actually at this point now? I am less confident than you are...
And what is it going to read? How will it distinguish anything it reads today from AI generated content? Don't you see you've just set up the exact same circumstance that the article talks about?
That's a qualitative leap from what it does now, and there's no clear basis for believing the current trajectory of these systems can make that leap.
I’m a copywriter (I know, my days are numbered), so I didn’t just plop the generated text in unchanged, but it’s still primarily GPT’s content. I did a nice job, imho, and improved the article greatly. It’s been up for several days now, so I think it has a good chance of staying long term.
Still, it’s a funny situation. On the one hand, I did something I’ve done many times before: researched a topic and added to a wiki article. I always feel gratified when I contribute to Wikipedia. I’m adding to the sum of human knowledge in an incredibly direct way.
But, the information I added this time will be used to train future LLMs and thus “poison the well” with generated content. The verdict is still out on just how bad generated content will be for training new models. But I definitely feel slightly conflicted about whether I did something that is a net positive or negative.
Another factor is that generative data distributed online may be quite high quality (because people find it interesting enough to share), so it's plausible this could actually improve model performance. Some LLMs have been trained with data from other models with good results, e.g. supposedly Bard with GPT-4 prompt-response pairs and GPT-4 with Whisper transcripts of YouTube (https://twitter.com/amir/status/1641219919202361344/photo/1). Of course, there could be trolling or misinformation that "poisons" the data, and that is a problem (whether synthetic or organic)!
Human beings are given to (1) repeating tropes (2) arguing incoherently (3) missing the point, etc.
Think of all the books and movies from before 2023. Isn't there a lot that is formulaically wrong/misleading/suboptimal in there?
So this might not really be "poisoning the well" -- a more interesting area to look at would be, how can we make GPT-n aware of its own gaps and use this knowledge rather than just let the user know that it knows its gaps?
Creative content has been continuously devalued for decades now. The whole reason you can find so much music, creative writing, and art, available for free online is because this content is essentially worthless until you can build a brand or clientele to monetize it, and the only way to do that is to broadcast it for free to as many people as possible.