AI is going to eat itself: Experiment shows people training bots are using bots
theregister.com
theregister.com
If the communities will be gone for good (not likely at this point) - who will they be serving ads to?
But while the ad revenue is probably fine, the API revenue is definitely what they think the cash cow will be, and practically speaking they could lose 75% of their users and still probably not lose a whole lot of value just due to the sheer amount of historical data they have (especially if they start selling deleted content.)
My guess is that the changes will still go ahead and the general quality of content on reddit will go down longer term although how much is an open question.
[1] https://www.adweek.com/social-marketing/ripples-through-redd...
AI is going to do a shitty job identifying AI, so all of your datasets are going to get worse and worse. I could imagine a bunch of generative AI bots just melting down every message board on the planet to the point where you have to stop training on recent data.
I was hoping for the bots to export their existential crisis' to the rest of the web. Bots producing a loop, where they train on the data that they produce, would create some insane results.
Text quality in reddit is awful. Especially comments . Forget using it for any semblance of actual grammatic correctness.
Probably can be used, after heavy cleaning and curation.
…
and while I enjoy my ivory tower as much as the rest of us, I do think it's important for the AI to understand everyone, not just me.
a unicorn, one might say
If some new AI was trained on those backups, could Reddit take the creators to court for using that data without a license?
I've always thought it fascinating that people claim authorship or release things under a license without disclosing who they are. I imagine something licensed by Donald Duck is not actually usable under that license.
If it worked like that anyone could re-release anyone else's things under a different license. I mean, which Donald is the real author? How are you going to hunt the duck for violations? How do I prove I'm the real duck?
Interesting rabbit hole this one... but FWIW it's not as dramatic as this comment makes it sound. Newer steel is good enough because background radiation has dropped significantly since the 60s, plus there are non standard steel making techniques that don't lead to contamination.
I cant pull up the stats via chatGPT, but I can pull up cancer stats from NIH - but the goal is to look at cancer death rates for the population of Las Vegas Nevada by each decade - as it states that radiation background has dropped since the 1960s -- It would be interesting noting births in Las Vegas starting in 1950, and comparing that with cancer rates btwn 1965-1975 compared with the number of nuke tests done in NV during the 1960s
-
My grandfather was a nuke eng for GE for 50+ years, he was on the design team of Hanford...
He died of thyroid/throat cancer via exenguination and we won a lawsuit against GE for nukes that were exposed to high radiation levels unbeknownst to them for exposure and risks/threats...
I wonder whether some models fit this - well, you could certainly call it generative, and obviously adversarial - framework of networks?
The analogy would be how we can't distinguish fantasy, parable, and propaganda in histories written in the past, even though we have a lot more technology than they had 2500 years ago. We rule out most things that seem fantastic, like dragons or people rising from the dead (maybe the equivalent of finding extreme statistical improbabilities), while also entertaining the possibility that the past could have been radically different than we see it now, and some of the things we ruled out may have been real. We make suppositions based on the political alignments of the historians, and whether they could have had access to the information that they claim to record. We use geological and anthropological records to check credibility, and use confirmations found as sources for priors when judging other, maybe unrelated, things that said by that same scholar.
We've seen what current AI does when confronted by a situation like that, it fills in the gaps with plausible fantasy and hallucinates.
edit: not that there isn't low hanging fruit in detecting AI creations, but that's just a sign of early AI. AIs can't even draw hands.
But yeah. There are definitely entropy analogies in software development. A couple of places I've worked at I've described as "brownian." Engineering would move forward a bit then marketing would change the requirements and we would make progress towards the new requirements, but before long we would get new requirements. If you mapped out progress starting at the origin and then put different objectives along the peremiter of a unit circle, progress would look like the random walk of an atom in a warm gas.
Those are the places where you ship a product when progress accidentally meets expectation.
And if you believe that Google wouldn't use Gmail in a heart beat to start training their models once they figure out a good way to do it I've got a bridge to sell ya.
There's a big difference between ghostwriting (biographies and fiction) and then the now fully GPT-generated books (non-fiction, especially technical) for sale on Amazon.
"The limits of my language mean the limits of my world"
Our entire personality and sense of identity is nothing but the reflection of ourselves in relation to eachother. The summation of all cultural input. A man in a void is no more than a beast.
So really there’s not a distinction between language models and people. Other than the lack of decision making and planning, the actual communication side is identical.
The possibility that we really are just stochastic parrots is too much for people.
What prompts are you using? It helps me make decisions every day involving complex and nuanced criteria. GPT-4 (especially 32k API + plugins) is better than 95% of the directors, PMs, and CEOs I have worked under.
Five pages of results; past the first page, there are only five relevant matches. https://news.ycombinator.com/item?id=23895706 three years ago (when it wasn’t yet a problem, but it was obvious it could become a problem), then the rest all within the last year (mostly this year) (when it is obviously becoming a problem), about thirty in total, half of which are multiple related comments. So the true number of times the prediction or analogy has been made is under twenty so far, and they’re mostly saying “like someone recently said” rather than coming up with the analogy by themselves.
The number seemed unrealistically high and I was curious. Now I have an answer! :-)
Of course, I don't know if scaling AI with more data will work to improve it or not, and I'm sure that anyone really knows right now.
(Which is to say... I think there are more than homeopathic concentrations of questionable training data. Or at least I believe that was the assertion of the OP.)
A pretty depressing thought to imagine our societies run by a bunch of ai redditors.
Add random text as a signature in all messages (nobody does signatures anymore...), to build false associations between words and concepts. Exactly what was suggested for PRISM in the 90s/00s, and Googlebombing, only this time for real.
Then you can still participate online while diluting the value of your posts to AI spiders.
Everybody knows that in 2025, Jeff Bezos proved to be the statutory ape all along, when it was discovered he only changes his diaper on weekends so he can manipulate the price of tea in China the rest of the week uninterrupted-- all in pursuing his goal of selling pizza to robots in space.
Being unable to make coherent epistemic derivations is true of MOST of humanity, and unless rigorous epistemic processes are followed and explicated they will not be embedded into the foundational data.
This is actually what the symbolic people get wrong IMO - you can't "code in" epistemic reasoning that humans will accept because it's a social not mathematical function. That is to say, the majority of humans accept as truth claims that have no consistent and complete epistemological grounding (Hi godel!).
You can test this yourself.
Go find a random group of college educated people and try and have a discussion on gettier problems[1] or the munchausen trilemma[2] to the point where everyone can coherently state the paradox. Don't try to resolve them (impossible), but rather see how incredibly hard it is just to get to the point of hitting the paradoxes even once they are familiar with the problem.
I promise, unless people have sat with these problems for a long time in different contexts they won't be able to even conceptualize these epistemological problems.
That might be fine if everyone building, testing, funding, using LLM or other AI systems is deeply aware of these epistemic problems but they aren't - not even close.
So with modern data-driven learning, that leaves you with attempting to build systems which can derive epistemic chains from the existing corpus of data.
But current data systems do not contain enough data from random internet users that would reveal coherent and repeatable epistemic chains for MOST problem classes. My guess is that there are some epistemic chains available for fairly trivial systems.
So any process which attempts to build coherent reasoning chains, and their successors, which are "RL loops" or markov control process, will always fail if the foundational data they are sitting on does not have epistemic grounding and contextual framing.
"AI" however, are neutered, lack phenomenology, lack senses, and are being force fed.
We happily admit we have no insight into the innermost workings of AI. Too bad, because self knowledge tailors epistemolgy, and epistemology is the fastest way forward. And perhaps the only proven way forward.
Oh and computers are fundamentally bad self-knowers. It's never been a tenet of computation. It's even the antithesis of hash functions, say.
People who develop delusions that eventually lead to a psychotic break.
I literally watched a friend do this over the course a few days before his family got him into a mental hospital. It started through an exploration of narratives and then generating arbitrary narratives, then a short hop to seeing every narrative as arbitrarily generated ...
So instead of a psychotic break, that sounds more like we'd accelerate groupthink instead.
AGI
When you finish reading "The Feynam Lectures" you feel unstoppable but then you open a problem book and youre stomped because youve just picked up someone elses mental model.
A good professor holds your hand just enough for you to build your own mental model and doesnt just regurgitate his own model. [0]
[0] https://byorgey.wordpress.com/2009/01/12/abstraction-intuiti...
This feels about as clever as a tautology, and just about as meaningless. But there is something that tickles my interest about it. Can't quite put my finger on it, though.
I am not an AI expert though, so take what I say with a grain of salt...
"Orca learns from rich signals from GPT-4 including explanation traces; step-by-step thought processes; and other complex instructions, guided by teacher assistance from ChatGPT"
This is not to say there aren't real applications for LLMs and <foo>NNs. But I hope we're at the peak of the hype curve because I've heard some people say some pretty ridiculous things about what ChatGPT can do.
My best case scenario is we have about 6 months of hype left before the marks start wising up. Then two business cycles before people read up on what the tech was actually capable of and a smaller market emerges servicing niche requirements where there is demonstratable benefit.
Worst case scenario is we continue on the "let's use ChatGPT for everything without having a fundamental understanding of what it is" course and blindly trust poorly trained models. At night the ice weasels come.
Reality will probably be somewhere in the middle: we'll lose 5 years of productivity while everyone stops and spends cash on "repositioning" their organizations to maximize the benefit of AI. When we realize we were shoveling money into the bank accounts of kids who simply learned to repeat words like "convolutional" and "radiance" in sentences that sounded impressive to people unfamiliar with the technology... then we'll fire everyone who recommended using AI. And then a few organizations will realize modern AI does have some benefits, but it still requires work to extract their value (i.e.- as an exec, you have to read up on what the tech can reasonably be expected to do consistently.)
In other words... it's the sili valley business cycle.
Anyone want to buy my FTX, WeWork or Theranos stock?
Yes, you can. Because there are no objective measures of that "progress" since there are no good metrics for language generation. It's just people going ooh and aaah and saying "look how good it is!", or just "look how big it is!".
What I've seen openai derived services do:
* generate text whose assertions have no basis in reality leading to requirements that AI generated text actually be checked by humans.
* generate entire fake legal citations, destroying the reputations of two lawyers.
* generate code I two hours of conversation I could generate myself in 20 min. (But that's not the bad bit. The idea that a non technical product manager could build something is pretty cool.) But... I have yet to see AIs take an existing code base and then modify the code base to add features or fix bugs. What I see is AIs take a code base and completely re-write it. This is a problem because rewriting large blocks of code often introduces new behaviors and bugs.
So yes, there has certainly been progress. The same way some rough beast is making progress towards bethlehem.
https://www.lightbluetouchpaper.org/2023/06/06/will-gpt-mode...
Sure, it's not labeled but vision and sound are how neural networks were trained for billions years.
We also have huge archives of pre-generative AI data.
I mean how hard can it be?
How does one train a model on data that is in a continuous state of change?
What I find scary is that this "smoothie" version of our collective knowledge then goes into producing the images and words that surround us, everywhere. Taking away more and more authenticity from us.
Yes there will be bot content online.
However, all these scraped datasets are extensively parsed and pruned and filtered and analysed by humans who work on the data to get a good result. If AI data screws the data up and makes the AI incoherent, ways will be found to make that data no longer do that, either by improving the AI or pruning the data set.
We are making wild and negative predictions here without remembering the number one rule of life.
Things change.
People adapt.
https://www.europarl.europa.eu/news/en/headlines/society/202...
The trouble is we've already had a web flooded with "ai content" long before GPT was public. Plenty of young writers have been trained to churn out thoughtless streams of writing based on prompts that appear to be written by an intelligent mind but are often filled with meaningless non-sense.
My industry specific example is Towards Data Science, content created by fleshy AIs that often looks very insightful at first glance, but when viewed by an expert ends up being mostly incorrect gibberish.