"The limits of my language mean the limits of my world"
Our entire personality and sense of identity is nothing but the reflection of ourselves in relation to eachother. The summation of all cultural input. A man in a void is no more than a beast.
So really there’s not a distinction between language models and people. Other than the lack of decision making and planning, the actual communication side is identical.
The possibility that we really are just stochastic parrots is too much for people.
What prompts are you using? It helps me make decisions every day involving complex and nuanced criteria. GPT-4 (especially 32k API + plugins) is better than 95% of the directors, PMs, and CEOs I have worked under.
Five pages of results; past the first page, there are only five relevant matches. https://news.ycombinator.com/item?id=23895706 three years ago (when it wasn’t yet a problem, but it was obvious it could become a problem), then the rest all within the last year (mostly this year) (when it is obviously becoming a problem), about thirty in total, half of which are multiple related comments. So the true number of times the prediction or analogy has been made is under twenty so far, and they’re mostly saying “like someone recently said” rather than coming up with the analogy by themselves.
The number seemed unrealistically high and I was curious. Now I have an answer! :-)
Interesting rabbit hole this one... but FWIW it's not as dramatic as this comment makes it sound. Newer steel is good enough because background radiation has dropped significantly since the 60s, plus there are non standard steel making techniques that don't lead to contamination.
I cant pull up the stats via chatGPT, but I can pull up cancer stats from NIH - but the goal is to look at cancer death rates for the population of Las Vegas Nevada by each decade - as it states that radiation background has dropped since the 1960s -- It would be interesting noting births in Las Vegas starting in 1950, and comparing that with cancer rates btwn 1965-1975 compared with the number of nuke tests done in NV during the 1960s
-
My grandfather was a nuke eng for GE for 50+ years, he was on the design team of Hanford...
He died of thyroid/throat cancer via exenguination and we won a lawsuit against GE for nukes that were exposed to high radiation levels unbeknownst to them for exposure and risks/threats...
a unicorn, one might say
If some new AI was trained on those backups, could Reddit take the creators to court for using that data without a license?
I've always thought it fascinating that people claim authorship or release things under a license without disclosing who they are. I imagine something licensed by Donald Duck is not actually usable under that license.
If it worked like that anyone could re-release anyone else's things under a different license. I mean, which Donald is the real author? How are you going to hunt the duck for violations? How do I prove I'm the real duck?
I wonder whether some models fit this - well, you could certainly call it generative, and obviously adversarial - framework of networks?
The analogy would be how we can't distinguish fantasy, parable, and propaganda in histories written in the past, even though we have a lot more technology than they had 2500 years ago. We rule out most things that seem fantastic, like dragons or people rising from the dead (maybe the equivalent of finding extreme statistical improbabilities), while also entertaining the possibility that the past could have been radically different than we see it now, and some of the things we ruled out may have been real. We make suppositions based on the political alignments of the historians, and whether they could have had access to the information that they claim to record. We use geological and anthropological records to check credibility, and use confirmations found as sources for priors when judging other, maybe unrelated, things that said by that same scholar.
We've seen what current AI does when confronted by a situation like that, it fills in the gaps with plausible fantasy and hallucinates.
edit: not that there isn't low hanging fruit in detecting AI creations, but that's just a sign of early AI. AIs can't even draw hands.
But yeah. There are definitely entropy analogies in software development. A couple of places I've worked at I've described as "brownian." Engineering would move forward a bit then marketing would change the requirements and we would make progress towards the new requirements, but before long we would get new requirements. If you mapped out progress starting at the origin and then put different objectives along the peremiter of a unit circle, progress would look like the random walk of an atom in a warm gas.
Those are the places where you ship a product when progress accidentally meets expectation.
And if you believe that Google wouldn't use Gmail in a heart beat to start training their models once they figure out a good way to do it I've got a bridge to sell ya.
If the communities will be gone for good (not likely at this point) - who will they be serving ads to?
But while the ad revenue is probably fine, the API revenue is definitely what they think the cash cow will be, and practically speaking they could lose 75% of their users and still probably not lose a whole lot of value just due to the sheer amount of historical data they have (especially if they start selling deleted content.)
My guess is that the changes will still go ahead and the general quality of content on reddit will go down longer term although how much is an open question.
[1] https://www.adweek.com/social-marketing/ripples-through-redd...
AI is going to do a shitty job identifying AI, so all of your datasets are going to get worse and worse. I could imagine a bunch of generative AI bots just melting down every message board on the planet to the point where you have to stop training on recent data.
I was hoping for the bots to export their existential crisis' to the rest of the web. Bots producing a loop, where they train on the data that they produce, would create some insane results.
There's a big difference between ghostwriting (biographies and fiction) and then the now fully GPT-generated books (non-fiction, especially technical) for sale on Amazon.
Text quality in reddit is awful. Especially comments . Forget using it for any semblance of actual grammatic correctness.
Probably can be used, after heavy cleaning and curation.
…
and while I enjoy my ivory tower as much as the rest of us, I do think it's important for the AI to understand everyone, not just me.
Of course, I don't know if scaling AI with more data will work to improve it or not, and I'm sure that anyone really knows right now.
(Which is to say... I think there are more than homeopathic concentrations of questionable training data. Or at least I believe that was the assertion of the OP.)
A pretty depressing thought to imagine our societies run by a bunch of ai redditors.