OpenAI and Reddit Partnership
openai.com
openai.com
I have to assume because of the growing phenomenon of Reddit being the only valuable results returned from Google, it's becoming more valuable for companies to astroturf Reddit content that will appear in Google. So there must be either bot-run accounts, or some kind of business of karma farming to later sell off accounts with karma that won't get flagged for spam when they start endlessly posting about how great [PRODUCT] is.
> u can google ur questions about the world and get the CIA FBI answer or u can add "reddit" to the end of it and learn the truth
Examples from the current top rated askreddit monthly thread:
Scientific breakthrough
https://www.reddit.com/r/AskReddit/comments/1c9kenq/what_sci...
https://old.reddit.com/r/AskReddit/comments/wrbrs9/what_scie...
https://www.reddit.com/r/AskReddit/comments/x17fuf/what_scie...
Travel Tips
https://www.reddit.com/r/AskReddit/comments/10iptoa/what_is_...
https://www.reddit.com/r/AskReddit/comments/ueoys0/what_trav...
https://www.reddit.com/r/travel/comments/16lzwfc/what_are_yo...
https://www.reddit.com/r/AskReddit/comments/1iqvhc/whats_you...
If used this way low-quality content is GOOD for OpenAI's training as long as it gets accurately downvoted.
Thus, most top comments are jokes.
Anything informed but controversial will be buried down deep and can be only found by sorting for..controversial.
it's the same as facebook and twitter. gradually it turns into boomer memes which get more engagement from the increasingly mainstream audience
i'm not a luddite, or a "hippie". i'm going to keep paying for my GPT-N+1. I just want it in writing that if they can pirate everything, so can you and I. It's only fair.
As far as I'm aware, the people telling us not to pirate content are the people currently suing OpenAI for (allegedly) using pirated content to train their models.
Sadly I'm afraid that the ToS is going to f*k us on this
Big corporations/governments set the rules.
Rules are for small people.
My life got better when I switched from worrying if I was acting in a nonconformist way to worrying that I might be perceived as a conformist.
It's good to have a balanced opinion about technology. You're not a luddite for thinking so. Think for yourself.
Edit: Thought you were the same person. Comment still applies regarding conformity.
(Related question: Given how frequently the answer that Google sends me to is a Reddit comment that's been deleted... Has anyone gotten data on whether high-value contributors were disproportionately pissed off by Reddit, enough to throw away their legacy of contributions? The people who care the most, care the most? That's something we might guess if we thought through it ahead of time, was but I still surprised to see what looked like that in action.)
note that this terminology doesnt directly say if openai is training on reddit data. there's a difference between relying on reddit as a search engine (retrieved on demand) or as a pretrain corpus. this language leaves some wiggle room
either way. there has never been a better time to abandon reddit so ... see you all on Lemmy and here.
Seems more like an empty gesture than anything principled.
Junior ML engineers can graft on PR-friendly "AI safety" filters all day, but someone knows what evil lurks in the heart of gigabytes of floating point numbers.
I don't want one company to have the monopoly to train an AI on everything that me and all my friends post on a particular platform. If we're going to decide that training AIs on scraped data is fine, then everyone should have equal access to the dataset. Otherwise it's just a massive data grab and a massive transfer of power to whoever wins this data race, enacted by some platform owners hoping to monetize their users even more
my expectations are that it's tiny and these news outlets are playing up emotional headlines. saying "may go down as well as" without at least ballpark numbers just invites personal bias to fill in the blanks.
IMO: Renegotiating for "public approval" will always be something computing organizations have to engage with when making advances in tech that's downstream of publicly created records, knowledge, etc.
(That being said, I personally think this announcement is overall a win for healthier data flow, in particular because these kind of deals make the value of data explicit.)
"knowing" you are content, and being told straight away might lead to different reactions.
That's why HN is now the only public internet site that I participate in. I should quit even this, but I just can't. Addiction is a terrible thing.
Looks like ChatGPT will be able to dynamically query the Reddit Data APIs and retrieve new info, sort of a la Grok/Twitter. Interesting, and seems quite useful.
In the long term I can see reddit using AI to replace user generated content to capture paying customers.
I'm not exactly sure how they seek to convert the current marketing to some kind of paid solution. I can guess they may do some kind of prioritisation of content if the submitter is a paying customer. AI can satisfy most users so brands may pay to play.
[1] https://blog.google/inside-google/company-announcements/expa... [2] https://openai.com/index/openai-and-reddit-partnership/
Hearing this line from OpenAI sounds like the death knell for the Open Web.
First off - is Reddit really "the open web" or just a part of it? It's absolutely privately owned, there should be no confusion mentioning "Reddit" and "open" in the same breath.
Secondly - if keeping the internet open is crucial, why does it neccessitate incorporating open content into a commercial product? Is it not fair to worry that this is an attempt to commoditize the free transfer of information?
Truly, what a braindead piece of marketing copy to write. Sam ought to should be ashamed of the monster he's created.
What does the phrase "open web" really mean? If I search for it the best answer is this: https://en.wikipedia.org/wiki/Web_standards - Reddit meets this as I can use reddit in a standards compliant web browser. Their auth and api's are theirs to do as they please including ignoring open auth and api standards.
To me the open web is similar to an OS - I can run proprietary programs which speak proprietary protocols along side open source programs and protocols. The platform is "open" - the applications might not be. Buyer beware.
I know everyone has complained about OpenAI misunderstanding their own nom-de-plume, but this extends to how they view the world. It's conceptual and literal openwashing, and you don't need any OSI definitions to understand why it's wrong.
I wonder what this means? Will OpenAI be investing research and engineering into creating models that are optimised to create ads that lead to high engagement? Is this a going to be a new revenu stream model for OpenAI?
but if you use the text-to-speech reader it reads "Keeping the internet open is crucial, and part of being open means public Reddit content... " (emphasis mine)
Surely if you are making a genAI bot for medical/legal applications the likelihood of hallucinations/misinformation is way higher if Reddit posts are included than if it were solely trained on official medical/legal documents