Reddit is OpenAI’s moat
cyberdemon.org
cyberdemon.org
https://www.semianalysis.com/p/google-we-have-no-moat-and-ne...
Original author is an ML researcher, and the crux of his argument is that most weights in a LLM are significantly overdetermined. Once you have ingested several terabytes of natural language, you know how to generate natural language.
The remaining misses are facts that it has never seen, usually because they are so obvious that nobody thinks to write them down explicitly. And so more training data doesn't necessarily help LLM performance, unless you're either ingesting either extremely basic facts that are so obvious most adult discourse overlooks them, or you're ingesting expert knowledge that's highly specific and only discussed in a few forums. Reddit data could perhaps help with the latter, but a.) Reddit is usually not the place to go for expert discourse and b.) there are other better sources of data for it. You'd usually be better off training on trade publications, scientific journals, or fandom than Reddit.
Also the LLaMa/RedPajama approach makes this stupidly easy, because you can pass around patch sets to the model trained on specific mini-corpora, and then update the weights appropriately. Hence why the author of the Google memo believes neither Google nor OpenAI has a viable moat.
The “patches” he refers to are LoRA and are treated as deus ex machina. They’re not, ex. playing with Stable Diffusion we can see they’re additive but they’re not nearly as good as training the original model on the data.
(Disclaimer: Googler)
How are you defining "nearly as good", can you be more specific?
It's obvious full fine-tuning > [PEFT tuning] but to my understanding the gap isn't that significant, as reported in various papers. (specifically with respects to language models, I'm not familiar with diffusion models).
Its not expert research but Reddit can be used to find probably-real-first-hand-experience on a given subject. Which is enough to prefix "reddit" to a lot of google searches. It just needs to be a marginal improvement to Google blogspam for it to have some intrinsic value.
Blogs are fine, I don't mind reading them.
The classic recipe that is 95% down the page with every possible word you could google in the preceding 95% in the form of a fake story.
Bold claim.
There are plenty of engineers and doctors on reddit answering nuanced questions that arent talked about in any journals.
That is what is missing. The absurdly specific stuff that reddit gets. Sure you can answer with nonsense that sounds realistic, but you could also answer with the exact text that solved the problem in the past.
Its the difference between asking the question "How do I save money at the grocery store?" and getting generic answers like "Make a budget" and getting specific answers like "Buy fresh rather than processed".
There are also plenty of people answering nuanced questions with complete BS.
It's not that bold to claim that as when taken as a whole the signal-noise is worse on Reddit than in a publication.
But you get very different kinds of knowledge from each of those. They are not really comparable.
You won't find any useful explanation about how to solder surface electronic components in a PCB in a book, and you won't find a deep enough explanation about how to dimension an uncoupling capacitor for an amplifier in Reddit.
If I ask an LLM to write me an essay in the style of a caveman, the LLM simulates a redditor simulating a fictional caveman (who speaks english but with short words and bad grammar). After all, that's what was in the training data.
If I ask the same LLM to write me an essay in the style of a Harvard professor - who's to say the same thing doesn't happen?
So I can see how, when training a model, you might want it to know what's really written by a Harvard professor, and what merely professes to be.
There are lots of really good subreddits that have expert feedback, but I think the issue is that going through all that and separating the wheat from the chaff is going to be a much more involved process than feeding in those other sources of expert opinion.
more than anything, AI/ML/DL enables people to make claims that cannot be refuted by any application of the scientific method. it's literally "not even wrong" in exactly the same way as people claim string theory is.
Obviously you need a page of fine print to make a claim strictly falsifiable, but complaining about the absence of fine print in casual discussion is absurdly uncharitable unless you have reason to believe that agreeable fine print couldn't be drawn up and I'm 99% sure that in this case you have so such reason.
are you a string theorist? are you sure that you understand the claims made by string theory and how falsifiable they are?
conversely
>fine print couldn't be drawn up and I'm 99% sure...
where does your 99% surety derive from?
As someone trained both in mathematical physics and DL, i'm here to burst your bubble: the falsifiability of both disciplines is exacty the same because they're both premised on knowledged derived of models with an astronomical number of free parameters.
just to be clear: i'm not poopooing LLN here, i.e. the models themselves are fine, i'm talking about absurd meta-claims (like ops about generating language) derived of observing many such models.
Maybe
> we don't know what discrete jump in model architecture might come next
This is true, but as long as these models are based on creating sentences based on what sentences that it's seen have looked like, and not on fetching and understanding verifiable facts, these services will do more harm than good overall.
If anything, these models need two sources of training data:
1. The standard language model (as it is now) to be able to generate and process queries and provide understandable answers
2. A database of verifiable factual information that it can query in order to prevent it from completely hallucinating information and then asserting that it is verifiable factual information when asked [0].
Until we can solve the AI hallucination problem, these systems are going to require users to be much more careful with information they're given than most people can manage right now.
[0] https://www.cbsnews.com/news/lawyer-chatgpt-court-filing-avi...
This is exactly where I think Reddit's value lies. I disagree, though, that people don't go to Reddit for it. Here are a couple recent queries where I appended reddit to my search query: * best places to visit from London * best mattress sold in UK * should I bring king mattress from US to UK
(Can you tell I just moved to London?)
Why do people go to Reddit? Because it's guaranteed (at least for now) that the answer came from humans who are not trying to sell you something. They may be wrong but... it's opinions.
No no no no! This is not guaranteed AT ALL on Reddit! Why do you think Redditors are constantly accusing each other of being shills? Because Reddit users are often corporate representatives doing "native marketing" or whatever they call it, and they are not upfront about it, because they are preying on the naivete of people like you!
Seeing the belief expressed that Redditors aren't trying to sell you something is.. distressing to me. Half of Reddit is some kind of ad, if not for a physical product it's for a political party or cause, or some celebrity's personal brand.
Please have a little cynicism online.. please, for the sake of everyone..
What it also ignored (along with some of the comments here) is data ranking. Google didn't just build a search engine by crawling more of the web -- many search engines before it had already done that. Google managed to rank what's relevant and what isn't. Relevancy is hard. Similarly, not all scientific publications are ranked equally. Or for that matter, even publications with a lot of peer reviews or citations can become obsolete through new discoveries.
Reddit's data has value in that it can fill in a lot of the gaps left by more qualitative sources and furthermore the data is user-ranked by a trusted community. This also has implications for specialised querying, for example training on just r/fitness could be fairly useful for that community.
As a side note, other valuable data stores are not just text but voice/video as well. YouTube and podcast transcripts are readily available, for example to Google. Data and ranking is valuable all over again.
- The entire corpus of data the community had curated over the last XX years
- The "goodwill" that these platforms had developed towards third party developers in allowing developers to work with their data
- Potentially large amounts of traffic that would normally come to their sites via Google (e.g. site:reddit.com), that is now available instantly (and customized) via ChatGPT
Despite Reddit's probably closer connections to OpenAI than other startups through Y-Combinator and Sam Altman, I wonder how keen they are to actually work with a company that potentially destroyed a ton of their value, right before they were ready to IPO.
A lot of commenters here on HN have already pointed that out already too. Unless they build a literal brick wall (paywall) around the site, that data can and will get scraped if the intention is to use for a model.
You could get it down to a science where you only scrape any new data whenever you train the next model.
[0]: https://old.reddit.com/r/reddit/comments/145bram/addressing_...
How would that work from a legal perspective, though?
Let's say there's no paywall and Reddit's terms of use disallow unauthorized commercial use of their data. Wouldn't that be a violation of Reddit's terms and liable to some legal procedure?
I think there are probably issues to address with scraping it blindly:
- Can Reddit imprint its data somehow? A watermark?
- Can Reddit prove that certain type of information appeared on Reddit first and thus that serves as proof its data was used without authorization?
If OpenAI can't work around this, I'm not sure they would be willing to cross any lines in terms of copyright, they've already done it with ChatGPT and I am guessing rules are only going to get stricter on this topic.
To me this feels like its opening up the door for the elimination of copyright as any algorithmic layer interjected between scrapped data and end users could claim to be "inspired".
It won't, with the LinkedIn vs HiQ precedent.
Does Google have a special agreement with Reddit (and all other sites?) or is it legally "fair use" to reproduce web pages that are available freely online?
The only question is copyright, but I find it hard to argue that LLM training is not sufficiently transformative in 99% of cases.
Reddit on mobile browser is a case study of insane dark patterns
Click to sort comments while not logged in? A popup appears asking to log in, with no close button. You have to click out of the box, but that’s not easily apparent
View an 18+ subreddit? Let them browse for 30 seconds, then ask they log in
Visit the site? Ask to login or use their app.
At this point, I’m more motivated to do anything but use their app if they’re this hellbent on getting me to download it.
The thing is, to me, reddit would never be a viable enterprise without the volunteers moderation. If reddit had to pay for that they’d never ipo. And if they think the moderators will stay and be exploited they’re wrong. Thus if they go through with this they’ll fail as a public company. There is no there where they going.
I'm convinced that this would have happened with or without OpenAI, especially with the mirage of an IPO on the horizon. Controlling the client to show ads and siphon data is just too valuable. Maybe the OpenAI thing pushed them to speed up the process.
The irony if lowtax and :10bux: ends up being right all along about how to build a community...
- When they started, I believe they boasted creating fake accounts to mimic engagement to help grow community
- It is painfully obvious the amount of political, gov sponsored, and corporate astro turfing that “gets through”. Both human and bot comment farms are real and have no doubt been artificially bolstering ideas and content for years
As I see it…
They have enabled this / looked the other away at this behavior forever in exchange for engagement. The super high tech AI chat bots are probably going to be welcomed for their clever content (hence why CEO could care less about its users anymore).
Their real value has always been having a controlling voice and being able to push viral ideas. The data thing is all hype noise IMO. It was never going to be a serious part of their IPO. Big messy gross company.
IMO, the future will be using LLMs with live search results - which of course will probably require a funding model better than the terrible Display Ads setup we have for such content now. So best of luck to Reddit on that front - I hope your ipo fails
N of 1, but I vastly prefer the Google Generative Results over chatGPT. Quality may not always be the same, but a “chat” seems like an awkward metaphor for finding info, and of course, chat GPT has no links to content when I’m worried about hallucinations.
I’m true google-search fashion, GenSearch has a lot less deep-in-the-weeds technical answers and will push you to simpler results. Eg. If you want to know what chemicals have a similar absorption properties to methane… you’re better off with ChatGPT or traditional search.
Main reason such sites came into existence was there was too much info on the Internet.
These sites where attempts at simplifying the numerous websites and blogs ppl had to manually discover and track.
They have meandered around that problem, and got totally distracted by all kinds of other problems (many self created).
Sure it will - asymptotically.
But it took 18 years to grow Reddit. Even if a solid alternative emerges in a fraction of that time -- that's still a multi-year gap. Plus we're at a significant knowledge deficit (for us lowly humans, never mind LLMs) if Reddit's archives don't re-emerge sometime soon.
(Google search has been braindead for some time now -- so we can leave that source out of the equation).
SO had publicly available, no-auth-required data dumps. This makes it difficult for them to know who is using their data. However, this surely isn't the case for Reddit who offered only API endpoints for this content, and I'm guessing you couldn't use .json to get the whole site (rate limits, etc). I wouldn't be inclined to believe that Reddit would miss a new major API user.
This is purely speculation disregarding Hanlon's razor, but I'm thinking that the API pricing comes down to killing two birds with one stone.
* Sama got to train his LLM on Reddit and some best-of-the-internet content there such as r/bodyweightfitness, informed discussions on niche topics etc for free. The catch-up players face prohibitive pricing.
* Third-party apps get killed, bringing the UX to Reddit's control. I think this is more important to Reddit than ad revenue, as they could've simply built an SDK for probably less than this PR nightmare will cost us.
* Their new development platform, however, hints at the Reddit app supporting serving "redditor-made apps" which "can be seamlessly reused between communities".
The description (and the idea to have apps in your app) weirdly reminds me of WeChat apps, and given that Tencent is a major shareholder in Reddit, I would consider the possibility that apps are something they're pushing. That idea has no chance of success without the UX being completely in Reddit's hands, even then it's questionable how it would work on Reddit.
Spez couldn't actually be that detached from Reddit?
Snoop Dogg owns more of reddit than Tencent. The Tencent investment is very overblown.
Basic math days Snoop owns twice as much.
Further, you could use the Reddit API to injest the full firehose of all site data in real time without violating rate limits. This is [one of the ways] how Pushift made their datasets.
Reddit was a lot more open that any other site approaching their size!
Who is "us"?
- next step is to stop third-party apps that generated these data
- then they let the moderators show their power
I’m not sure if people care about a CEO being exposed as a liar nowadays, but maybe some former Yahoo managers have another idea on how to destroy more value.
Specifically the ones that ended up in charge of tumblr, so they can suggest "ban adult content on a platform famous for its adult content"
I'm not so sure that's the case, for two reasons:
1) I have not stopped appending "reddit" to my searches or stopped visiting StackOverflow or other Stack* sites. ChatGPT is simply now an additional tool, and while there's some overlap in use cases there's also plenty I can do with *GPT that I couldn't with other sites.
2) To the extent that a training corpus relies on these (or any other) data sources, OpenAI has just made them much much more valuable! It's sort of like a mining company discovering that someone has an extremely valuable use case for decades of the mining company's tailings. This "someone" may have been allowed to haul away some of it without payment to the mining company, but now they know its value, now they can build a very lucrative business selling access to what remains, and what will be generates in the future.
That's all highly simplified though. It's of course much more complex than this, much more complex than can be captured in an HN discussion, but we can explore the outlines a bit, and even disagreement will reveal more and more of it.*
Is it?
When I worked in networking, and later web dev work I found Reddit to be a TERRIBLE place for Q & A type situations.
Answers on Reddit are often skewed by truthy answers from people with limited perspective in the industry who are surprisingly sort of militant about a given topic.
For example the networking topics often were “small shop” focused and users unaware of very common enterprise level practices, and when they saw someone mention it they would react really poorly. The result was advice tailored to small shops that have limited staff, so you got ideas that “work” but would perform poorly at scale (at best) and at times could present security risks, miss opportunities where more performance was needed, better solutions were available with a sizable staff and planning. They weren’t wrong, but the answers were skewed.
Subreddits are weird, they might be about a topic but the community often moves that topic/ has a specific pov within that topic… people without those views tend to abandon those subs and those who remain aren’t necessarily “right”. Reddit isn’t stack overflow where being correct is at least on the surface valued where we can run code and see if it works. Top answer on Reddit could be wrong and anyone taking issue with it may simply not answer. Heck there are whole subreddits dedicated to wallowing in misinformation…
Now I don’t know if Q&A is the start and end of LLM training (not my area of expertise at all, I am open to the possibility that what I’m talking about isn’t a problem at all) but Reddit as a source makes me wonder what the results would be.
It's not great for business stuff because there are way more people using e.g. AWS for their hobby project than running multi-million $ businesses on it.
What are the alternatives? Facebook and Discord are walled gardens. Quora is a shithole. The rest of the internet is blogspam and significantly more dubious than your average Reddit comment!
I could spend hours just neutrally asking someone questions about a particular view they hold, why they hold it, etc. I think you gain a lot more insight into how people think this way than through an online forum.
I do.
Or your desire for an umbrella that doesn't break in the wind, or the best public estimate of when Starship will launch, or any of the myriad other desires that cause people to append "reddit" to their searches.
The only consistently good advice I see on Reddit is "call a local expert/call a lawyer." Anything else is almost certainly abject nonsense.
In Twitter you can follow people whose content you find useful. If you're a Blender user, for example, you can follow 5 Blender tech artists and your feed will be now full of cool tips and art.
With friends, with more or less the same skillset and hobbies as you have, you can ask them whatever you wanna ask. Heck, you'll learn a lot from them without even asking!
I'd just avoid using Reddit or forums to decide if I'm gonna use or buy something.
For example, if you ask what's the best OS on the internet, people will say Linux. Best DAW? Reaper. Best game engine? Godot. Best headphones? whatever is new and costs 100$ in Amazon.
Meanwhile, friends who make money and are in the industry and have businesses just prefer Windows, FL Studio and Unity. They aren't on the Reddit saying that the software or hardware they use - they just use them and make money. :)
The point you make about "professionals don't have time to use Reddit" definitely applies though. /r/synthesizers for example is a great sub if you want to spend thousands on hardware. Not so good for making real music! Although I do feel like there are a lot of Unity game devs on there. Maybe you just need the right community.
Walled garden is marketing speak. Don't use it. Use curated/controlled/whatever described the closed nature.
It spins what would otherwise be a negative term into one that gives images of delicious fruits, vegetables and beautiful flowers.
It's actually funny how much this is true. They used to be a quality source of info, but these days the default option is to show "all related" so you don't even get direct answers to the question you searched for, ads from top to bottom (and even somewhere in between) and just hordes of bots answering nonsense. At least they made it official with Chatgpt integration now lol.
I'm curious to what your take is on Chat-GPT's response to the same sorts of questions?
I also work in a specialist field, and I've come to understand that the instinct on of a particular question is going to give me a bad answer on Reddit or SO, is also finely tuned to about when GPT either starts hallucinating, or providing bad info.
But for most things outside of that Reddit is by far the best online source that actually represents answers you would get from real people. Questions like what is the nightlife like in x city don't have a single true answer and thrive due to Reddit's the diversity of thought.
I dont think we need a llm that only responds with "Your question was marked as a duplicate"
Training in this sense is about how people use language not what is the right answer. Among other sources - the exact thing you describe is a contributor to some of the things within the (in my opinion misnamed) internal taxonomy of 'hallucinations'.
Right or wrong - reddit is human beings using (mostly english) language to speak to each other, and so if you want to train an ML model to do that - its a great source.
- the voting system is about what people want the truth to be, not the truth.
- the users can be very easily gamed in comments to vote one way or another just by the initial voting being negative or positive (aka most just vote with the trend)
- the opinions all come from a bias of urban and major metro area people. It's painfully obvious they don't understand any mindset outside of this one.
- it wasn't always this way.
I agree with you, the quality went down as the quantity went up, and it may be possible to use the comments, but using the vote data as a signal of truth is going to produce a terribly narrow minded AI.
Probably the best example is legal type advice.
Personal example: The whole situation with internet archive and their e-book lending process. I was vocal that they didn’t have a leg to stand on and would lose. The comment was just about the legal situation.
I was downvoted to hell for that opinion and identified as some sort of book publisher advocate/ internet archive hater. As was anyone else predicting a loss or just raising questions.
Personally I love the internet archive, but I didn’t think they had a chance at winning.
But go by the comments and it was a slam dunk win because of some idea of “freedom” the comments and voters had. And yet they lost the lawsuit…
Interestingly, here on HN the response seem to be a generally more balanced. That’s not an endorsement of HN all the time, but it was an interesting contrast.
Yes, I know reddit can be valuable for niche and non-contentious topics but such examples are becoming increasingly rare. Snark and combativeness pervades the site and it seems to be a function of the voting system and monkey-see, monkey-do from how commenters are rewarded in the default subreddits for this kind of behavior.
I dread AI models being taught to comment like a redditor.
Equally said of HN. Folks with a few years of dev or startup experience asserting this or that all over the place. Im sure less experienced folks read these "truthy" answers and take them as gospel in some cases.
That's kinda sorta exactly how ChatGPT answers stuff though, no?
I don't really want the corpus of Reddit data influencing the output of a generative AI model...because it's Reddit, after all...
Even though I am pretty sure it is already included in the training dataset already...
I could be wrong, though...
But the Reinforcement Learning from Human Feedback (RLHF) is also one of the key tools to getting useful outputs.
the opinions i run into real life can be very different then with people in the real world. the communities online are made up of the kinds of people who spend their time online, and the content you see on reddit is generally from the people who spend enough time on reddit that they want to browse new. these arent average people. only a minority professionals are actually engaged in reddit. even those ive seen run off because they don't agree with the acceptable opinions
Hegemony, perhaps.
I think you messed that sentence up, but I get what you're trying to say....
And I think it's this. The opinions you get from people 'IRL' are not apt to be as strong as the ones online, and or will run into the regency bias.
For example, it's very unlikely you'll actually meet someone that has used 10 different coffee makers because they wanted to see which one was best. Online on some subreddit, you're very likely to meet someone who has done exactly that. Of course those people with strong opinions are the ones that are apt to post most online.
So who's option is wrong? Neither. That's why they are opinions.
No, I've noticed what he's talking about, and it's not the strength of the opinion, it's what the opinion is. Reddit has a moral system that's completely misaligned with real world morality, where having a child or being autistic makes you a bad person, even if you didn't do anything wrong. You can find some really weird takes on r/AITA, which ironically points out that the subreddit has a fucked up sense of morality in its highest rated post.
Not sure if you've noticed but younger generations that go out and do things are big into having children. Now, some groups on reddit are more extreme on that, but the general trend of Americans at least is to tell the act of having kids to screw off.
---
But to answer your question, it's not Reddit that has this moral system, it's social media in general that has a moral system that does not match reality (kinda). The loudest idiots tend to get voted up, moderates disappear in the bulk of posts. Binary voting systems on sites tend to amplify this. Content suggestion systems tend to lift up contentious posts for engagement. Welcome to the internet.
But coming back to (kinda). This is becoming reality. Behavior IRL affects behavior online. Online behavior affects IRL behavior. People don't talk to their neighbors these days in most places. Communities are spread all over the earth.
not with a bang, but with a "this. take my upvote my good sir"
horrific
Reddit in particular seems like a terrible set of training data. Pretty much any opinion is going to have a counter-opinion somewhere in the thread. Synthesizing multiple comments together requires nuance based on circumstances that I don't think I would trust to an automated process- which heuristics were applied, what was the reputation of the people on either side, etc.
Hell, I don't even stop at the first recipe I find if I'm looking for something new to cook for dinner. I look at a couple variations on a dish first.
LLMs aren't automatically disincentivized from training on blogspam so they're not going to avoid it either.
ChatGPT is great for distilling that experience down to just a simple recipe for whatever it is I'm looking for.
Tangentially: props to AnyList for their amazing plugin that scrapes recipes off of sites like this and stores them in an easy-to-use format.
My prediction: By 2025, you'll be able to purchase sponsored sentences and sponsored paragraphs in ChatGPT's (or whoever's on top) output.
Search, on the other hand, doesn't answer your question.
It's devilishly difficult to get citations in there, you're not going to do it with langchain, but its possible (c.f. Bing).
Python x langchain x LLMs makes it very easy to create demos so there's been an initial influx of meh stuff, I'm very excited for 6 months from now.
I've used for rather obscure queries and liked the summary the AI wrote as well as the links to dive deeper. I imagine that's what AI search will look like across all providers before long.
I threw in a question Wikipedia does kinda answer but leaves a lot of details off, and it enumerated a set of possible answers, with my favored option linking into a forum where it looks like less than a dozen people post, but with people that tried all variations of it and know all of the details.
DDG and Google would never show me that site (yeah, I've tried).
If you're good at googling the flow is: Ask the question > Clock which result isnt spam and click it > Figure out how to dismiss the cookies gate without accepting the cookies > Dismiss the google login box > Dismiss the popover pushing you to install an app > Scroll the page or ctrl+F to find the answer
With ChatGPT it's just type your question and your answer is appearing right away.
Of course I understand ChatGPT shouldn't be used for this because it will lie to you and make things up. But I am saying that's why you'll see people who don't know that preferring it over Google.
Build a product that's more convenient, and people will use it. Google is so full of shit now, and even answer boxes are below 4 ads, it's just way more comfortable and efficient to ask chatgpt.
This morning I wanted to know how many calories there are in a breaded chicken breast. Chatgpt told me in 3 seconds after asking. Google would have been way more time. (sidenote: I also hate google, so I'd much rather use chatgpt anyway)
And I’m not talking about fringe Q Anon type stuff: the other day I was looking up the specifics of China’s “Blue Sky Initiative” climate change policies and the only thing Goog/DDG would show me (despite several attempts at rephrasing the query) was Western industry think tanks bellyaching about how the policies effect profits. It took me a good ten minutes of refining my search before I got an English translation of the actual policy bullet points.
I can’t imagine this being such a PITA on 2013 era Google.
However, I’ve seen plenty of bad information in google answer boxes too. And finding it in actual search results is going to be way more time. It’s not a life and death question.
I'm learning typescript and ran into a weird typescript construct the other day, I threw it into chatgpt and asked "what is this" and it explained it to me. I'm not entirely sure if pasting in a bunch of braces and parentheses into google would find me the same result.
AI offers none of the above. Ask the same question twice in a row, maybe you'll get a different answer, maybe you won't. You won't know which of the two answers are hallucinations, which are true factoids it trained on, and which are bogus things it had been trained on. There's literally no providence for the results- no trail, no references, nothing, because it's really a sophisticated game of "Whose Line" where everything's made up and nothing matters.
I don't mean this as a way of thumbing my nose at people who don't know this stuff, but rather, to point out what a massive failure the completely nonexistent consumer education surrounding ML products has been.
The appeal of a conversational response was too great to wait for a solution to the problem of training data quality.
Search requires work on the part of the user to distinguish between good links and bad. AI is an oracle just tells you what you're looking for.
Now you and I might think this is a terrible way to evaluate the veracity of information. But think about the new generation of mobile-native users who were raised on simplistic discourse in tweet-length messages, and would rather watch a 1-min video on a topic than scan search results for 30 seconds.
For this group, searching for information where more than 2 clicks is required is going to be a "too complicated", anda "bad user experience".
LLMs are oracles that arrange words in a probabilistic order that are grammatically correct and may be factually correct. Unfortunately there's no way to evaluate the probability of confabulation with any of the LLM chat bots. The distribution of occurrences confabulation is also not regular or predictable nor is it fixed. So you can't ever say "ChatGPT is bad at X" because it can be bad today, good tomorrow, then bad the next day.
If you want truth, you don't want language, you want references to reviewed work. You also want things like 'show your work' chains of though. These are really different things.
If I tell GPT "make up a story" I don't want it coming back and saying, sorry I can only tell the truth.
On the other hand, some random people's opinions in a Reddit community with -apparently- no further agenda seem somewhat more honest.
Not that the answer is better but it gives you new data points in your search.
Basically, it's not one or the other, you can use both tools and that's probably why it makes sense to include Reddit in AI models (which do this job for you automatically)
Reddit has a lot of other private information though. The entire corpus of moderated / censored comments is the internet's best alignment dataset. The real upvote/downvote numbers make for perfect RLHF.
Hell, Reddit could remove ads and replace them with paid RLHFing bots. Charge OpenAI money to selectively deploy AI bots to certain QA subreddits. All AI generated comments get deleted within a week. That way it doesn't lead to comment spam, but the bot still gets amazing upvote/downvote feedback and a ton of direct comment-reply feedback. 3rd party apps can't 'ad block' it either, since only Reddit knows which comments are fake.
(psst, reddit, if you're gonna steal my idea, I want to be paid for it)
That said, I'm not totally sold on the idea that Reddit has truly unique and valuable data that's a cut above what can be found elsewhere. After all, several of the so-called "glitch tokens" in GPT's tokenizer are Reddit user names that occur over and over in threads of people just counting. See: https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldm...
I don't think I have a refutal, honestly. OpenAI aside, Reddit clearly has realized that their data is valuable, perhaps more valuable than their ads business. This is why my piece is frankly very speculative: the strategy aligns really well and the financial interests are there, but I wish I had a great argument why my theory is more likely than simply Reddit trying to capture the value of their data.
They’re cutting off third party apps to increase their control over the user experience, to show more ads.
The relationship with OpenAI might also be a thing, but I think it’s the ad numbers they need to juice for their IPO.
Of course, I think another good explanation is that, OpenAI aside, Reddit simply wants to get paid for their data, which they now realize is valuable. Scraping aside, that might be the real goal here.
The fact is that many, many internet businesses today not make sense. Many of them are bleeding money with no possible path to profitability ahead. Mostly due to the fact that people's expectations are very different and I would argue even skewed. Clearly hosting Reddit is expensive, but who really pays for it? People don't like ads, don't like to pay for services, don't like to be reminded that they can provide donations (Wikipedia). What else is there? Hoping the customers have the suddenly conscience to fund the product?
I foresee many internet companies dying over the next few years with no real replacement to crop up. But people always forget and history repeats itself; so maybe we will just see an endless reincarnations of similarly failed products.
I think it is about forcing people to watch ads, particularly keeping reddit safe for "dark patterns".
It's an awful experience to follow a link from a search engine for reddit because reddit systematically injects ads into results and discussion to try to trick you into clicking them and injects even more blended links to irrelevant discussions to increase your chances of getting confused and clicking on an ad accidentally.
Third party clients destroy all that.
People who want to harvest text from reddit can just do it on the web. (Pro tip: ignore the API and implement your own 'IPA' that just works like a web browser; APIs are almost always nerfed in some way; if you have to crawl 20 web sites odds are a generic web crawler will work for 19 of them; if you are using APIs you will have to implement something different for all of them, maybe 3 out of 20 will have some feature in the API such as a complex and poorly documented authentication process that will take a few hours.)
APIs are not a gift, they usually are an attempt to take access away. (Considering Hacker News, there is no API to make a post or get your upvotes. It's like 20 lines of Python to code up an "IPA" for either against the web interface.)
Also, if reddit is in any dataset of these models then it's no wonder that I'm not intrigued by artificial """intelligence""" at all.
"They could buy out more 3rd party clients, which would somewhat appease the community. This would be a terrible move IPO-wise."
Why would this be a terrible move IPO wise? If Reddit has two major apps, one for casual users and one for power-users, and it has enough control to eventually put (tasteful) ads in both, how is that bad for their IPO? I don't think it's likely, given how much goodwill they've already burned with the community, but from a pure IPO perspective it seems like a solid go forward strategy
Or if that's too much work for you, just download one of them many archives that are already freely available online.
> just not recent enough
thats a very big deal for a lot of the content you find on reddit which users might want to get answers from from an AI.
https://tracker.archiveteam.org/reddit/
I guess either OP doesn't know about them or he assumes that AI researchers won't think of using them.
But like you said, the models would ideally be trained on many other data sources as well.
> Most of Reddit’s current data has been scraped anyway, so the game is to protect Reddit’s data going forward.
But yes, Pushshift archives of all posts and comments until the recent ban [1] are freely available for download [2]
[1] https://old.reddit.com/r/pushshift/comments/135tdl2/a_respon... The ban was followed by allowing the parent non-profit of Pushshift (Network Contagion Research Institute) to use the API provided access is restricted to a use-case Reddit has care for: mod tools. Reddit hasnt replaced those with its own just yet. The rest of us are shut out.
The intention of the API change is to kill third-party app users who cannot be as effectively monetized as first-party app users.
Wikipedia, in itself, accounts for a good 50% of the positive image i have of the internet.
They could come as a white knight, restoring the API access but for apps only, conveniently blocking competitors while satisfying the users.
How many billions will AI be worth in the future? What's the cost of the potential reputation damage for having a edgy brand? All that matters is the operation can be expected to be profitable.
Better: what would be the synergies between the xbox audience and the reddit audience? The "degenerates" are clients like any others, maybe also with more disposable income.
One might argue these communities are too valuable to generative AI to allow them to flail under current substandard management.
There are many synergies that aren't properly considered by a naive take.
I use site:reddit.com with most of my queries. I now prefer bing to google. Bing market share is growing.
Coughing up a few B to avoid reddit going on with an IPO would be a very strategic move on many fronts.
Also, Microsoft doesn't have a generic social network to mine data from.
They may prefer to stick to the professional world, but then acquiring a few high value sites just to close their API would be another move.
How much $$$ do you think it would take for YC to say "hell yes!" to put HN under Microsoft control?
How valuable would it be to add a few conditions like a very generous giveaway of azure credits to HN companies in return of exclusivity (no more S3 or google cloud for the new batch!)
They already have LinkedIn, which is far more valuable than HN for this purpose.
There's a lot of value for keeping in touch with the zeitgeist (ex: sqlite innovations, llama etc) and finding new trends on HN
There's about 0 value from using linkedin.
It's quite naive if someone think all the major tech big boys who have at least moderate ambitions about AI haven't already at least archived Wikipedia, Reddit and StackOverflow.
No, no... jeez, you pervs... not the content.
The upvote/downvote data for every image will probably train any generative AI on exactly what makes a good porn video/pic - and we all know sex sells. The dirty comments, and etc.
On top of that, they'd make really useful leverage if you can connect any one of these accounts to their real world counterparts. I bet an AI with a large enough dataset could make short work of that.
Blackmailing people who participated in /r/jailbait? I would bump that up a 0 or two.
Sorry, just thinking about how an AI would go about acquiring resources.
Please tell me that admin guy wasn't part of that.
Maybe they want to wait for it to IPO and then drop 90% instead.
Agree with the viewpoint. Hacker News is heavily moderated in a non-user-friendly way which results in some comments being nonsensical after moderation editorials and also disturbs searching for content by a title one remembers.
Sure the API made things convenient, and scraping content will be a bit of an arms race, but scraping public postings that don't require a login to view still seems like a bridge over the moat. It's tempting to refer to the recent victory of LinkedIn over HiQ, but there's an important distinction in that ruling: It pertained to use of logged in accounts and did not explicitly overturn the prior ruling that public-facing data was fair game.
For now it's all still unknown territory. I would be skeptical of anyone who adamantly affirms a position that scraping for training is allowed/not-allowed because it is not yet settled law in the US where these companies are headquartered or the rest of the world where they operate.
This triggered me to review a thought I had.
VC's are very keen with attending board meetings without being directors, retaining the right to appoint a director without carrying that out, and exerting control thought other means. I think some tightening of the concept of shadow directors is in order. Shadow directors are supposed to have the same legal responsibilities as properly appointed directors, yet they currently carry all the power without the reponsibilty
I don't get what the author is trying to say here.
If the deadline was extended, why would it "get rid of 3rd party apps"? I assume because some (Apollo, RIF) said they're going to close anyway even if the deadline was extended? If so, the community will not be "intact".
If the author think the death of 3rd party won't affect the community, then Reddit don't need to extend the deadline.
And then one fine day; some random guy has a local/private LLM trained to browse his facebook groups as him, his account gets deactivated meta; he sues arguing in court about right of using his own tool to mine personal data.
Eventually all B2C AI companies will be just slick interfaces over open models that are too large to run on commodity hardware. B2B AI is in a bit better shape, both because compiling niche business data sets can be expensive, and there won't be giant public data sets produced from their output.
I assembled one of my previous gaming PCs 10 years ago and installing 32 GB of RAM wasn't a problem back then.
But you can't even buy a consumer GPU with 32 GB of VRAM. Data center cards are considerably more expensive.
There's no technical reason to not have 100+GB consumer GPUs today.
Ok, looking closer at the article I see
> The important piece is that it’s easiest for OpenAI to get the data (given that companies with co-investors help each other)
which seems like a pretty weak basis to form a moat (if Reddit IPOs as is their plan, sama’s influence will be not nearly as strong)
To forbid LLMs simply could be a legal paragraph in Reddit’s API terms of service. Good actors would abide, bad actors would still crawl and scrape the HTML. Practically the same situation with the closed API from July forward, but without the drama.
Killing 3rd-party clients seems the far more likely motivation.
Besides, if someone really wanted to get to the data they could just scrape it. Google, Bing & co index it after all. Bit of a pain in the ass but not impossible
It wields reddit as meaningful source of openAI's performance, and then goes down a rabbit hole from there.
OpenAI's moat is the thousands of 'human in the loop' contractors that they hired in south america... for years... Not reddit...
They shut down i.reddit and reddit.compact so theres no optimized mobile web format.
Now they shut down api access to apps. Its as clear as day. I don't think they give a sh*t about AI
I don't think anyone is making any credible guarantee that any AI model is unbiased.
The RefinedWeb / Falcon model has proven this is not true.
This turn of phrase almost always is followed by an argument you don't have to bother reading because it is wrong.
> Comments should get more thoughtful and substantive, not less, as a topic gets more divisive.
I can tell you after over a decade doing this job that the title moderation on HN is probably the single most beneficial thing we do. Without that, the front page would consist of linkbait, sensationalism, and rage. In other words, HN wouldn't exist.
I don't like or dislike any particular title or headline. I dislike the policy of selectively altering an author or publisher's chosen headline in general.
> I can tell you after over a decade doing this job that the title moderation on HN is probably the single most beneficial thing we do. Without that, the front page would consist of linkbait, sensationalism, and rage.
I appreciate the insight. Is the plan to keep this policy in place indefinitely in order to protect us from ourselves? I find it very depressing that a community like HN cannot be trusted to look at an unedited list of headlines. Do you think there is any hope of eventually developing something like a collective immunity (or even resistance) to clickbait? If it can't happen here, what does that say about the prospects for media literacy more broadly?
HN's quality, such as it is, very much includes "selectively altering an author or publisher's chosen headline" - publishers and authors (and headline specialists, at media sites) are incentivized to take advantage of the default dynamics to get attention. It's their job to sex up their titles and our job to knock them down to size, if we can. That takes sustained moderator attention.
That's not to criticize the community. HN's system consists of community, software, and moderation, and the community is by far the most valuable of the three. But all three must be active in order to keep HN from deteriorating (or at least to stave it off - https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...). Moderation's job isn't to determine discussion, it's to serve the community by jiggling the system out of its failure modes when it gets stuck in one of them.
> The answer to any headline with a ? at the end is “no.”
There's a popular meme that says that, but it's not true.
Edit: after reading everyone's objections I think the easiest thing to do is just reverse all of this. Sorry for the trouble!
In the strongest terms, this is unacceptable.
Same point if you mean OpenAI.
Titles on HN aren't the property of the submitter nor of the author of an article (though the author's opinion carries more weight). They're shared pointers and we edit them to be as neutral as possible. Mods do this all the time and have been doing it since HN began 16 years ago. It's primarily because of that practice that HN's front page looks completely different from Reddit's or any other similar forum's, and that's a crucial factor in the identity of the site.
If we didn't have this rule, people could (and would) use the title field to get away with anything. I'm not saying that's what you were up to, just that this is an important rule and bog-standard HN moderation.
Edit: I guess we went too deep in the thread, so we're resorting to edits. that was a rhetorical question. It's sort of like the Twitter community notes feature: making the edit, and its reason clear, builds trust. Otherwise, you just look like you're doing something shady (given that, again, the people in the article literally own HN).
In terms of YOUR opinion about the quality of my article – that is what upvotes are for. I guess you have a much heavier hand into what we're allowed to read than I realized! Again, given your incentives, that's worrying.
One sign of a linkbait title is when it claims a lot more than the article delivers. Usually such articles begin by walking back their own title, since it has served its purpose (getting attention) and is now a handicap (because the text has to deliver on what the title promised). I'm not saying you did this intentionally! But your article follows just this pattern, since by the time you get to the meat of it, the claim has become weaker and in fact has been turned into a question: "what if we thought of Reddit as, functionally, subservient to OpenAI?".
Another thing we often do is replace baity titles with representative language from the article itself. In this case we could do that by using your own question instead of appending a question mark to the original title. That is, we could make the HN title something like "What if we thought of Reddit as subservient to OpenAI?". I think I'll go ahead and do that, since using representative language from the article itself is arguably less of an intervention in this case.
p.s. I just put back your original title. It was a borderline case and not worth arguing over (but maybe the discussion was clarifying or of interest to some readers). Generally when there are strong protests about a title edit we just either roll it back or figure out something that people are happy with.
However, the staff at HN are doing themselves a disservice and legitimizing criticism with this sort of action. If you feel this content is grand and speculative then downvote it or flag it like anyone else. If you feel it's actively damaging or strongly against the ethos of HN then delete it or sticky a comment at the top of the discussion to give your position and encourage people to flag/downvote.
The author makes a fair observation and does speculate as part of their writing, but speculation happens all the time on HN.
Reddit can be a moat for OpenAI even if the intention wasn't deliberately made.
You might be underestimating how much curation/moderation we do of HN's front page. We do a lot. The front page is not generated by votes and flags alone (https://hn.algolia.com/?dateRange=all&page=0&prefix=false&so...). Upvotes and flags are important but the front page is a complex combination of those plus software action plus moderator action. If we didn't have the latter, HN would basically be impossible.
It's a pity—I'd love to manage without moderator action because the latter is tedious and it would be much more fun to write code to do it or just let the upvotes have their way.
You can keep replying - you just need to click on the comment's timestamp to go to its page and the reply box will appear there.
> given that, again, the people in the article literally own HN [...] Again, given your incentives, that's worrying.
What people and what incentives are you talking about here?
> that is what upvotes are for
HN has never operated by upvotes alone. It would be a completely different site if it were (https://hn.algolia.com/?dateRange=all&page=0&prefix=false&so...). HN a curated, moderated site and always has been, and we've always been completely open about this.
This is why I think the editorializing should be made clear: even if your title edits are in good faith – which I totally buy – you want to protect community trust.
It should take only a brief glance at https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que... to decide whether HN is protecting Reddit - the last two weeks have seen a tsunami of threads, almost all negative. Discussion of OpenAI isn't that lopsided but I wouldn't call it positive. sama isn't involved with YC in any way as far as I know. (Edit: I just asked, and it's a bit more complicated than that - he speaks at batch events occasionally, has ownership in past funds like most former employees do, invests in YC startups sometimes, etc.)
I appreciate the reference to good faith! that is kind of you. FWIW, the reason we don't mark titles as edited, or comments as edited when commenters edit them, or any of that kind of thing, is because HN's UI has always been minimal and in particular has always avoided ceremonial or bureaucratic trappings. This is the sort of quality that dies by a thousand cuts if you start adding details, and I don't think that would be the right tradeoff. It's just not in the overall spirit of the site. That doesn't mean we're trying to hide anything—we're always happy to answer questions about anything on HN, and I spend most of my time doing so.
That said, other comments of yours (where you speak about the work it takes to give HN the “HN feel”) give me a thought: HN is less a link board than it is an actual, edited, journal (lack of a better word). I mean, links are submitted by users, and upvotes are used as a signal, but unlike Reddit, HN is edited. That’s a good thing, because clearly we find this site useful.
The article advances a position and argues for it, which is reflected in the original headline formatted as a statement. It doesn't ask a question and then provide answers; the editorialized headline misrepresents the article's content.
EDIT: You've now rather spitefully changed the title to something that's even less representative of the contents—plucking a quote doesn't make it representative—and removes the intended allusion to recent conversations around the Google moat memo.
On the contrary, I changed it to the article's own statement of what it is about. (Note that it happens to be a question - in other words, the mod who put the question mark in the title was on the right track.) More explanation here: https://news.ycombinator.com/item?id=36328562.
Replacing a baity title with representative language from the article itself is just about the most common thing mods do on HN. If you think that HN is worth being a regular reader of, then you're benefiting from this practice whether you feel aggravated by it or not; it's foundational to this community.
I know people love to hate HN's title editing practices - I've been hearing about it for well over 10 years - but of all the things we do on HN this is the one I feel most confident about and am happy to defend until the cows come home. This may seem counterintuitive (how can the best thing be what people love to hate?) but there's a simple reason for that: people notice the cases they dislike and overlook the others.
dang, I see the value in what you're going for, but I think making the edit clear would go a long way. Consider it a feature request.
The new title sucks, but it doesn't matter – the story is off the front page for some reason.
Time will tell if the move will work out for Reddit in the long term.
They haven't taken any steps to stop scrapping. They made access to the data via api extremely expensive. Calling an API is not crawling/scrapping. You can still crawl/scrape, well once/if the mod protest is over. And also stop calling it scrapping, it web crawling.
I also actually wonder about the validity of it as training data. I've done a few experiments with fine tuning models, with a few hundred thousand samples curated from hundreds of millions of threads <https://huggingface.co/winddude/pb_lora_7b_v0.1>. they are interesting, because they end up being so sarcastic. People tend to either be short, or overly opinionated.
At the very least someone would have to do a lot of pre-processing, which would make it a transfomative work anyways.