I don't really want the corpus of Reddit data influencing the output of a generative AI model...because it's Reddit, after all...
Even though I am pretty sure it is already included in the training dataset already...
I could be wrong, though...
I don't really want the corpus of Reddit data influencing the output of a generative AI model...because it's Reddit, after all...
Even though I am pretty sure it is already included in the training dataset already...
I could be wrong, though...
Reddit in particular seems like a terrible set of training data. Pretty much any opinion is going to have a counter-opinion somewhere in the thread. Synthesizing multiple comments together requires nuance based on circumstances that I don't think I would trust to an automated process- which heuristics were applied, what was the reputation of the people on either side, etc.
Hell, I don't even stop at the first recipe I find if I'm looking for something new to cook for dinner. I look at a couple variations on a dish first.
LLMs aren't automatically disincentivized from training on blogspam so they're not going to avoid it either.
ChatGPT is great for distilling that experience down to just a simple recipe for whatever it is I'm looking for.
Tangentially: props to AnyList for their amazing plugin that scrapes recipes off of sites like this and stores them in an easy-to-use format.
My prediction: By 2025, you'll be able to purchase sponsored sentences and sponsored paragraphs in ChatGPT's (or whoever's on top) output.
Search, on the other hand, doesn't answer your question.
It's devilishly difficult to get citations in there, you're not going to do it with langchain, but its possible (c.f. Bing).
Python x langchain x LLMs makes it very easy to create demos so there's been an initial influx of meh stuff, I'm very excited for 6 months from now.
I've used for rather obscure queries and liked the summary the AI wrote as well as the links to dive deeper. I imagine that's what AI search will look like across all providers before long.
I threw in a question Wikipedia does kinda answer but leaves a lot of details off, and it enumerated a set of possible answers, with my favored option linking into a forum where it looks like less than a dozen people post, but with people that tried all variations of it and know all of the details.
DDG and Google would never show me that site (yeah, I've tried).
If you're good at googling the flow is: Ask the question > Clock which result isnt spam and click it > Figure out how to dismiss the cookies gate without accepting the cookies > Dismiss the google login box > Dismiss the popover pushing you to install an app > Scroll the page or ctrl+F to find the answer
With ChatGPT it's just type your question and your answer is appearing right away.
Of course I understand ChatGPT shouldn't be used for this because it will lie to you and make things up. But I am saying that's why you'll see people who don't know that preferring it over Google.
Build a product that's more convenient, and people will use it. Google is so full of shit now, and even answer boxes are below 4 ads, it's just way more comfortable and efficient to ask chatgpt.
This morning I wanted to know how many calories there are in a breaded chicken breast. Chatgpt told me in 3 seconds after asking. Google would have been way more time. (sidenote: I also hate google, so I'd much rather use chatgpt anyway)
And I’m not talking about fringe Q Anon type stuff: the other day I was looking up the specifics of China’s “Blue Sky Initiative” climate change policies and the only thing Goog/DDG would show me (despite several attempts at rephrasing the query) was Western industry think tanks bellyaching about how the policies effect profits. It took me a good ten minutes of refining my search before I got an English translation of the actual policy bullet points.
I can’t imagine this being such a PITA on 2013 era Google.
However, I’ve seen plenty of bad information in google answer boxes too. And finding it in actual search results is going to be way more time. It’s not a life and death question.
I'm learning typescript and ran into a weird typescript construct the other day, I threw it into chatgpt and asked "what is this" and it explained it to me. I'm not entirely sure if pasting in a bunch of braces and parentheses into google would find me the same result.
AI offers none of the above. Ask the same question twice in a row, maybe you'll get a different answer, maybe you won't. You won't know which of the two answers are hallucinations, which are true factoids it trained on, and which are bogus things it had been trained on. There's literally no providence for the results- no trail, no references, nothing, because it's really a sophisticated game of "Whose Line" where everything's made up and nothing matters.
I don't mean this as a way of thumbing my nose at people who don't know this stuff, but rather, to point out what a massive failure the completely nonexistent consumer education surrounding ML products has been.
The appeal of a conversational response was too great to wait for a solution to the problem of training data quality.
Search requires work on the part of the user to distinguish between good links and bad. AI is an oracle just tells you what you're looking for.
Now you and I might think this is a terrible way to evaluate the veracity of information. But think about the new generation of mobile-native users who were raised on simplistic discourse in tweet-length messages, and would rather watch a 1-min video on a topic than scan search results for 30 seconds.
For this group, searching for information where more than 2 clicks is required is going to be a "too complicated", anda "bad user experience".
LLMs are oracles that arrange words in a probabilistic order that are grammatically correct and may be factually correct. Unfortunately there's no way to evaluate the probability of confabulation with any of the LLM chat bots. The distribution of occurrences confabulation is also not regular or predictable nor is it fixed. So you can't ever say "ChatGPT is bad at X" because it can be bad today, good tomorrow, then bad the next day.
the opinions i run into real life can be very different then with people in the real world. the communities online are made up of the kinds of people who spend their time online, and the content you see on reddit is generally from the people who spend enough time on reddit that they want to browse new. these arent average people. only a minority professionals are actually engaged in reddit. even those ive seen run off because they don't agree with the acceptable opinions
Hegemony, perhaps.
I think you messed that sentence up, but I get what you're trying to say....
And I think it's this. The opinions you get from people 'IRL' are not apt to be as strong as the ones online, and or will run into the regency bias.
For example, it's very unlikely you'll actually meet someone that has used 10 different coffee makers because they wanted to see which one was best. Online on some subreddit, you're very likely to meet someone who has done exactly that. Of course those people with strong opinions are the ones that are apt to post most online.
So who's option is wrong? Neither. That's why they are opinions.
No, I've noticed what he's talking about, and it's not the strength of the opinion, it's what the opinion is. Reddit has a moral system that's completely misaligned with real world morality, where having a child or being autistic makes you a bad person, even if you didn't do anything wrong. You can find some really weird takes on r/AITA, which ironically points out that the subreddit has a fucked up sense of morality in its highest rated post.
Not sure if you've noticed but younger generations that go out and do things are big into having children. Now, some groups on reddit are more extreme on that, but the general trend of Americans at least is to tell the act of having kids to screw off.
---
But to answer your question, it's not Reddit that has this moral system, it's social media in general that has a moral system that does not match reality (kinda). The loudest idiots tend to get voted up, moderates disappear in the bulk of posts. Binary voting systems on sites tend to amplify this. Content suggestion systems tend to lift up contentious posts for engagement. Welcome to the internet.
But coming back to (kinda). This is becoming reality. Behavior IRL affects behavior online. Online behavior affects IRL behavior. People don't talk to their neighbors these days in most places. Communities are spread all over the earth.
not with a bang, but with a "this. take my upvote my good sir"
horrific
On the other hand, some random people's opinions in a Reddit community with -apparently- no further agenda seem somewhat more honest.
Not that the answer is better but it gives you new data points in your search.
Basically, it's not one or the other, you can use both tools and that's probably why it makes sense to include Reddit in AI models (which do this job for you automatically)
But the Reinforcement Learning from Human Feedback (RLHF) is also one of the key tools to getting useful outputs.
If you want truth, you don't want language, you want references to reviewed work. You also want things like 'show your work' chains of though. These are really different things.
If I tell GPT "make up a story" I don't want it coming back and saying, sorry I can only tell the truth.