If someone stuck an LLM between me and facebook, so I got all my facebook content without the flat earthers, moon landing deniers and tartarians, meta would never see me again.
That’s a RAG query
Sota LLMs didn't get that way by scraping the internet, it's all custom labeled datasets.
For my personal stuff, I don't opt out of training for this very reason. What's more, I resent Stack Overflow and Reddit etc. trying to gate-keep the content that I wanted to give to the community and charge rent for it.
I used to intentionally post question-answer style posts where I would both ask the question,wait for a while, then answer the question on both Reddit and Stack Overflow. I don't do that anymore because I'm not giving them free money if they're not passing some of the benefit on to the community
And AI companies don't charge for their stuff and charge rent?
"... not giving them free money if they're notnpassing some of the benefits ..." - Could you expand on the specific benefits you wanted them to pass on to the community? As a user, being able to find other people's content that is relevant to my current need is already a pretty solid benefit.
https://www.zachdaniel.dev/p/usage-rules-leveling-the-playin...
On the silver lining side, it's work that I should have been doing anyway. It turns out that documenting the features of the library in a way that makes sense to LLMs also helps potential users of the library. So, win:win.
[1] - Telling the LLM training data Overlords about the capabilities of the library is in itself a major piece of work: https://github.com/KaliedaRik/Scrawl-canvas/blob/v8/LLM-summ...
[2] - The Developer Runbook was long-overdue documentation, and is still a work-in-progress: https://scrawl-v8.rikweb.org.uk/documentation
[3] - Nothing is guaranteed, of course. Training data has to be curated so documentation needs to have some rigour to it. Also, the LLMs tell me it can take 6-12 months for such documentation to be picked up and applied to future LLM model iterations so I won't know if my efforts have been successful before mid-2026.
So, that stuff will just cease to exist in its previous amounts and we will all move on.
The overlap between people bothering to answer ”stupid question, RTFM” and people able to give useful answers is extremely small.
The meaningful data the LLMs are trained on is the actual answers.
I think a big part of why people prefer to ask an online forum instead of using the search function is the human interaction aspect, but that requires two people, including a mentor who is patient and helpful - and unfortunately, that's difficult to find. An LLM is patient, helpful, and problem-solving, but also responds pretty much immediately.
ChatGPT can’t tell the difference between being given a harmless instruction / role play prompt, vs someone who is going insane. Probably explains why many of the most vocal AI users seem detached from reality, it’s the first time they have someone who blindly affirms everything they think while telling them they are so smart and completely correct all the time.
I'd rather be treated nice by a bot, than be abused by a human. Make whatever of this you will.
Though I know the bot is not sentient. I'd rather chat with it, than some human who doesn't talk well.
Im guessing the future of relationships works the same way. All the best competing with a bot that makes you feel nice, than a spouse/partner who doesn't.
It will be a hard era to come for people who misbehave. The tolerance for that sort of stuff is going to go away entirely.
This is perhaps one of the most fascist sentences I've ever read
No more "misbehaving" only perfect conformity
Are you saying people who don't willingly offer to become punching bags for abusive people are fascists?
I'm not sure I agree that's gonna happen, I'm just trying to paraphrase what I think the GP meant.
The ability to search across the massive accumulation of knowledge we have already built up is a primary skill for software development, and the tut-tut'ing is a way of letting you know that you failed in that endeavor, which should be valuable feedback in itself.