I think a big part of why people prefer to ask an online forum instead of using the search function is the human interaction aspect, but that requires two people, including a mentor who is patient and helpful - and unfortunately, that's difficult to find. An LLM is patient, helpful, and problem-solving, but also responds pretty much immediately.
Sota LLMs didn't get that way by scraping the internet, it's all custom labeled datasets.
For my personal stuff, I don't opt out of training for this very reason. What's more, I resent Stack Overflow and Reddit etc. trying to gate-keep the content that I wanted to give to the community and charge rent for it.
I used to intentionally post question-answer style posts where I would both ask the question,wait for a while, then answer the question on both Reddit and Stack Overflow. I don't do that anymore because I'm not giving them free money if they're not passing some of the benefit on to the community
And AI companies don't charge for their stuff and charge rent?
"... not giving them free money if they're notnpassing some of the benefits ..." - Could you expand on the specific benefits you wanted them to pass on to the community? As a user, being able to find other people's content that is relevant to my current need is already a pretty solid benefit.
https://www.zachdaniel.dev/p/usage-rules-leveling-the-playin...
On the silver lining side, it's work that I should have been doing anyway. It turns out that documenting the features of the library in a way that makes sense to LLMs also helps potential users of the library. So, win:win.
[1] - Telling the LLM training data Overlords about the capabilities of the library is in itself a major piece of work: https://github.com/KaliedaRik/Scrawl-canvas/blob/v8/LLM-summ...
[2] - The Developer Runbook was long-overdue documentation, and is still a work-in-progress: https://scrawl-v8.rikweb.org.uk/documentation
[3] - Nothing is guaranteed, of course. Training data has to be curated so documentation needs to have some rigour to it. Also, the LLMs tell me it can take 6-12 months for such documentation to be picked up and applied to future LLM model iterations so I won't know if my efforts have been successful before mid-2026.
So, that stuff will just cease to exist in its previous amounts and we will all move on.
If someone stuck an LLM between me and facebook, so I got all my facebook content without the flat earthers, moon landing deniers and tartarians, meta would never see me again.
That’s a RAG query
The overlap between people bothering to answer ”stupid question, RTFM” and people able to give useful answers is extremely small.
The meaningful data the LLMs are trained on is the actual answers.