Are large language models a threat to digital public goods?
arxiv.org
arxiv.org
"Second, we investigate whether ChatGPT is simply displacing simpler or lower quality posts on Stack Overflow. To do so, we use data on up- and downvotes, simple forms of social feedback provided by other users to rate posts. We observe no change in the votes posts receive on Stack Overflow since the release of ChatGPT. This finding suggests that ChatGPT is displacing a wide variety of Stack Overflow posts, including high-quality content."
Can anyone tell me how they are related and what it means?
PS: I also asked ChatGPT (4). Here's what it says https://chat.openai.com/share/7029a1e5-63d0-4cec-bf76-10b0b5...
The moderation team and community in general on Stack Overflow is so toxic I'm not even sure you could control for that effect well enough to arrive at this conclusion. I would argue people are leaving because it's easier to ask ChatGPT your question than be flamed and banned for asking how to do something. A half right answer from ChatGPT is better than getting marked duplicate and closed because the moron moderation team can't detect nuance.
> the moron moderation team can't detect nuance.
It is unwise to leave the detection of nuance to volunteer moderators. Rather, highlight the nuance in your question (or answer). That strategy had always landed me high value answers to my questions.But shallow questions where the official docs are too raw or are missing a few specifics? Stack Overflow was good for that before chatgpt.
So in 2023 I almost never visit SO, not because of GPT but because the people who write frameworks also supply you with documentation and implementation examples.
Again anecdotal, but I do use GPT quite a bit. Though rarely to help me figure something out. It writes my documentation, unless the code is very secret. It sometimes writes some basic code like auto-generating a docker file or basic API classes. Things that aren’t in competition with SO. I’m sure it’ll take over the role of SO for some parts, but to me personally, it seems far more likely that SO’s decline is linked to a range of other challenges as well.
they should ask people that ask questions, not people passively looking for answers
the experience asking a question on stackexchange sites is horrible! each tag is its own community with edicts and customs you have no idea about which derail your path to actually getting an answer, the auto-moderation system is completely broken because it thinks your question should be replaced with one from 2013, or the human moderators unilaterally fix your question which then causes the auto-moderator to think its similar to the one from 2013 and unceremoniously replaces it
With chatgpt you dont need to prove anything about what you tried, you dont need a reputation score to do anything, you dont need to be a steward of every upvote for all eternity, you dont need to engage in meta discussion to get your post out of deletion, and it answers faster. You don't need to be told to ask a separate question if you have any followups, with all of the same gamble of problems, while chatgpt already assumes what you will have a followup question about and just tells you all the gotchas. Because its read all the forum posts and has seen what the recurring issues are.
its better, faster, and cheaper time wise
its not a direct comparison, chatgpt is more like a pair programmer. If stackoverflow had an adhoc pair programming live session you could hop into it would be a better comparison.
I can't think of a way, buy it would be interesting to see if there is a similar effect on Wikipedia contributions.
My alternative hypothesis: ChatGPT is eliminating all the long-tail, garbage questions, so the vote patterns don't change because the missing questions weren't getting much of them anyway.
> Could it not be the case that SO is less adept at finding related/duplicate questions than ChatGPT? Given the later's facility with the language, I would expect it to be.
Given the later's facility with the language, I would expect it to be a better search engine. I would expect that ChatGPT is replacing Google as sort of tokenizer. I.e. search pipeline changes from "form natural question -> input to google -> go to SO -> if first few links do not yield answer post new question" to "form natural question -> input to chatgpt -> extract keyword tokens -> input to google -> go to SO".
There is an important bit in the article:
> > Using data on programming language popularity on GitHub, we find that the most widely used languages tend to have larger relative declines in posting activity.
Here we can form a hypothesis that reduction in post frequency comes from entry level posts with posters not knowing what to search for. Under this hypothesis ChatGPT has strongest effect on users using SO as knowledge base rather than Q&A forum. This user type distinction would affect posting frequency much more heavily than voting patterns.
Still, I find ChatGPT answers much harder to validate. In programming Q&A, the answer is usually a series of calls. When you try to apply the human solution (assuming it's at least somewhat correct), the error often lies in input data mismatch or some minor changes in requirements when the answer is not perfectly aligned. Rarely the calls themselves are wrong - they might be deprecated if it's an old answer, but generally, they do at least exist. With ChatGPT, you never know which particular part of the answer has been hallucinated.
Also, as SO has more stringing question requirements, you are forced to construct a minimal case to reproduce the error. Often, composing an SO question got me straight to the cause of the error.
It seems like these “deaths” are either temporary or something else will pop-up that continuously improves and trains an LLMs with proprietary data.
There’s def a shift in the social contract of the internet. We’re shifting from “publish a few things you know a lot bit in exchange to read stuff other smart people have shared” to “solve a problem with an LLM and train it in a manner that can be helpful to others”
What’s unclear is how the middle of that will be monetized. Search engines figured out how to do it for the “publish and share” paradigm and made a boatload of money off of it after paying for the massive infrastructure required to index everything. How will LLMs do it without killing the data that trains it?
I feel like this hasn't been the majority of internet users experience for ~20 years
It was, nevertheless, internet.
I kind of expect those projects would want ChatGPT to write the documentation. ;)
If you struggle with broad issues a HN post is probably a good place to start, assuming you can express your question well. I imagine you'll get the highest signal to noise ratio here.
Edit: btw nobody can pm you as you have no contact info in your bio afaict.
Ironically, internet discussion forums democratized specialist knowledge that you previously would have had to be in the right university or research circle to access. SO monetized that, and now we've moved on to complaints like this about further democratization of knowledge and (to quote another post) that knowledge being "stolen".
However, there are also links between jobs that no longer care, and GenZ is a smaller cohort, and also maybe more of an interest in trades as a path forward.
I'm not suggesting the paper linked here is right or wrong, but we should keep an open mind about all the other potential reasons why things suddenly shift. One random reason could be that during this overly oppressive year of layoffs engineers have either stopped needing to search or haven't had time to work on a project that needed it.
Also there's smarter IDEs which are likely the #1 reason I've stopped searching the internet for one-off small things like "what's the API to deep copy an array in Rust." Documented autocomplete has gotten really good and things like copilot are just icing on the cake.
https://web.archive.org/web/20230315152647/https://www.forbe...
This is bad at first glance because look what happened when smart phones became popular: no one remembers phone numbers of their family, people can’t remember directions when driving. Now you don’t even need the ability to generate ideas.
And it’s not clear to me that LLMs democratises information. LLM have guard rails, are potentially trained on copyrighted data. Need to be run by large companies…
This is a concern, though I'd argue unrelated to the concern about people no longer contributing publicly online, and one that was already present with SO. I've seen SO posts specifically reference or criticize the "copy paste" crowd that is just taking the answer and putting it into their code without thought. Some people will use any technology blindly.
Those coming to adulthood in the 90s were the first generation to really get to experience mass, endless, and cheap dopamine driven digital entertainment. Notably the [positive] Flynn effect is still in full swing in places like China and India, where digital entertainment market saturation has taken longer. So they'll effectively work as perfect experiments. If it reverses over the coming decades, as everybody over in these places now has their heads shoved in e.g. smart phones as much as anywhere else, it should be telling.
[1] - https://www.sciencealert.com/iq-scores-falling-in-worrying-r...
Implementation is when skill matters.
I can imagine post scarcity Star Trek like ideas but not implement them. My skills at CE and SWE allow me to implement reasonable solutions there.
Chasing ideas for the sake of chasing ideas is a form of bike shedding and premature optimization.
Something new is going on.
I use ChatGPT for work in a limited way. If you haven’t tried it, I suggest that you do. This potentially paradigm shifting technology.
Even then you'd want attribution to survive.
Eliminating or changing the incentives people have for doing things is the absolutely surest way to change behavior.
Postulating that all published work can be appropriated at will by certain private for-profit entities is the death knell of the knowledge economy as we know it.
I sincerely hope you are shitposting. So called Web 2.0, the platformized web has placed huge incentives to stifle propagation and dissemination of ideas. However, these incentives are somewhat distributed under human review, therefore there can be bubbles with opposing ideas. LLMs centralize censorship, by design.
Currently you can find "BrawndoLounge" and "ToiletWater" subreddits that heavily censor opposing ideas. Even if one is full of hate speech or whatever, there still are opposing ideas there. If LLM guardians decide that one of these sources is "toxic", in an LLM-led web that viewpoint simply vanishes.
Prove this.
I would not be surprised to find StackOverflow usage dropped significantly because of ChatGPT. It's simply a much more effective tool for getting help with typical programming problems. Not as good of a resource for expert-level or architectural advice, but that's okay with me. Basic "debugging via internet" is much easier to do with an interactive service with lots of knowledge.
It's often pretty helpful just pasting an entire error message with backtrace into ChatGPT and seeing what it thinks.
You can make fun of me all you like, but it's taken me decades to get good at this, and I'll be damned if some soft-skinned SV kid with a MacBook uses my work to power his mill.
When I felt like writing again after years of closing my previous blog, it though about it for a while before committing to https://bitecode.dev.
But eventually, I realized that
- I also write for myself, not just for others.
- People like reading things without having to prompt for it. So they will read the blog because it's nice and topics come to them even if they don't think that they need to know.
- ChatGPT doesn't have opinion. It tries very hard to be balanced. Your blog will have an opinion.
- There is more to the experience you provide than just knowledge. You can add tools, exercises, videos, etc. Which GPT cannot, for now, replicate.
- GPT can replicate style, but will not by default. People will come for your style as well. And pics. And design. And jokes.
- People value the interaction they feel when they content seems like there is a person behind it. They get attached. They develop sympathy.
- A blog puts things in context. "If you want to know that, you probably need to know that". It also gives information about what happen right now.
- Humans curate. In a world where creating crappy content is very cheap, a good filter has tremendous value.
So yes, you will be scanned, and replicated. Doesn't mean you don't have value in writing what you do.
And making something of value is nice.
This may change in 10 years or so. Maybe the LLM will be able to do all that. But depriving yourself of a rewarding experience right now for the fear of what might happen is not worth it.
Those that do make fun of you are victim blaming. A lot of these folks stating that if you put out in the public then it's not yours anymore sound like criminals to be fair.
It's the same issue as with search engines providing "information boxes": huge win for me, but a mortal enemy for those who want to monetize anything resembling intellectual property.
> An LLM based internet seems to remove the creator and shim itself in between for the sake of business.
This sounds bad, and in some cases it is, but in others it is not. Content farms and recipe sites have creators behind them too.
And this is why humanity is going down the tubes...because you want something, you derive value from what you want, and yet you do not care about giving something back to who makes it.
On the contrary - this is exactly how and why humanity built a technological civilization in the first place. Note that I didn't say
> because you want something, you derive value from what you want, and yet you do not care about giving something back to who makes it.
Yes, because it would be backward and limiting to do that. Note: I never said I don't want to give anything back - I said I don't personally care specifically about the author/publisher. I don't want to establish any relationship with them. If I'm paying them directly for something, I'm paying them for that thing - not also for relationship (which really is a sales channel), not also for being advertised to. If I'm paying some intermediary, then rewarding the maker is the intermediary's problem, not mine.
Consider: do you compensate directly, and have an active relationship with, the person who bakes your bread (hint: people selling bread in bakeries are not actual bakers)? The company who supplied them with flour? The farmers who supplied the flour-makers with grain? Do you pay delivery drivers directly? After all, you're deriving value from their labor directly. Etc. Then there's an entire army of people whose work benefits you directly, and whom you don't even think much about, and rely on being compensated from the common pool (e.g. taxes) or stochastically.
The whole point of money is to allow exchanging value without forcing parties into maintaining an ongoing relationship. That's a feature, not a bug. And if anything is driving humanity down the drain, it's the idea that you should, need, or are even entitled to capture all the value you produce.
The people in the bread supply chain get paid, the author of content we're discussing will never get paid by anyone, will never even get a bit if personal satisfaction from their analytics knowing last month x thousand people read that page and it hopefully helped them.
It's completely zero reward, even worse it's completely zero feedback of any kind!
This really is the doom of the web as we know it because for the first time ever there will be an active disincentive to put knowledge on it. I think much information will retreat to places like Discord or locked down login only versions of sites like Stackoverflow.
It's addressing GP's complaint about me pointing out the indirect and transactional nature of the interaction between information producer and consumer.
> the author of content we're discussing will never get paid by anyone
That's... not my problem? Bear with me here.
> will never even get a bit if personal satisfaction from their analytics knowing last month x thousand people read that page and it hopefully helped them.
Aha!
So we're talking specifically about content creators that publish for free in hopes of maximizing a number on their analytics? That's healthy neither for them nor the society at large. Or you mean people publishing content for free to make money off ads? Yeah, I don't mind that content to disappear entirely.
Note that outside of web publishing, it was never the expectation of an author to have any idea how many people read their work, much less get paid for every single "read event". They only got a lump sum or a fraction of first sale of a printed work - but had no insight or control over further circulation of parts of entirety of their works. Being able to resell your books or magazines, or give or lend it to friends, or borrow some from a library, are all good things.
All that was true before AI, and people found reasons to write new books, or to publish quality content on-line, for free and without advertising or telemetry. LLMs don't change that. If anything, they may reduce readership, not publication.
> I think much information will retreat to places like Discord or locked down login only versions of sites like Stackoverflow.
This has already been happening for the past couple years; LLMs, again, don't change anything here.
I can't tell if I have aged out of some ideal or is it that the individual creative efforts are being homogenized into pseudo answers for someone to sell.
In case of LLMs, no compensation is directed to the person authoring the information. While it may not be a problem for the consumer of the information, it removes any incentives for the people authoring the information to continue doing so, which has long term consequences.
So when you say "contribute to the internet", this is what I consume and.. I'm sure there are similar examples in every niche- fishing, golf, coding, AI art creation ...
No I don't see this as gloomy scenario, and content creators- the goonhammers and honestwargamers, creatives, are still going to get paid (a bit, they were never rich), maybe in new ways.
OTOH, there are also plenty of technical blogs full of advanced content that is not "fun" to produce on its own, that are written to interact with a community of professionals (or juniors), and that might wither if engagement with actual human beings is reduced.
And I maintain that your attitude is one that makes the world worse. We should know where things come from and not have an anonymous transactional attitude towards it. Our technological civilization as you call it has led us to destruction with only a minority benefitting.
BTW, my favourite place to get bread is one in which they actually sell and bake it in store (a tiny market with its own oven).
>I do not want to be forced or prodded to establish relationships with creators or communities. I do not want their ads and upsells.
I feel the same way I lurked on HN for 7 years before signing up. I enjoy the lurking aspects of the internet.
I usually think the internet can survive anything.
This seems like the same disruption that Uber promoted to avoid regulation. Imagine all the great ideas we can generate and claim they are original to the unauditable AI thought process.
The last worst thing to happen to the internet IMHO was SEO. I sort think a schism may develop to the reality of the broadcast world and the narrowcast world.
I definitely think that valid uses exist.
One of them is that creation costs of information are fixed, while its usefulness is unbounded, so it doesn't make sense to try and reward creators for each access/view/use, in perpetuity.
Secondly, there's a lot of information laundering going on - any random book I read carries between a few to few hundred references to prior written work. What I pay for the book goes to the author and the publishers, but AFAIK it doesn't go to any of the authors and publishers of works referenced in the book. Wikipedia takes this one step further, effectively turning all that information free.
Thirdly, AFAIK copyright explicitly does not cover information/knowledge - it covers specific works. So Google showing me an info box with a recipe scrapped from some site could technically fall afoul of the law - but an LLM generating me a recipe based on associations created from being trained on millions of recipes, this feels like it should be in the clear, at least from user's POV.
The new situation isn't the same as search as that wasn't there to hide information sources or to immediately convert information into useful things (texts, guides, etc.).
If this matters to you, then you shouldn't. But to flip this around: why should you care?
Unless you're doing some unique work targeting a global audience, the point when LLM gets trained on what you created is way outside space you'd normally care about. Trying to capture all the value your work generates does not lead to a good world.
Or maybe it's me who isn't profit-minded enough, but e.g. a lot of what I wrote on-line, including blog articles and commentary on Reddit and HN, has been used by search engines for free for a long time (over a decade, in some cases), and now is (most likely) part of the training corpora for LLMs. But I never believed, and still don't believe, that I'm entitled to some share of the gains LLMs (or search engines) make.
Valuable information in a way is becoming more valuable for the LLM provider, so I would expect a drop in high value information in the public domain.
In LLM land, they get no monetisation any more because nobody visits their sites, instead the LLM just regurgitates the answers they found.
The search engines actively supported these authors, by sending them people who needed the answers they had.
So in LLM land this information goes away because the feedback loop of the traveller creating information which earns them money to continue travelling goes away.
A LOT of the useful information on the web was built on similar feedback loops and they go away in LLM land.
I consider this to be a problem on its own, but it's not relevant here because:
> In LLM land, they get no monetisation any more because nobody visits their sites, instead the LLM just regurgitates the answers they found.
That can't possibly be true, because if it were, there wouldn't be any travel blogs anymore today. All that travel spam has been made redundant approximately around the time Flickr was created, and every interesting location ever has been photographed from every interesting angle in a way neither me, nor you, nor your favorite travel blogger could ever hope to match. All the information they post has also been posted many times over by travel bloggers that came before.
The point being: travel information and photography is worthless commodity these days. Travel bloggers (or Instagrammers, or whatever) are not in the business of selling information. They're selling dreams and personal experiences. The photos and information are necessary as delivery vector ("social object"), but by themselves are worthless and not the point. The point is entertainment, and ideally getting you trapped in a parasocial relationship with the travel blogger/grammer, which gives them a recurring revenue stream.
> In LLM land, they get no monetisation any more because nobody visits their sites, instead the LLM just regurgitates the answers they found.
It's the same model as with most other ad-monetized social media publishing. People will keep visiting them for the same reason they visit them now, and for the same reason they have their favorite youtubers and tiktokers. LLMs and other generative models don't change anything here, at least not short-to-mid-term, because they can't convincingly replicate human connection and keep it up for long.
(Also, I personally don't buy that travel instagrammers can actually sustain their travel lifestyle through ads and affiliate marketing and sponsorship deals. I suspect most are funded in some way, whether by family wealth or by services performed while traveling around.)
> The search engines actively supported these authors, by sending them people who needed the answers they had. (...) A LOT of the useful information on the web was built on similar feedback loops and they go away in LLM land.
Hard disagree. The only feedback loop this created in practice is the one that displaces quality information from the Internet - the combination of SEO and ad-based monetization means the most scummy players are the ones with most money to stalk every conceivable search query. The results speak for themselves: making a Google query for pretty much any topic of interest to general population will give you only content marketing sites - results that carry negative knowledge, as in if you waste your time reading them, you'll come more misinformed about the topic than you were before. If LLMs make all that go away, I'm 100% for it.
As for "A LOT of the useful information" - nope, can't think of a single case where ad/affiliate-supported site was a good information source, vs. just displacing a better free source.
https://stingynomads.com/annapurna-circuit-cost-planning/
Solid introduction and primer for the Annapurna Circuit, full of useful information from people who did it which has been kept up to date.
https://stingynomads.com/who-are-we/
> Today stingynomads.com is our full-time business and main source of income.
Now please tell me in what way is this not an example of an ad/affiliate supported site that provides a lot of useful information and what non ad/affiliate based resource has it displaced that was better? Cause I'm doubting someone would write up a better guide than that, publish it and not monetise it.
The way I see it, there are roughly three groups of information providers:
1. Those who do it pro bono - because they feel like its a worthwhile thing to do, or because they believe in by "pay it forward", or otherwise because they haven't even thought that what they share is worth trying to extract rent from.
2. Those who do it "for free", as a way to lure people to where they can expose them to ads, affiliate marketing, upsells, or other such schemes - making money by being predators using information as bait.
3. Those who just put up a paywall, being up front that they're selling information, not giving it away.
(There's also a weird "in-between" group of publishers that are almost like 1., except they're being funded out of marketing budgets of companies that figure providing quality information is good advertising.)
LLMs don't change anything for group #1. They may compete with group #3, but that's business as usual, not anything transformative. The group that's directly affected is #2, which also happens to be the group that produces all the garbage on the Internet, so I'm actually very happy to see them forced to find a more useful way of making money. Since group #2 produces "information" that's arguably negative knowledge on the net, it's likely to improve the amount of quality information you'll be able to find on-line.
Most people are paid for doing things every day, they don't get to create one thing and never work again. Expanding creators compensation laws is regressive and only helps a few elites survive job uncertainly, not the bulk of the people. We're better off limiting this sort of thing specifically to help everyone advance, share the knowledge.
(BTW. that you can even make a system this way is a huge breakthrough that's not being talked about enough.)
But even if they were a mere database indexing copies of other peoples' IP, then - copyright issues notwithstanding - the de-bullshittifying of information retrieval process alone would be service worth paying a lot of money for.
1) LLM providers harvest a common to create their product (don't think we disagree here much).
2) What happens next is where we diverge, I suspect: I think they will use their products to extract rents from that common while you think they will provide a fairly priced service.
Ultimately time will tell how the business model shakes out. Both could even be happening in sequence.
Phrased like this, I can't really disagree with you. I don't expect a business to play fair in general, when it has a profitable option to do otherwise.
I guess my objection is more that right now, I don't see LLMs creating any kind of disincentive to publish quality content. In my eyes, LLMs are not a substitute for quality content in the first place - I see them more like using quality content to create a tool that competes with ad-hoc and shitty content.
That's not to say LLMs won't be able to eventually provide high-quality information on their own - but at that point, we'll have more important problems to deal with, such as chunk of humanity being rendered obsolete.
How do you claim the right to learn from the works of others and then demand government regulation and forceful intervention to keep from having to share whatever paltry innovations you may develop?
The word "creation" is loaded. No one "creates" content. They discover it hidden in some idea-space... occasionally even two people might discover the same thing. The same melody, the same verse of a poem, the same fragment of art.
The idea that one should be rewarded, but the other is slandered the infringer is amusingly dumb.
> but an LLM generating me a recipe based on associations created from being trained on millions of recipes, this feels like it should be in the clear, at least from user's POV
But how will we entrench the rent-seekers?
I definitely believe that the I could get sued for my own ideas. The chance of lawyer claiming AI say it is their idea sounds horrible.
Originally I saw a website as a way to hang out a shingle. Until recently I was thinking I could just publish away and maybe someone would hire me based off the website.
Currently I don't feel the same way regarding publishing on the web. I will be me more guarded in what I share.
The internet started thanks to ample government funding for research. So have many other technologies, including AI.
I wonder if there's a way we could all somehow pool our resources and use that to pay for common goods that we all use. What would we call such a scheme?
Is this tongue in cheek? I think it's called the government and taxes! :)
I'm surprised the nerve the issue has struck in me.
As an intuition pump, when I write "2+2 = " and you mentally complete it with "4", should I chastise you for not completing it with "4, as per ${your elementary class math textbook} and ${that other book you read as a kid}, corroborated by ${your first math teacher} and ${your parent} quoting ${some other work}"?
It's roughly the same thing.
I.e. in case of both the student and an LLM, correct citation doesn't actually mean the idea originates from the cited work - only that the work contains this idea.
At the same time, could it not be construed as humanity leveling up? We've abstracted away that, now low, level of thought. Much like we don't have people pressing the button in the elevator. Most developers and writers are still better than an LLM, but its good enough to replace their input on simple tasks.
Writing is being commoditized much like other industries. The moment you release a new vacuum cleaner there's ten others that do the same thing and nobody is fussed over who invented it. People still know who to go for a premium vacuum though.
Maybe my simple scheme would be all it needs. Or maybe it needs some new breakthrough and right now nobody knows how to do it. I was hoping some resident expert could let me know.
First, it seems possible that if sources were in the training data like I described, then understanding of sources could be an emergent capability, just because the LLM reads "the source of the following is X."
Second, maybe a trainer LLM could be tasked with reading the trainee's answers and any sources it provides, and judging whether the source is correct.
But I'm no expert, hence my question.
In operation, the "trainer" could do the same thing in the background. And then of course, human users could also check up on the sources if they need to be sure of catching hallucinations.
Video Audio Images Heat Motion/acceleration Lidar RF Sonar Radar Network traffic Atmospheric pressure Wind vectors Magnetic fields System/application logs Electrical current UV X-ray Microwave Ionizing particle emissions
Anyway, I really like that name. Non-dataist AIs are clearly the best kind.
It's not clear what the effect will be of LLMs consuming more and more of their own output and/or waste products.
One possibility might be a feedback system that leads to superintelligence which eventually becomes incomprehensible to humans.
Another might be a feedback system that leads to increased garbage output that eventually devolves into incomprehensible noise and nonsense.
The two extremes might not be easily distinguishable.
Running a LLM does not seem to create much of a moat. Everybody is doing it now.
>> But since users interact privately with the model, these models may drastically reduce the amount of publicly available human-generated data and knowledge resources.
The premise seems a stretch when the models are simply providing a better way to find the information. It is essentially reducing the number of searches or low quality questions floating around.
For creator, it is just an aid to improve the creation. The models help them find duplicates better or simulate variations faster.
This paper appears to construct a case of walled garden by mis-representing user questions as the actual content/creation.
User search questions always remained private with the site owners (Google search, Bing etc.,)
(I'll note that proprietary models would be first-order disincentivised to stop this because such conversations can be used to better train other models, but of course they will benefit more generally for maintaining a commons of knowledge to draw from.)
Is it really removed if it is easier to get, and at a higher quality than a lot of SEO spam search results? On a primitive level I'm not sure it's much different than when computers & calculators replaced printed mathematical tables, such ones used that were used in artillery firing.
That said, there are other consideration of accessibility separate from the broader topic LLM's in general. Specifically, that LLM's truly suited as better alternatives to traditional content might end up gate-keeped by corporations charging high prices, an especially likely scenario if lawsuits regarding copyright rule in favor strict copyright protection of the work as not eligible for a fair use exemption. And/or if regulatory capture puts the barrier of entry into LLM's creation too high. Contrary as it might seem, losing the copyright battle might be a net benefit to Open AI and large competitors: It will be very difficult for more open alternatives to get the resources required to build open models.
I don't think use of LLM's, in themselves, constitutes a risk to public knowledge, in so far as they remain just as or more accessible than traditional content. That is the case right now, for some use cases, where I can get a much faster answer and immediately critique or get follow up responses to clarify things.
Right now, most of those services very strictly refuse to say anything even mildly controversial or pornographic. Because LLMs have an increasingly large amount of influence on society, this is quite concerning, because many things will simply get memoryholed if the creators of the models deem it's not "safe" or acceptable.