It's the same issue as with search engines providing "information boxes": huge win for me, but a mortal enemy for those who want to monetize anything resembling intellectual property.
> An LLM based internet seems to remove the creator and shim itself in between for the sake of business.
This sounds bad, and in some cases it is, but in others it is not. Content farms and recipe sites have creators behind them too.
The internet started thanks to ample government funding for research. So have many other technologies, including AI.
I wonder if there's a way we could all somehow pool our resources and use that to pay for common goods that we all use. What would we call such a scheme?
Is this tongue in cheek? I think it's called the government and taxes! :)
One of them is that creation costs of information are fixed, while its usefulness is unbounded, so it doesn't make sense to try and reward creators for each access/view/use, in perpetuity.
Secondly, there's a lot of information laundering going on - any random book I read carries between a few to few hundred references to prior written work. What I pay for the book goes to the author and the publishers, but AFAIK it doesn't go to any of the authors and publishers of works referenced in the book. Wikipedia takes this one step further, effectively turning all that information free.
Thirdly, AFAIK copyright explicitly does not cover information/knowledge - it covers specific works. So Google showing me an info box with a recipe scrapped from some site could technically fall afoul of the law - but an LLM generating me a recipe based on associations created from being trained on millions of recipes, this feels like it should be in the clear, at least from user's POV.
The new situation isn't the same as search as that wasn't there to hide information sources or to immediately convert information into useful things (texts, guides, etc.).
If this matters to you, then you shouldn't. But to flip this around: why should you care?
Unless you're doing some unique work targeting a global audience, the point when LLM gets trained on what you created is way outside space you'd normally care about. Trying to capture all the value your work generates does not lead to a good world.
Or maybe it's me who isn't profit-minded enough, but e.g. a lot of what I wrote on-line, including blog articles and commentary on Reddit and HN, has been used by search engines for free for a long time (over a decade, in some cases), and now is (most likely) part of the training corpora for LLMs. But I never believed, and still don't believe, that I'm entitled to some share of the gains LLMs (or search engines) make.
Valuable information in a way is becoming more valuable for the LLM provider, so I would expect a drop in high value information in the public domain.
In LLM land, they get no monetisation any more because nobody visits their sites, instead the LLM just regurgitates the answers they found.
The search engines actively supported these authors, by sending them people who needed the answers they had.
So in LLM land this information goes away because the feedback loop of the traveller creating information which earns them money to continue travelling goes away.
A LOT of the useful information on the web was built on similar feedback loops and they go away in LLM land.
I consider this to be a problem on its own, but it's not relevant here because:
> In LLM land, they get no monetisation any more because nobody visits their sites, instead the LLM just regurgitates the answers they found.
That can't possibly be true, because if it were, there wouldn't be any travel blogs anymore today. All that travel spam has been made redundant approximately around the time Flickr was created, and every interesting location ever has been photographed from every interesting angle in a way neither me, nor you, nor your favorite travel blogger could ever hope to match. All the information they post has also been posted many times over by travel bloggers that came before.
The point being: travel information and photography is worthless commodity these days. Travel bloggers (or Instagrammers, or whatever) are not in the business of selling information. They're selling dreams and personal experiences. The photos and information are necessary as delivery vector ("social object"), but by themselves are worthless and not the point. The point is entertainment, and ideally getting you trapped in a parasocial relationship with the travel blogger/grammer, which gives them a recurring revenue stream.
> In LLM land, they get no monetisation any more because nobody visits their sites, instead the LLM just regurgitates the answers they found.
It's the same model as with most other ad-monetized social media publishing. People will keep visiting them for the same reason they visit them now, and for the same reason they have their favorite youtubers and tiktokers. LLMs and other generative models don't change anything here, at least not short-to-mid-term, because they can't convincingly replicate human connection and keep it up for long.
(Also, I personally don't buy that travel instagrammers can actually sustain their travel lifestyle through ads and affiliate marketing and sponsorship deals. I suspect most are funded in some way, whether by family wealth or by services performed while traveling around.)
> The search engines actively supported these authors, by sending them people who needed the answers they had. (...) A LOT of the useful information on the web was built on similar feedback loops and they go away in LLM land.
Hard disagree. The only feedback loop this created in practice is the one that displaces quality information from the Internet - the combination of SEO and ad-based monetization means the most scummy players are the ones with most money to stalk every conceivable search query. The results speak for themselves: making a Google query for pretty much any topic of interest to general population will give you only content marketing sites - results that carry negative knowledge, as in if you waste your time reading them, you'll come more misinformed about the topic than you were before. If LLMs make all that go away, I'm 100% for it.
As for "A LOT of the useful information" - nope, can't think of a single case where ad/affiliate-supported site was a good information source, vs. just displacing a better free source.
https://stingynomads.com/annapurna-circuit-cost-planning/
Solid introduction and primer for the Annapurna Circuit, full of useful information from people who did it which has been kept up to date.
https://stingynomads.com/who-are-we/
> Today stingynomads.com is our full-time business and main source of income.
Now please tell me in what way is this not an example of an ad/affiliate supported site that provides a lot of useful information and what non ad/affiliate based resource has it displaced that was better? Cause I'm doubting someone would write up a better guide than that, publish it and not monetise it.
The way I see it, there are roughly three groups of information providers:
1. Those who do it pro bono - because they feel like its a worthwhile thing to do, or because they believe in by "pay it forward", or otherwise because they haven't even thought that what they share is worth trying to extract rent from.
2. Those who do it "for free", as a way to lure people to where they can expose them to ads, affiliate marketing, upsells, or other such schemes - making money by being predators using information as bait.
3. Those who just put up a paywall, being up front that they're selling information, not giving it away.
(There's also a weird "in-between" group of publishers that are almost like 1., except they're being funded out of marketing budgets of companies that figure providing quality information is good advertising.)
LLMs don't change anything for group #1. They may compete with group #3, but that's business as usual, not anything transformative. The group that's directly affected is #2, which also happens to be the group that produces all the garbage on the Internet, so I'm actually very happy to see them forced to find a more useful way of making money. Since group #2 produces "information" that's arguably negative knowledge on the net, it's likely to improve the amount of quality information you'll be able to find on-line.
Most people are paid for doing things every day, they don't get to create one thing and never work again. Expanding creators compensation laws is regressive and only helps a few elites survive job uncertainly, not the bulk of the people. We're better off limiting this sort of thing specifically to help everyone advance, share the knowledge.
How do you claim the right to learn from the works of others and then demand government regulation and forceful intervention to keep from having to share whatever paltry innovations you may develop?
(BTW. that you can even make a system this way is a huge breakthrough that's not being talked about enough.)
But even if they were a mere database indexing copies of other peoples' IP, then - copyright issues notwithstanding - the de-bullshittifying of information retrieval process alone would be service worth paying a lot of money for.
1) LLM providers harvest a common to create their product (don't think we disagree here much).
2) What happens next is where we diverge, I suspect: I think they will use their products to extract rents from that common while you think they will provide a fairly priced service.
Ultimately time will tell how the business model shakes out. Both could even be happening in sequence.
Phrased like this, I can't really disagree with you. I don't expect a business to play fair in general, when it has a profitable option to do otherwise.
I guess my objection is more that right now, I don't see LLMs creating any kind of disincentive to publish quality content. In my eyes, LLMs are not a substitute for quality content in the first place - I see them more like using quality content to create a tool that competes with ad-hoc and shitty content.
That's not to say LLMs won't be able to eventually provide high-quality information on their own - but at that point, we'll have more important problems to deal with, such as chunk of humanity being rendered obsolete.
I definitely believe that the I could get sued for my own ideas. The chance of lawyer claiming AI say it is their idea sounds horrible.
The word "creation" is loaded. No one "creates" content. They discover it hidden in some idea-space... occasionally even two people might discover the same thing. The same melody, the same verse of a poem, the same fragment of art.
The idea that one should be rewarded, but the other is slandered the infringer is amusingly dumb.
> but an LLM generating me a recipe based on associations created from being trained on millions of recipes, this feels like it should be in the clear, at least from user's POV
But how will we entrench the rent-seekers?
Originally I saw a website as a way to hang out a shingle. Until recently I was thinking I could just publish away and maybe someone would hire me based off the website.
Currently I don't feel the same way regarding publishing on the web. I will be me more guarded in what I share.
I'm surprised the nerve the issue has struck in me.
As an intuition pump, when I write "2+2 = " and you mentally complete it with "4", should I chastise you for not completing it with "4, as per ${your elementary class math textbook} and ${that other book you read as a kid}, corroborated by ${your first math teacher} and ${your parent} quoting ${some other work}"?
It's roughly the same thing.
I.e. in case of both the student and an LLM, correct citation doesn't actually mean the idea originates from the cited work - only that the work contains this idea.
>I do not want to be forced or prodded to establish relationships with creators or communities. I do not want their ads and upsells.
I feel the same way I lurked on HN for 7 years before signing up. I enjoy the lurking aspects of the internet.
I usually think the internet can survive anything.
This seems like the same disruption that Uber promoted to avoid regulation. Imagine all the great ideas we can generate and claim they are original to the unauditable AI thought process.
The last worst thing to happen to the internet IMHO was SEO. I sort think a schism may develop to the reality of the broadcast world and the narrowcast world.
I definitely think that valid uses exist.
And this is why humanity is going down the tubes...because you want something, you derive value from what you want, and yet you do not care about giving something back to who makes it.
On the contrary - this is exactly how and why humanity built a technological civilization in the first place. Note that I didn't say
> because you want something, you derive value from what you want, and yet you do not care about giving something back to who makes it.
Yes, because it would be backward and limiting to do that. Note: I never said I don't want to give anything back - I said I don't personally care specifically about the author/publisher. I don't want to establish any relationship with them. If I'm paying them directly for something, I'm paying them for that thing - not also for relationship (which really is a sales channel), not also for being advertised to. If I'm paying some intermediary, then rewarding the maker is the intermediary's problem, not mine.
Consider: do you compensate directly, and have an active relationship with, the person who bakes your bread (hint: people selling bread in bakeries are not actual bakers)? The company who supplied them with flour? The farmers who supplied the flour-makers with grain? Do you pay delivery drivers directly? After all, you're deriving value from their labor directly. Etc. Then there's an entire army of people whose work benefits you directly, and whom you don't even think much about, and rely on being compensated from the common pool (e.g. taxes) or stochastically.
The whole point of money is to allow exchanging value without forcing parties into maintaining an ongoing relationship. That's a feature, not a bug. And if anything is driving humanity down the drain, it's the idea that you should, need, or are even entitled to capture all the value you produce.
In case of LLMs, no compensation is directed to the person authoring the information. While it may not be a problem for the consumer of the information, it removes any incentives for the people authoring the information to continue doing so, which has long term consequences.
So when you say "contribute to the internet", this is what I consume and.. I'm sure there are similar examples in every niche- fishing, golf, coding, AI art creation ...
No I don't see this as gloomy scenario, and content creators- the goonhammers and honestwargamers, creatives, are still going to get paid (a bit, they were never rich), maybe in new ways.
OTOH, there are also plenty of technical blogs full of advanced content that is not "fun" to produce on its own, that are written to interact with a community of professionals (or juniors), and that might wither if engagement with actual human beings is reduced.
The people in the bread supply chain get paid, the author of content we're discussing will never get paid by anyone, will never even get a bit if personal satisfaction from their analytics knowing last month x thousand people read that page and it hopefully helped them.
It's completely zero reward, even worse it's completely zero feedback of any kind!
This really is the doom of the web as we know it because for the first time ever there will be an active disincentive to put knowledge on it. I think much information will retreat to places like Discord or locked down login only versions of sites like Stackoverflow.
It's addressing GP's complaint about me pointing out the indirect and transactional nature of the interaction between information producer and consumer.
> the author of content we're discussing will never get paid by anyone
That's... not my problem? Bear with me here.
> will never even get a bit if personal satisfaction from their analytics knowing last month x thousand people read that page and it hopefully helped them.
Aha!
So we're talking specifically about content creators that publish for free in hopes of maximizing a number on their analytics? That's healthy neither for them nor the society at large. Or you mean people publishing content for free to make money off ads? Yeah, I don't mind that content to disappear entirely.
Note that outside of web publishing, it was never the expectation of an author to have any idea how many people read their work, much less get paid for every single "read event". They only got a lump sum or a fraction of first sale of a printed work - but had no insight or control over further circulation of parts of entirety of their works. Being able to resell your books or magazines, or give or lend it to friends, or borrow some from a library, are all good things.
All that was true before AI, and people found reasons to write new books, or to publish quality content on-line, for free and without advertising or telemetry. LLMs don't change that. If anything, they may reduce readership, not publication.
> I think much information will retreat to places like Discord or locked down login only versions of sites like Stackoverflow.
This has already been happening for the past couple years; LLMs, again, don't change anything here.
I can't tell if I have aged out of some ideal or is it that the individual creative efforts are being homogenized into pseudo answers for someone to sell.
And I maintain that your attitude is one that makes the world worse. We should know where things come from and not have an anonymous transactional attitude towards it. Our technological civilization as you call it has led us to destruction with only a minority benefitting.
BTW, my favourite place to get bread is one in which they actually sell and bake it in store (a tiny market with its own oven).
Maybe my simple scheme would be all it needs. Or maybe it needs some new breakthrough and right now nobody knows how to do it. I was hoping some resident expert could let me know.
First, it seems possible that if sources were in the training data like I described, then understanding of sources could be an emergent capability, just because the LLM reads "the source of the following is X."
Second, maybe a trainer LLM could be tasked with reading the trainee's answers and any sources it provides, and judging whether the source is correct.
But I'm no expert, hence my question.
In operation, the "trainer" could do the same thing in the background. And then of course, human users could also check up on the sources if they need to be sure of catching hallucinations.
At the same time, could it not be construed as humanity leveling up? We've abstracted away that, now low, level of thought. Much like we don't have people pressing the button in the elevator. Most developers and writers are still better than an LLM, but its good enough to replace their input on simple tasks.
Writing is being commoditized much like other industries. The moment you release a new vacuum cleaner there's ten others that do the same thing and nobody is fussed over who invented it. People still know who to go for a premium vacuum though.
Video Audio Images Heat Motion/acceleration Lidar RF Sonar Radar Network traffic Atmospheric pressure Wind vectors Magnetic fields System/application logs Electrical current UV X-ray Microwave Ionizing particle emissions
Anyway, I really like that name. Non-dataist AIs are clearly the best kind.
It's not clear what the effect will be of LLMs consuming more and more of their own output and/or waste products.
One possibility might be a feedback system that leads to superintelligence which eventually becomes incomprehensible to humans.
Another might be a feedback system that leads to increased garbage output that eventually devolves into incomprehensible noise and nonsense.
The two extremes might not be easily distinguishable.