I know, I know, the ad banner-funded web is a mess and I wouldn't mourn its demise either. But it worries me that it's an entirely open ended question for what actually replaces it.
I know, I know, the ad banner-funded web is a mess and I wouldn't mourn its demise either. But it worries me that it's an entirely open ended question for what actually replaces it.
It builds a paragraph answering your query but it has a lot of footnotes that link directly to websites.
Ex:
"geopolitical reason for palm oil being banned and why it's bad for health"
The EU has banned palm oil in biofuels due to its negative impacts on health[1] and its geopolitical implications, such as favoring alternative crops grown in Europe[2]. Palm plantations are also a major factor of deforestation[3], leading to the loss of habitat for endangered species[4]. Indonesia's President Joko Widodo recently announced a ban on the export of palm oil, which could backfire due to its importance in the global market[5].
1nih.gov 2weforum.org 3theconversation.com 4triplepundit.com 5carnegieendowment.org
Like for me, it gave be a footnote to this page: https://www.reddit.com/r/OpenAI/comments/109h24i/gpt_4_is_co...
When I asked about chatGPT.
I like to think of ChatGPT and the like as an on-demand personalized Wikipedia: a good starting point, comes with strings attached, not always correct (but some are fine with it).
ChatGPT doesn't link sources yet but I saw that the beta test context search from Kagi had them.
I feel like there is something here, but I wonder if second order effects make it trickier than it appears atm
So if you have a low effort website that's factual and text based, you're going to get your lunch eaten by GPT, if you have a higher effort website (subscription gated with lots of multimedia content and user engagement) you'll be fine.
Think of all the blogspam recipe sites that are going to run into trouble when ChatGPT learns to cook well. Lots of text, little additional value, no community. There still will be America's Test Kitchen because people on the upper end of the value curve don't just want a recipe, they want pictures + video of that recipe being made and a place where they can ask questions and get answers.
I guess to get the bots to tell your lies.
There was congressional testimony by the founder/owner of "Celebrity Net Worth" about how Google made it impossible for them to stay in business. Whenever somebody would search "How much is <celebrity X> worth?", the answer would just show up directly on the Google results page. There was still an attribution link to Celebrity Net Worth, but nobody ever clicked on it anymore, so the result was Celebrity Net Worth had to shut down.
You can certainly argue fairly whether sites like CNW deserve to exist in the first place, but it's not hard to see how there is still a huge financial problem when ALL the ad revenue goes to the search engines and they don't even leave any of the slim scraps to the publisher sites.
Same with stable diffusion, AI art. And it'll be the same with LLMs.
Eventually the whole internet would be flooded with cheap AI generated content and clearly AIs need human generated content to train on so it'll be the snake eats itself.
Geniune human-generated content will retreat to account-gated networks and private group chats where everyone knows one another. The rest of the internet will just be incestuous AI-generated chum.
In a way, it'd be an improvement. Genuine connection doesn't scale, so let's be honest about keeping it away from random parasites online.
This is fair—also in fairness, whenever my Google speaker answers a question, it always tells me which site the answer is from.
A chat bot, like search index, need to updated for new and current events to stay relevant. I can’t see why it can’t be deceived by spams
WebText and WebText2 referenced in their papers are corpuses based on Reddit submissions which had a 22% weight in their training model.
https://openwebtext2.readthedocs.io/en/latest/
This is larger than Wikipedia (3% weight) or either of their two book corpuses (8% each).
The only other data included was a filtered set from Common Crawl (weighted 60%).
---
I was imprecise with my language before but hopefully that at least provides some clarity.
For the time being but there is no reason AI couldn't produce genuine original content. Real life human artists also use previous content for inspiration.
The single most intellectually valuable website on the entire Web is very likely Library Genesis, where the only Web content is a catalog of books you can pirate by clicking a link, and it's, like, a lot more valuable than any other site (even Wikipedia). It may well be more valuable, in those terms, than the entire rest of the Web combined.
If serious book publishers survive a while longer and if non-fiction books aren't overrun with dubiously-accurate AI bullshit, things won't actually change all that much, I think.
As far as written content goes, the (public) Web is most useful for opinions or product discovery, and even those can already hardly be trusted because of all the marketing astroturfing. AI garbage barely changes that already-toxic dynamic.
Video's another matter—some video content on the Web is great and has ~no at-least-as-good replacement anywhere else, in any other medium. But it's also all but completely monopolized by Youtube and hardly participates in or factors into the broader Web.
I would enthusiastically welcome a web that isn't based on firehose-advertising and outright deception/lies.
I would too but I don't see how it happens. The awesome web that used to be was built on the backs of unpaid volunteers. Maybe that could return but even if it did all that wonderful volunteer work would get funneled through Google or OpenAI so investors can make a fat profit from it. Feels fundamentally wrong to me.
How that pans out in practice remains to be seen.
This is partially a BS answer. As long as websites are running Google Ads, they will have an incentive to be crawled. Fewer clicks > no clicks (which is what would happen if the site was set to 'noindex'.
Google also pays news publishers to license their content; $1bn alone just for Google News Showcase [0]
Are you saying Google pays them per scrape or they get paid only when users click through?
I think you are correct. These are two distinct products.