AI bots strain Wikimedia as bandwidth surges 50%
arstechnica.com
arstechnica.com
Now we see models showing the user it's "thinking" as some sort of intelligent agent, when really, ploys like that is to prop up the AI sector's stock price.
Surely LLMs in this form can't be the future of AGI?
Arguably this is using Wikipedia exactly for what it's designed for, although in an unexpectedly resource-intensive way. I bet just adding a web query cache for most frequently visited URLs on the side of the LLM provider could mitigate most of the negative effects here.
Like making charts out of CSV data, writing cute postcard messages to my girlfriend, or passive aggressive complaints to my landlord.
(87 points, 1 day ago, 97 comments) https://news.ycombinator.com/item?id=43555898
(47 points, 1 day ago, 45 comments) https://news.ycombinator.com/item?id=43562005
I get that it could be incompetence in one or two places, but as far as I know all the big providers fail here... so am I missing something obvious? Is it really just greed and vc money burning scrapefest?
>I get that it could be incompetence in one or two places, but as far as I know all the big providers fail here...
Is there any indication that the scraping traffic is from the "the big providers"? The article doesn't mention this.
My working theory is that rather than the big AI companies, this is mainly a rash of Asmall I startups (globally) who don't know any better and are just writing poor tools, and also think that "unique data" gives them an advantage. So this is more about a flood of capital to (frankly) incompetent groups, that will empty out when the boom busts. But I can see the arguments on both sides. It would be great to know the facts!
https://meta.wikimedia.org/wiki/Wikimedia_Foundation_Annual_...
> Key result FA2.1: Generate 5,000,000 views from short-form video content across all owned channels by the end of H1.
Wikipedia is actively saying it should pivot to short-form video to engage young people...
It's probably both cheaper and more accurate (for fast-changing content) to them to just hit Wikipedia servers every time than to special case search results pointing there and keeping a local copy.
Second sentence in the article.
How many large models are realistically being trained on all of Wikimedia, compared to the number of people using AI search, possibly fetching sites very inefficiently (e.g. using a headless browser spun up from scratch, i.e. with a cold cache, every single time).
Welcome to the new dark ages.