Perplexity's grand theft AI
theverge.com
theverge.com
There's a distinction that is being missed between:
* Web crawlers automatically accessing pages, such as recursively following links to index them for search engines
* A tool accessing a URL in direct response to a user request
robots.txt is only intended for the former. For instance, archive.is state:
> [Why does archive.is not obey robots.txt?] Because it is not a free-walking crawler, it saves only one page acting as a direct agent of the human user. Such services don't obey robots.txt (e.g. Google Feedfetcher, screenshot- or pdf-making services, isup.me, …)
Perplexity do (as far as I've been able to find) respect robots.txt for their scraping. What the investigations confirmed it was ignored for was users entering a URL to summarise.
There is still a conversation to be had. I think users should be able to browse the web with whatever tool they want to - whether that's a standard browser, a minimal reader mode that doesn't show ads, or a statistical summarisation tool - but it does pose a problem for sites that rely on standard browsers letting them show users advertisements, get them to sign up to mailing lists, collect data, etc. if users decide to choose tools that don't allow that.
AI that answers questions using online sources isn't a web crawler, but a web researchers.
Once we get to the point of this AI running locally, or you just have a browser extension that grabs the data and sends it to the server, what is the difference between what it is doing now and that.
If I use an AI and I give it a specific URL I think the expectation is that it is acting like a user and not like a crawler so the robots.txt is meaningless.
That doesn't mean that perplexity is free of any blame here, if they are getting around paywalls or then using this loophole to gather up the data to train later. That isnt ok. I am also concerned about the reports of them using third party crawlers to get around blocking them.
But, purely on what this seems to be focused on and the example given. I don't see anything wrong with it when it is triggered based on a specific action of the user. And should be the expected behavior IMO.
Headline grabbing articles like this I feel like undermine the valid criticism of AI because it can be used to paint a picture of a vendetta against AI or worse as an example of why these arguments are wrong as a way to discredit large swaths of criticism.
We really need to be more mindefull about how we criticize the real problems with the technology.
This type of summarization and aggregation seems to be exactly what consumers want.
Perplexity AI is lying about their user agent https://news.ycombinator.com/item?id=40690898
Perplexity AI is lying about their user agent - https://news.ycombinator.com/item?id=40690898 - June 2024 (531 comments)
Ron Howard voiceover: They do not.
On his recent appearance on the Lex Fridman episode, Perplexity CEO Srinivas admitted that they'd abuse Twitter's academic grant program and autogenerated thousands of grant applications with GPT.
If he so laughingly reminisces about doing that, I wonder what kind of other dodgy shit they do behind closed doors.
just my 2 ct
Unless it has already been done.
EDIT: You still have to rely on a server to index the web pages of interest.
Of course opinions vary on HN but generally it feels like a "ad-blocking is cool, paywall hopping is cool, but the server also has the right to not give us the information".
Given this, this is pretty sensational reporting by The Verge -- calling it theft, comparing it to crypto, breaking the "foundations of the internet" (... the Internet is about sharing) etc. This is a non-story, we already have these tools, this just adds an AI summary layer on top.
And in-text ads can already be removed by summarisation tools (even from before the AI era!) or by apps like Boring Report.
This is a story. With ad blockers, the assumption was that the original page as a whole is valuable, but users just don't want to pay the cost of the ads and other obtrusive elements. But now we're opening up the question of what exactly is it on that page that's of actual value - is it just the text, or only some of the text, or only the "facts" expressed on the page, or maybe something even more abstract, such as how the page corroborates a fact on another page.
I find this to be a really interesting question, both on a practical level and on an epistemological level - what is it that we as readers actually (want to) get from a resource?
“Leveraging” other people’s content is basically par for the course. Whether it’s training or google news or Google books or stability A.I. images it’s all doing the same just to different degrees
Perplexity is a bullshit machine
> Though Forbes has a metered paywall on some of its work, the premium work — like that investigation — is behind a hard paywall
No it's not, see for yourself! I can read it just fine...
https://www.forbes.com/sites/sarahemerson/2024/06/06/eric-sc...
We have seen other models that don't rely on protection flourish - Wikipedia, open source, scientific publications, open weights models, and even fashion. They all are permissive, and thriving.
And isn't Google stealing the content from the whole internet and making a huge profit from it? Why the double standard. Perplexity has links to sources too.
Pay-gating copyrighted content is different indeed from public-posted social media content. In the instance of the Forbes theft, Perplexity offered links to other copycats, but never the original [costly] reporting.
Just consider the fact that chatGPT as 180M users and probably solves 1B tasks per month. It collects contextual data, guidance and feedback on a massive scale, going into the next model with some PII and copyright mitigations.
We will go to LLMs for advice because they will have the best data, refreshed by hundreds of millions of users. LLMs will naturally collect the most up to date feedback from the world, they don't even need to do anything, just wait for users to copy paste the relevant data and provide iterative guidance.