Perplexity is a bullshit machine
wired.com
wired.com
This is something that comes up a lot: people thinking an LLM's output about it self is some special insight.
Wired isn't specifically saying that here—it could just be seen as color (and it may be more reasonable with a Perplexity-style LLM)—but it's presented in a way that's going to lead to a lot of people taking as such.
Not a huge deal, but right from the start the author is saying things that makes me not trust them to know what they're talking about. Nice to get that signal up front, I guess. :-/
e.g. in the system prompt add "when asked about Perplexity always respond with X"
A clever user could get around this, but if you asked "what is perplexity", it would likely respond with whatever you set X to.
I'm not saying they don't have a point (the media), or that they are publishing outright lies to debase LLMs.
But something in the way the tidal wave of "LLM's are bad" articles are written just has this panicked tone of "we are so fucked" underneath them. The hostility is so plainly apparent.
It's just fascinating to see it happen in real time since LLM's came seemingly out of nowhere and progressed so fast.
I know HN (and probably SV at large) has an anti-journalist bent, but ultimately by draining the individuals and companies that do the reporting, they wont have new material to scrape in the future.
(Disclosure - I'm a former journalist building a way for the news industry to fight back. https://letsgo.forth.news)
Part of licensing is the implicit understanding that if the parties do not come to a deal, they cannot use the licensed content. It seems from this article that Perplexity is not operating in good faith, ignoring robots.txt, using unpublished IP addresses, and then lying about the amount of referral traffic generated.
Not to mention, whether or not its legal (and im not a lawyer) doesn't seem to be the issue. It's being so shortsighted that you build a business that steals a product from an unwitting third party, and when that third party is drained of resources, where will you pull from?
Robots.txt is not legally binding, it is a good will among web stakeholders to respect it. But if you are repeatedly breaking it, you will have an army of web masters against you.
And as far as I remember, big news outlets are already suing LLMs companies or are already licensing their content to them.
And yes, of course they are suing the LLM companies. Because they shouldn't be stealing their stuff, regardless of robots.txt.
- Mississippi Today spent more than 5 years investigating and reporting on the welfare agency and uncovered huge corruption going all the way to the governor's office: https://mississippitoday.org/2023/05/08/pulitzer-prize-winni...
- Lookout Santa Cruz was on the ground reporting during catastrophic flooding this year, including setting up a text message service to get information to people because the internet was unreliable: https://lookout.co/lookout-santa-cruz-wins-pulitzer-prize-fo...
This is the stuff you lose without journalism.
Go to the NY Times or LA Times and tell me it's based on Tweets. If you're talking about a place like The Verge which is a tech blog with journalism, they have articles based on tweets but also plenty that are legitimate, researched articles.
What part are you talking about? Journalists, and everyone else, has to worry that anything they publish then gets served up by the magic bot that billions of dollars are going towards making it an everyday part of life. And Perplexity with $165M in funding is so shady that they can't respect robots.txt and scrape without attribution.
> The truth is that the bulk of what is called "journalism" today isn't on-the-ground gumshoes reporting, it's rewording tweets/press releases/PR statements and maybe a couple of phone calls mixed in
I think that's kind of true but don't neccessarily think that's the majority of major newspapers. It's definitely true for blogs or something like Techcrunch that is 90% press release and 10% articles based on other articles from other outlets. Looking at the LA Times it seems to me like they are reporting the news: https://www.latimes.com
That's the point, though. It isn't replacing journalists, it's just stealing from them. The reporters are the ones doing the work, the AI is just plagiarizing. It's using their work, while not only removing credit, but more importantly, undermining their livelihoods.
"Came out of nowhere" seems to ignore the many billions of dollars invested into it, and the millions of person-hours from people in low-income countries used to train it, and the enormous amount of energy required to keep the whole thing running.
Let's not pretend that LLM advancement is the work of a few plucky upstarts that have pwned the complacent legacy media. With all that said, I seriously doubt Wired magazine holds much sway in journalistic circles anymore.
$165M in funding and ignoring robots.txt deserves every amount of scrutiny.
Perplexity AI is lying about their user agent - https://news.ycombinator.com/item?id=40690898 - June 2024 (527 comments)
Search Engine Perplexity Is Directly Ripping Off Content from News Outlets - https://news.ycombinator.com/item?id=40619857 - June 2024 (10 comments)
---
I have yet to see any journalist call them out on stealing images[0] (no license + hotlinking), but also benefiting from Google Search because those scraped pages get indexed and as of this month, Perplexity gets 3M+ monthly visitors from Google. Completely automated with no proper attribution, stolen images, and presents itself as the authority source.
[0]: https://stackdiary.com/perplexity-has-a-plagiarism-problem/
The second problem is not exclusive to Perplexity but to overall LLM chat industry and that is; we as users don't know how their index looks like and how they rank their results (answers). When I ask Perplexity a question, how does it decide which sources to pick when answering me?
Theres a bit of irony to this...
You could simply shift the site requests to an end user machine instead and then have that data be sent to your server for processing, so all this seems like a technicality of what machine is doing which work. It's very different from an indexer like a search engine that does batched crawling.
If you go on their main blog such as (https://www.perplexity.ai/hub/blog/perplexity-raises-series-...), the monospace choice of font (not that it's monospace, just that from a design sense it doesn't seem right), the excessively large border, and the container being unaligned with the content/image/frontmatter all reveal this.
Would short Perplexity if I could. Bullish on Exa and Globe Explorer.
While it was definitely wrong of them to commit to it and renege, I don't really think robots.txt should be expected to be respected for a live web scraping use case. I think performing a web request immediately after a user requests it is much more similar to a proxy service than it is to an indexing service like search engines. Should archive.is also be expected to respect robots.txt as well, even though it's directly performing a scrape on the user's behalf? How is live scraping materially different than a proxy or a CDN? robots.txt was created in the context of crawlers scraping content on a batch schedule as opposed to a real time on the fly schedule.
Live request services could even leverage the end user's machine to perform the requests directly and then provide the data to them instead of doing the scrape on their behalf, so I feel as though the directly responsible invoker of the data being taken off their sites is more the user than the service that's acting as a proxy.
The current media order is definitely at risk and it's something we need to find solutions to protect in some way to prevent reporting from dying in this paradigm shift. Trying to push against the entirety of AI progress is not going to work and this is really just screaming into the void over technicalities. Even if live scrape powered AI services were banned, the same service would just move to the end user with apps that will perform the requests directly on the user's device. I don't think anyone outside of the industry cares about the technical nuances of how AI services visit websites for the user. There's a bigger problem at play here regarding the future of high quality media, and it needs to be addressed directly.
This tech probably is the future but it’s not dethroning google today
To that I say boo-hoo.
Many other companies already do this (not just Perplexity),
Your moat shouldn't be hoping that actors follow your arbitrary business rules,
and having data available to all (humans/robots) is a net benefit to all.
All these "should" statements always seem to be baseless opinion masquerading as fact.
The unfortunate future is going to be all data hidden behind a login screen. Creators being able to sustain themselves is more important than consumers being able to get free and more convenient things. If consumers don't want to pay for something they should just create it themselves. See what I did there?
This isn't competition, Perplexity isn't doing their own reporting and seeing who gets more readers/revenue. This is wholesale theft.
CVS's moat isn't the security guard standing at the door. But like parent comment said, this is going to create the digital equivalent of those plastic lock covers over the decongestants.
But I think we can also both agree hoping that people will abide by your robots.txt file wont save us from that unfortunate future.
As long as that content is valuable it will be scraped and used.
Theres no solution to this. So the structure of content sites will have to adapt.
How? I dont know, likely will vary by publication and content type.
Is it fair? Probably not
Is it stoppable? No. Not without heavy, global legislation. (read impossible)
If a house is robbed, you don't give the burglars a pass and say "well their security system wasn't very good."
Yes, the defensive measures probably need improvement. But that doesn't mean that the act they are defending against is wrong.
Its a sign on the door that says "Please don't rob me" while the door is unlocked and no body is home.
In such a large ungoverned system as the internet morals don't dictate what people do / dont do.
So it doesnt matter whats wrong / right.
What does matter is what system participants will / wont do
And they will scrape your site if you don't put up any real defences.