Perplexity Deep Research
perplexity.ai
perplexity.ai
These things have the reasoning skills of a toddler, yet we keep fine-tuning their writing style to be more and more authoritative - this one is only missing the font and color scheme, other than that the output formatted exactly like a research paper.
Perplexity, OTOH, has almost completely replaced Google for me now. I'm asking it dozens of questions per day, all for free because that's how cheap it is for them to run.
The emergence of reliable tool use last year is what has sky-rocketed the utility of LLMs. That has made search and multi-step agents feasible, and by extension applications like Deep Research.
Yet what's essentially "cat [62 random files we googled] > prompt.txt" is now being confidently presented with academic language as "62 sources". This rubs me the wrong way. Maybe this time the new AI really is so much better than the old AI that it justifies using that sort of language, but I've seen this pattern enough times that I can be confident that's not the case.
That's not a very charitable take.
I recently quizzed Perplexity (Pro) on a niche political issue in my niche country, and it compared favorably with a special purpose-built RAG on exactly that news coverage (it was faster and more fluent, info content was the same). As I am personally familiar with these topics I was able to manually verify that both were correct.
Outside these tests I haven't used Perplexity a lot yet, but so far it does look capable of surfacing relevant and correct info.
I boycotted ai for about a year considering it to be mostly garbage but I’m back to perplexifying basically everything I need an answer fo
(That said, I agree with you they’re not really citations, but I don’t think they’re trying to be academic, it’s just, here’s the source of the info)
No, these AI companies are burning through huge amounts of cash to keep the thing running. They're competing for market share - the real question is will anyone ever pay for this? I'm not convinced they will.
There are plenty of open questions in the AI space around unit economics, defensibility, regulatory risks, and more. "Will people pay for this" isn't one of them.
The obvious crutch of this new AI stack reduced go-live time from 3 weeks to 3 days. Well worth the cost IMHO.
The leadership of every 'AI' company will be looking to go public and cash out well before this question ever has to be answered. At this point, we all know the deal. Once they're publicly traded, the quality of the product goes to crap while fees get ratcheted up every which way.
There seems to be such variance in the utility people find with these models. I think it is the way Feynman wouldn't find much value in what the language model says on quantum electrodynamics but neither would my mom.
I suspect there is a sweet spot of ignorance and curiosity.
Deep Research seems to be reading a bunch of arXiv papers for me, combining the results and then giving me the references. Pretty incredible.
I have to say I am really underwhelmed. It sounds all authoritative and the structure is good. It all sounds and feels substantial on the surface but the content is really poor.
Now people will blame me and say: you have to get the prompt right! Maybe. But then at the very least put a disclaimer on your highly professional sounding dossier.
Prove it or fight me
They're optimizing for the sales demo. Purchasing managers aren't reading the output.
But I admit that I didn’t spend huge amount of time designing the prompt.
Basically how progress of everything ever looks like.
The next huge jump will have to again make a qualitative change, such as enabling AI to handle a new class of tasks - tasks that fundamentally cannot be represented in text form in a sensible fashion.
But I do expect the next qualitative change to come from this area. It feels exactly like what is needed, but it somehow isn't there just yet.
Every time I see a comment about someone getting excited about some new AI thing, I want to go try and see for myself, but I can't think of a real world use case that is the right level of difficulty that would impress me.
Many of the AI companies ride on the hype are being overvalued with idea that if we just fine-tune LLMs a bit more, a spark of consciousness will emerge.
It is not going to happen with this tech - I wish the LLM-AGI bubble would burst already.
I ran Perplexity through some of my test queries for these.
One query that it choked hard on was, "List the college majors of all of the Fortune 100 CEOs"
OpenAI and Gemini both handle this somewhat gracefully producing a table of results (though it takes a few follow ups to get a correct list). Perplexity just kind of rambles generally about the topic.
There are other examples I can give of similar failures.
Seems like generally it's good at summarizing a single question (Who are the current Fortune 100 CEOs) but as soon as you need to then look up a second list of data and marry the results it kind of falls apart.
The first was Gemini Deep Research: https://blog.google/products/gemini/google-gemini-deep-resea... - December 11th 2024
Then ChatGPT Deep Research: https://openai.com/index/introducing-deep-research/ - February 2nd 2025
Now Perplexity Deep Research: https://www.perplexity.ai/hub/blog/introducing-perplexity-de... - February 14th 2025.
I, for one, am glad they are standardising on naming of equivalent products and wish they would do it more (eg. "reasoning" vs "thinking", "advanced voice mode" vs "live")
and BTW - I post an exact same spirit comment an hour ago... So I guess Today's copycat ethics aren't solely for products- but also for comment section . LOL.
Don't get me wrong - I don't mind to be copied on the Internet :), but I find this behavior quite rude, so I just mentioned it.
"Since google, everyone trying replicate this feature... (OpenAI, HF..) It's powerfull yes, so as asking an A.I and let him sythezise all what he fed.
I guess the air is out of the ballon from the big players, since they lack of novel innovation in their latest products."
I'd say the important differences are that simonw's comment establishes a clear chronology, gives links, and is focused on providing information rather than opinion to the reader.
We're mere months into these things, though. These are all version 1.0. The sheer speed of progress is absolutely wild. Has there ever been a comparable increase in the ability of another technology on the scale of what we're seeing with LLMs?
Google nor bing can find this
The only link I have found is a reproduction of the article[1], but I am unable to access the full text due to a paywall. I no longer have access to academic resources or library memberships that would provide access.
My Google search query was:
pussification of silicon valley inurl:upside
which returned exactly one result.I suspect the article's low visibility in standard Google searches, requiring operators like 'inurl:', might be because its PageRank is low due to insufficient backlinks.
[1] https://www.proquest.com/docview/217963807?sourcetype=Trade%...
Perhaps it’s softnuked in the eu or something?
I reckon it might be triggered by the word 'pussification' to refuse to return any results related to that.
If you're using a corporate account, it's possible that your account manager has enabled SafeSearch, which you may not be able to disable.
Local censorship laws, such as those in South Korea, might also filter certain results.
"Did you miss anything?"
"Can you fact check this?"
"Does this accurately reflect the range of opinions on the subject?"
Taking the output to another LLM with the same questions can wring out more details.
And their 2.0 Thinking model is great for other things. When my task matters, I default to Gemini.
I'm puzzled as to how that would work, when people talk about quick changes in model behavior. What exactly is being adjusted? The model has already been trained. I would think it's just randomness.
And fine tuning.
Choose your fighter...
High level overview: https://www.datacamp.com/tutorial/fine-tuning-large-language...
More detail: https://www.turing.com/resources/finetuning-large-language-m...
Nice charts: https://blogs.oracle.com/ai-and-datascience/post/finetuning-...
The big platforms also seem to employ an intermediate step where they rewrite your prompt. I've downloaded my ChatGPT data and found substantial changes from what I wrote. Usually for the better. Changes to the way it rewrites changes the results.
Now, as we can finally access it, Google has a chance to get back into the race.
I think the people who think anybody is close to OpenAI don't have pro subscription
I was interested in getting simple summary data on the outcome of the recent US election and asked for an approximate breakdown of voting choices as a function age brackets of voters.
Gemini adamantly refused to provide these data. I asked the question four different ways. You would think voting outcomes were right up there with Tiananmen Square.
ChatGPT and Claude were happy to give me approximate breakdowns.
What I found interesting is that the patterns if voting by age are not all that different from Nixon-Humphrey-Wallace in 1968.
I tried DeepSeek, it's fine, had some downtime, whatever, I'll just stick with 4o. Claude is also fine, not noticeably better to the point where I care to switch. OAI has my chat history which is worth something I suppose - maybe a week of effort of re-doing prompts and chats on certain projects.
That being said, my barrier to switching isn't that high, if they ever stop being close-to-tied for first, or decide to raise their prices, I'll gladly cancel.
I like their API as well as a developer, but it seems like other competitors are mostly copying that too, so again not a huge reason to stick with em.
But hey, inertia and keeping pace with the competition, is enough to keep me as a happy customer for now.
Also they just updated 4o recently, it's even better now. o3-mini-high is solid as well, I try it when 4o fails.
One issue I have with most models is that when they're re-writing my long scripts, they tend to forget to keep a few lines or variables here or there. Makes for some really frustrating debugging. o1 has actually been pretty decent here so far. I'm definitely a bit of a power user, I really try to push the models to do as much as possible regarding long software contexts.
You can also use tools like litellm and openrouter to abstract away choice of API
Luckily work just gave me access to ChatGPT Enterprise and O1 Pro absolutely smoked a really hard problem I had at work yesterday, that would have taken me hours or maybe days of research and trawling through documentation to figure out without it explaining it to me.
It’s a well documented Microsoft process but I didn’t even know where to begin as it’s something I hadn’t used before. I gave it the authorization policy (which was AND logic, and was async so it’d reject it any of them failed) said “how can I have this support lots of attributes” and it just straight up wrote the authorization filter for me. Ran a few tests and it worked.
I know this is basic stuff to some people but boy it made life easier.
I'm not particularly impressed with the examples they provided. Queries like "Top 20 biotech startups" can be answered by anything from Motley Fool or Seeking Alpha, Marketwatch or a million other free-to-read sources online. You have to go several levels deeper to separate the signal from the noise, especially with financial/investment info. Paperboys in 1929 sharing stock tips and all that.
Happened with ChatGPT - a chat oriented way to use Gen AI models (phenomenal success and a right level of abstraction), then code interpreter, the talking thing (that hasnt scaled somehow), the reasoning models in chat (which i feel is a confusing UX when you have report generators, and a better ux would be just keep editing source prompt), and now deep research. [1] Yes, google did it first, and now Open AI followed, but what about so many startups who were working on similar problems in these verticals?
I love how openai is introducing new UX paradigms, but somehow all the rest have one idea which is to follow what they are doing? Only thing outside this I see is cursor, which i think is confusing UX too, but that's a discussion for another day.
[1]: I am keeping Operator/MCP/browser use out of this because 1/ it requires finetuning on a base model for more accurate results 2/ Admittedly all labs are working on it separately so you were bound to see the similar ideas.
They also seem to ignore usurpers, like Anthroipic with their MCP. Anthropic succeeded in setting a direction there, which OpenAI did not follow, as I imagine following it would be a tacit admission of Anthropic's role as co-leader. That's in contrast to whatever e.g. Google is doing, because Google is not expressing right leadership traits, so they're not a reputational threat to OpenAI.
I feel that one of the biggest screwups by Google was to keep Gemini unavailable for EU until recently - there's a whole big population (and market) of people interested in using GenAI, arguably larger than the US, and the region-ban means we basically stopped caring about what Google is doing over a year ago already.
See also: Sora. After initial release, all interest seems to have quickly died down, and I wonder if this again isn't just because OpenAI keeps it unavailable for the EU.
They are the loudest dog, not the fastest. And they have the most to lose.
But for the query "what made the Amiga 500 sound chip special" it wrote a fantastic and detailed article: https://www.perplexity.ai/search/what-made-the-amiga-500-sou...
For me personally it was a great read and I learnt a few things I didn't know before about it.
Might have just gotten lucky, but as they say "this is the worst it will ever be"^
^ this is true and false. True in the sense that the technology will keep getting better, false in the sense that users might create websites that take advantage of the tools or that the creators might start injecting organic ads into the results
I'm just starting to wonder where we as the entrepreneurs end up fitting in.
Every majorly useful app on top of LLMs has been done or is being done by the model companies:
- RAG and custom data apps were hot, well now we see file upload and understanding features from OAI and everyone else. Not to mention longer context lengths.
- Vision Language Models: nobody really has the resources to compete with the model companies, they'll gladly take ideas from the next hot open source library and throw their huge datasets and GPU farm at it, to keep improving GPT-4o etc.
- Deep Research: imo this one always seemed a bit more trivial, so not surprised to see many companies, even smaller ones, offering it for free.
- Agents, Browser Use, Computer Use: the next frontier, I don't see any startups getting ahead of Anthropic and OAI on this, which is scary because this is the 'remote coworker' stage of AI. Similar story to Vision LMs, they'll gladly gobble up the best ideas and use their existing resources to leap ahead of anyone smaller.
Serious question, can anyone point to a recent YC vertical AI SaaS company that's not on the chopping block once the model companies turn their direction to it, or the models themselves just become good enough to out-do the narrow application engineering?
See e.g. https://lukaspetersson.com/blog/2025/bitter-vertical/
If suddenly agentic stuff works really well... Then that breaks that world. I think there's a chance it won't though. I suspect it needs a substantial innovation, although bitter lesson indicates it just needs the right training data.
Anyway, if agents stay coherent, my startup not being needed any more would be the last of my worries. That puts us in singularity territory. If that doesn't cause huge other consequences, the answer is higher level businesses - so companies that make entire supply chains using AI to make each company in that chain. Much grander stuff.
But realistically at this point we are in the graphic novel 8 Billion Genies.
Likely they throttle and do a lot of waiting for nothing during those five minutes. Can help with stability and traffic smoothing (using "free" inference during times the API and website usage drops a bit), but I think it mostly gives the product some faux credibility - "research must be great quality if it took this long!"
They will cut it down by just removing some artificial delays in few months to great fanfare.
Similarly, their inference GPUs have some capacity. Spreading out the traffic helps keep high utilization.
But lastly, I think there is just a marketing and psychological aspect. Even if they can have the results in one minute, delaying it to two-five minutes won't impact user retention much, but will make people think they are getting a great value.
I do wonder if this will push web publishers to start pay-walling up. I think the economics for deep research or AI search in general don't add up. Web publishers and site owners are losing traffic and human eyeballs from their site.
In the meantime, I hope the bean counters are keeping track of revenue vs LLM use.
It seems like chain of thought combined with search. Seems like it looks for 30 some references and then comes back with an overview of what it found. Then you can dig deeper from there to ask it something more specific and get 30 more references.
I have learned a shitload already on a subject from last night and found a bunch of papers I didn't see before.
Of course, depressed, delusional, baby Einsteins in their own mind won't be impressed with much of anything.
Edit: I just found the output PDF.
"How to do X combining Y and Z" (in a long detailed paragraph, my prompt-fu is decent). The sources it picked were reasonable but not the best. The answer was along the lines of "You do X with Y and Z", basically repeating the prompt with more words but not actually how to address the problem, and never mind how to implement it.
It's powerfull yes, so as asking an A.I and let him sythezise all what he fed.
I guess the air is out of the ballon from the big players, since they lack of novel innovation in their latest products.
Also, I'd compare with the output of phind (with thinking and multiple searches selected).
Real research needs several more levels of depth of contextual knowledge than the model is currently doing for any prompt. There is so much background information that people working in my field know. The model would have to first spend a ton of time taking in everything there is to know about the field and several related fields and then correlate the sources it found for the specific prompt with all of that.
At the current stage, this is not deep research but research that is remarkably shallow.
Reminds me of when Altman went to TSMC and bloviated about chip fabs to subject matter experts: https://www.tomshardware.com/tech-industry/tsmc-execs-allege...
It's very confusing.
Once chat gpt added web browsing, I largely stopped using perplexity
Lately they've been making a string of moves thought that smell of desperation though.
The guy at the helm also has a very weird body language/physiognomy, sometimes it seems he's just about to slip into a catatonic state.
I have no idea what made investors pour hundreds of millions into this guy/pitch, perhaps a charitable impulse? That money is dead, though.