The Deep Research problem
ben-evans.com
ben-evans.com
I have a decent idea of where to look to find comp information for a given municipality. But there are a lot of Chicagoland suburbs and tracking documents down for all of them would have been a chore.
Deep Research was valuable. But it only did about 60% of the work (which, of course, it presented as if it was 100%). It found interesting sources I was unaware of, and assembled lots of easy-to-get public data that would have been annoying for me to collect that made spot-checking easier (for instance, basic stuff like the name of every suburban Village Manager). But I still had to spot check everything myself.
The premise of this post seems to be that material errors in Deep Research results negate the value of the product. I can't speak to how OpenAI is selling this; if the claim is "subscribe to Deep Research and it will generate reliable research reports for you", well, obviously, no. But as with most AI things, if you get paste the hype, it's plain to see the value it's actually generating.
No it’s not. It’s that it’s oversold from a marketing perspective and comes with some big caveats.
But it does talk about big time savings for the right contexts.
Emphasis from the article:
“these things are useful”
Of course there are straightforward ways, in terms of UX, to make verification orders of magnitude easier – i.e. inline citations – but TFA argues that OpenAI isn't quite there yet.
Ultimately, if a research agent requires us to verify significant AI synthesized-conclusions as TFA argues, I'd argue research agents actually haven't automated tricky and routine work that keep us thinking about our research at lower level than we would like.
> As with all LLM output, all of these things are presented in the same fluid, confident-sounding style: you have to know the material already to realize when your foot has gone through what was earlier solid flooring.
Deep dive.
"Yes, that's exactly right"
It gets in the trenches, braves the cookie popups, email signups, inline ads and overly clever web design so I don’t have to. That’s enough to forgive its attempts to create research narratives. But, I hope we can figure out a way to train a heaping spoonful of humility into future models.
You dreamed of this? Why not dream of a web where you don’t have to brave a veritable ocean of crap to get what you want? It may surprise you to learn such a web existed in the not too distant past.
> and "websites" just turn into data-only endpoints made for AI to consume.
As is already the case with humans, that only serves users to the extent that the websites' veracity is within the intelligence's ability to verify — all the problems we've had with blogspam etc. have been due to the subset of SEO where people abuse* the mechanisms to promote what they happen to be selling at the expense of genuine content.
AI generated content is very good at seeming human, at seeming helpful. A "review website" which is all that, but with fake reviews that promote one brand over the others… a chain of such websites that link to each other to boost PageRank scores… which are then cross-linked with a huge number of social media bots…
Will lead to a lot of people who think they're making an informed choice, but who were lied to about everything from their cornflakes to their president.
* Tautologically, when it's not "abuse", SEO is only helping the search engine find the real content. I've seen places fail to perform any SEO including the legitimate kind.
>A "review website" which is all that, but with fake reviews that promote one brand over the others… a chain of such websites that link to each other to boost PageRank scores… which are then cross-linked with a huge number of social media bots…"
functionally different than a majority of news outlets in the US parroting the exact same story, verbatim?
This is extremely dangerous to our democracy. https://www.youtube.com/watch?v=_fHfgU8oMSo
* please note, i am glibly linking that youtube video as a completely transparent example of what i am talking about. There are many (many) more examples. don't read into the content of the video so much, and just think about the implications.
AI is automation, and the existence of fully automated propaganda doesn't deny the existence of manual propaganda before it.
> This is extremely dangerous to our democracy
Indeed, the existing manual kind of propaganda is already dangerous. Always was.
Even so, manual propaganda can at least be fought by grass-roots movements, by humans being human. This is why freedom of speech is valuable.
Hard for real humans to counter an AI that can be even moderately conversational at a cost of just USD 0.99/day to fully saturate someone's experience of the world.
But ads can be put in AI.
And since AI is so great and you don't want to bother going to the website and clicking through a cookie banner, as another commenter mentioned, you just won't ever know the competitor exists
I guess that's technically an ad but it's so much more subversive that "ad" doesn't really do it justice. It'll be like product placement but worse somehow
Have you dealt with people paying for ads? They want your rapt attention, not some possibly subliminal suggestion. The "good" thing about chat AIs is that they can, with a little work, make it nearly impossible to block. You want to use it for free, you get ads.
The summaries are probably still wrong but you do you, at least this would save you the step of reading bullshit and boiling a pond to generate a couple links
I think Deep Research (and tools like it) offer an even stronger illustration of that same effect. Anything that can produce a well-formatted multiple page report with headings and citations surely must be of PhD-level intelligence, right?
(Clearly not.)
Google Gemni models seem to lead...hopefully the metrics aren't being gamed.
Deep research will dramatically improve as it’s a process that can be replicated and automated.
By conflating technology’s evolving development path with a basic exponential decay function, the analogy overlooks the crucial differences in how innovation actually happns.
And many haven't
> A mathematician and a physicist agree to a psychological experiment. The mathematician is put in a chair in a large empty room and a beautiful naked woman is placed on a bed at the other end of the room. The psychologist explains, "You are to remain in your chair. Every five minutes, I will move your chair to a position halfway between its current location and the woman on the bed." The mathematician looks at the psychologist in disgust. "What? I'm not going to go through this. You know I'll never reach the bed!" And he gets up and storms out. The psychologist makes a note on his clipboard and ushers the physicist in. He explains the situation, and the physicist's eyes light up and he starts drooling. The psychologist is a bit confused. "Don't you realize that you'll never reach her?" The physicist smiles and replied, "Of course! But I'll get close enough for all practical purposes!"
Is that it? Is it sexist because the physicist and mathematician are attracted to the naked woman?
It's this mismatch which has contributed heavily towards society's whiplash over the last decade.
Presumably Deep Research has a bunch of weird multi-LLM-agent things going on, maybe there's something about their architecture that makes it more likely for mistakes like that to creep in?
https://www.ben-evans.com/benedictevans/2025/1/the-problem-w...
Claude gets that question right: https://claude.ai/share/7bafaeab-5c40-434f-b849-bc51ed03e85c
ChatGPT treats a PDF upload as a data extraction problem, where it first pulls out all of the embedded textual content on the PDF and feeds that into the model.
This fails for PDFs that contain images of scanned documents, since ChatGPT isn't tapping its vision abilities to extract that information.
Claude (and Gemini) both apply their vision capabilities to PDF content, so they can "see" the data.
I talked about this problem here: https://simonwillison.net/2024/Jun/27/ai-worlds-fair/#slide....
So my hunch is that ChatGPT couldn't extract useful information from the PDF you provided and instead fell back on whatever was in its training data, effectively hallucinating a response and pretending it came from the document.
That's a huge failure on OpenAI's behalf, but it's not illustrative of models being unable to interpret documents: it's illustrative of OpenAI's ChatGPT PDF feature being unable to extract non-textual image content (and then hallucinating on top of that inability).
This is an unfortunate example though because it undermines one of the few ways in which I've grown to genuinely trust these models: I'm confident that if the model is top tier it will reliably answer questions about information I've directly fed into the context.
[... unless it's GPT-4o and the content was scanned images bundled in a PDF!]
It's also why I really care that I can control the context and see what's in it - systems that hide the context from me (most RAG systems, search assistants etc) leave me unable to confidently tell what's been fed in, which makes them even harder for me to trust.
Code is one thing, but if I have to spend hours checking the output, then I'd be better off doing it myself in the first place, perhaps with the help of some tooling created by AI, and then feeding that into ChatGPT to assemble into a report. By showing off a report about smartphones that is total crap, I can't remotely trust the output of deep research.
I don't share this enthusiasm, things are better now because of better integrations and better UX, but the LLM improvements themselves have been incremental lately, with most of the gains from layers around them (e.g. you can easily improve code generation if you add an LSP in the loop / ensure the code actually compiles instead of trusting the output of the LLM blindly).
And note, the system is now directly competing with "interns". Once the accuracy is competitive (is it already?) with an average "intern", there'd be fewer reasons to hire paid "interns" (more expensive than $200/month). Which is maybe a good thing? Fewer kids wasting their time/eyes looking at the computer screens?
Of course, humans also make mistakes. There is a percentage, usually depending on the task but always below 100%, where the work is good enough to use, because that's how human labor works.
There is a huge gap from 85% to 99.99%.
but replacing that with a random number(/token) generator is more reliable to someone, then more power to them.
there is value to be had in the output of this tool. but personally i would not trust it without going through the sources and verifying the result.
They can make mistake in understanding something and will be able to explain those mistakes in most cases.
LLM and Human mistakes ARE NOT same.
Categorically not true and there’s so many examples of this in every day practice that I can’t help but feel you’re saying this to disprove your own statement.
Tell me, does an LLM know when it lies?
An LLM doesn't know when it lies, but a human also doesn't know when they are (innocently) wrong.
You won't give any job to that kind of person.
I insist that human hallucinations are NOT SAME as LLMs.
You might be amazed but most probably very shocked.
The "deep research" features were much more effective at getting me to pay for both subscriptions than in any valuable data collection. The former, I suspect, was the goal anyway.
It is very concerning that people will use these tools. They will be harmed as a result.
Compared to what exactly? The ad-fueled, SEO-optimized nightmare that is modern web search? Or perhaps the rampant propaganda and blatant falsehoods on social media?
Whoever is blindly trusting what ChatGPT is spitting out is also falling for whatever garbage they’re finding online. ChatGPT is not very smart, but at least it isn’t intentionally deceptive.
I think it’s an incredible improvement for the low information user over any current alternatives.
And there are plenty of logical problems that many humans can’t solve. Does that mean they’re not capable of reasoning?
At what point would you say something has reasoning? I’d argue that it’s more about how good something is at reasoning, rather than saying it is or isn’t capable of reasoning in absolute terms.
AI slop already produces many plausible-sounding articles used as infotainment and in academia. We already know this slop adds much noise to the signal and that poor signal slows actual research in both cases. But until now, the slop wasn't masquerading specifically as research! It was presented as an assistant, which provides no accuracy guarantees. “Research” by the word’s common meanings does.
This is why it will do harm. There is no doubt in my mind. And I believe OpenAI knows it. They have quite smart engineers, certainly clever enough to figure it out.
And if you think that all published “research” was guaranteed to be accurate before AI tools became available, then I think you should start looking more critically at sources yourself.
AI companies promising their LLMs will now do “research” won’t help.
And research that’s done outside of the academia (like business or independent thinker research) will be more muddied, with more people misled.
The reality is that for now it is not possible to leave the human out of research, so I think the best LLM can only help curate sources and synthesize them, but cannot reliably write sound conclusions.
Edit: this is something elicit.com recognized quite early. But even when I was using it, I was wishing I had more control over the space over which the tool was conducting search.
It seems to be at intern level according to the author - not bad if you ask me, no?
Did he try to proceed as with an intern? ie. was it a dialogue? did he try to drop in this article into prompt and see what comes out?
For skeptics my best advise is – do your usual work and at the end drop in your whole work with prompt to find issues etc. – it will be net positive result, I promise.
And yes they do get better and it shouldn't get dismissed – the most fascinating part is precisely just that – not even their current state, but how fast they keep improving.
One part which always bothers me a bit with this type of arguments – why on earth are we assuming that human does it 100% correctly? Aren't humans also making similar mistakes?
IMHO there is some similarity with young geniuses – they get tons of stuff right and it's impressive however total, unexpected failures occur which feel weird – in my opinion it's a matter of focused training similar to how you'd teach a genius.
It's worth taking step back and recognizing in how many diverse contexts we're using (like now, today, not in 5 years) models like grok3 or claude3.7 – the goalpost seem to have moved to "beyond any human expert on all subjects".
Even if they do the math right and find the data you ask for and never make any “facts” up, the sources of the data themselves carry a lot of context and connotation about how the data is gathered and what weight you can put on it.
If anything, as LLMs become a more common way of ingesting the Internet, the sources of data themselves will start being SEOed to get chosen more often by the LLM purveyors. Add in paid sponsorship, and if anything, trust in the data from these sorts of Deep Research models will only get worse over time.
It is in many ways a workaround to Google's SEO poisoning.
Doing very deep research requires a lot of context, cross-checking data, resourcefulness in sourcing and taste. Much of that context is industry specific and intuition plays a role in selecting different avenues to pursue and prioritisation. The error rates will go down but for the most difficult research it will be one tool among many rather than a replacement for the stack.
But the article goes into exactly how Deep Research fell exactly for the same SEO traps.
These deep research things are a waste of time if you can't trust the output. Code you can run and verify. How do you verify this.
1. Only do searches that result in easily verifiable results from non-AI sources.
2. Always perform the search in multiple products (Gemini 1.5 Deep Research, Gemini 2.0 Pro, ChatGPT o3-mini-high, Claude 3.7 w/ extended thinking, Perplexity)
With these two rules I have found the current round of LLMs useful for "researchy" queries. Collecting the results across tools and then throwing out the 65-75% slop results in genuinely useful information that would have taken me much longer to find.
Now the above could be seen as a harsh critique of these tools, as in the kiddie pool is great as long as you're wearing full hazmat gear, but I still derive regular and increasing value from them.
My current research workflow is:
* Add sources to NotebookLM
* Create a report outline with NotebookLM
* Get Perplexity and/or Chatgpt to give feedback on report outline, amend as required.
* Get NotebookLM and Perplexity to each write their own versions of the report one section at a time.
* Get Perplexity to critique each version and merge the best bits from each.
* Get Chatgpt to periodically provide feedback on the growing document.
* All the while acting myself as the chief critic and editor.
This is not a very efficient workflow but I'm getting good results. The trick to use different LLMs together works well. I find Perplexity to be the best at writing engaging text with nice formatting, although I haven't tried Claude yet.
By choosing the NotebookLM sources carefully you start off with a good focus, it kind of anchors the project.
Maybe good for wider subject areas, longer reports, or where some editorial nuance helps.
I do that a lot, too, not only for research but for other tasks as well: brainstorming, translation, editing, coding, writing, summarizing, discussion, voice chat, etc.
I pay for the basic monthly tiers from OpenAI, Anthropic, Google, Perplexity, Mistral, and Hugging Face, and I occasionally pay per-token for API calls as well.
It seems excessive, I know, but that's the only way I can keep up with what the latest AI is and is not capable of and how I can or cannot use the tools for various purposes.
I combine all the 'slop' from the three of them in to Gemini (1 or 2 M context window) and have it distill the valuable stuff in there to a good final-enough product.
Doing so has got me a lot of kudos and applause from those I work with.
For work tasks, there are several different variants of Gemini that are tuned for different things, just as OpenAI has.
I'll admit I'm surprised you need to combine all these LLMs to get a decent result on this kind of queries but I guess you go deeper than what I can imagine on these topics.
We built an alternative to do Deep Research (https://radpod.ai) on data you provide, instead of relying on Web results. We found this works a lot better in terms of quality of answers as the user can control the source quality.
That's probably how we should all be using LLMs.
Isn't this a very valuable product in itself? Whatever happened to the phrase "When there is a gold rush, sell shovels"?
Plenty of humans regularly make similar mistakes to the one in the Deep Research marketing, with more overhead than an LLM.
Ithink more certainty communication could help. Especially when they talk about docs or 3rd party packages etc. Regularly even Sonnet 3-7 just invents stuff...
It works surprisingly well. It also provides its reasoning about the quality of the sources, etc. (This is using GPT-4o of course, as it's the only mature GPT with web access)
I highly recommend adding this to your default prompt.
What do you mean by this exactly? That it makes you feel better about what its said, or that its assessment of its answer is actually accurate?
It's a conversation with AI, it's good to know its thought process on how certain it is of its conclusions, as it isn't infallible not is any human.
Arguably this applies just as well to Bayesian vs Frequentist statisticians or Molecular vs Biochemical Biologists.
I dont Google anything. Google maps, yeah. Google, no.
Everything I want to know is much better answered by ChatGpt Deep Research.
Ask a Question, Drink a Chai, Get a Great, Prioritised, structured Answer without spam or sifting through ad ridden SEO pages.
It is a game changer, and at one point the will get rid of the "drink a chai" wait and it will kill the Google we know now.
I was trying to convey that it had found some sources that, if I came across them naturally, I probably would have immediately recognized as fringe. The sources were threading together a number of true facts into a fringe narrative. The AI was able to get other sources on the true facts, but has no common sense, and I think ended up producing a MORE convincing presentation of the fringe theory than the source of the narrative. It sounded confident and used a number of extra sources to check facts even though the fringe narrative that threaded them all together was only from one site that you'd be somewhat apt to dismiss just by domain name if it was the only source you found.