ask a loaded, "filter question" I more or less know the answer for, and mostly skip the prose and get to the links to its sources.
Not to different from a lot of consulting reports, in fact, and pretty much of no value if if you’re actually trying to learn something.
Edit to add: even the name “deep research” to me feels like something defined to appeal to people who have never actually done or consumed research, sort of like the whole “phd level” thing.
For sure it's probably missing stuff that a well payed lawyer would catch, but for a project with zero budget it's a massive step up over spending hours reading through search results and trying to cobble something together myself.
Whereas with real legal advice, your lawyer will carry Professional Indemnity Insurance which will cover any costs incurred if they make a mistake when advising you.
As you say, it's a reasonable trade-off for you to have made when the alternative was sifting through the legislation in your own spare time. But it's not actually worth very much, and you might just as well have used a general model to carry out the same task and the outcome would likely have been much the same.
So it's not particularly clear that the benefits of these niche-specific models or specialised fine-tunes are worth the additional costs.
(Caveat: things might change in the future, especially if advancements in the general models really are beginning to plateau.)
I wrote it back when AI web search was a paid feature and I wanted access to it.
At the time Auto-GPT was popular and using the LLM itself to slowly and unreliably do the research.
So I realized a Python program would be way faster and it would actually be deterministic in terms of doing what you expect.
This experience sort of shaped my attitude about agentic stuff, where it looks like we are still relying too heavily on the LLM and neglecting to mechanize things that could just work perfectly every time.
My point was it's silly to rely on a slow, expensive, unreliable system to do things you can do quickly and reliably with ten lines of Python.
I saw this in the Auto-GPT days. They tried to make GPT-4 (the non-agentic one with the 8k context window) use tool calls to do a bunch of tasks. And it kept getting confused and forgetting to do stuff.
Whereas if you just had
for page in pages: summarize(page)
it works 100% of the time, can be parallelized etc.
And of course the best part is that the LLM itself can write that code, i.e. it already has the power to make up for its own weaknesses, and make (parts of itself) run deterministically.
---
On that note, do you know more about the environment they ran this thing in? I got API access (it's free on OpenRouter), but I'm not sure what to plug this into. OpenRouter provides a search tool, but the paper mentions intelligent context compression and all sorts of things.
Then I can further interrogate the information returned with a vanilla LLM.
I use it dozens of times per day, and typically follow up or ask refining questions within the thread if it’s not giving me what I need.
It typically takes between 10sec and 5 minutes, and mostly replicates my manual process - search, review results, another 1..N search passes, review, etc. Initially it rephrases/refines my query, then builds a plan, and this looks a lot like what I might do manually.
Besides I might give other large deep research models a try when needed.