Open Deep Research
github.com
github.com
Few points:
- open Deep Research is not a production app, but it could easily be productionized (would need to be faster + good UX).
- As the GAIA score of 55% (not 54%, that would be lame) says, it's not far from the Deep Research score of 67%. It's also not there yet: I think the main point of progress is to improve web browsing. We're working on integrating vision models (for now we've used a text browser developed by the Microsofit autogen team, congrats to them) because it's probably the best way to really interact with webpages.
- Open Deep Research is built on smolagents, a library that we're building, for which the core is having agents that write their actions (tool calls) in code snippets instead of the unpractical JSON blobs + parsing that everyone incl OpenAI and Anthropic use for their agentic/tool-calling APIs. Don't hesitate to go try out the lib and drop issues/PRs!
- smolagents does code execution, which means "danger for your machine" if ran locally. We've railguardeed that a bit with our custom python interpreter, but it will never be 100% safe, so we're enabling remote execution with E2B and soon Docker.
It's something I'd be happy to explore a bit if it's of interest.
Those remote interfaces may also work with local VMs for isolation.
Alternatively, PyPy is actually fully sandboxable.
On Linux, you can also use `seccomp.` See, for instance, https://healeycodes.com/running-untrusted-python-code
oh super cool! i've usually heard it the other way - people develop LLM-friendly web scrapers. i wrote one for myself, and for others there's firecrawl and expand.ai. a full "text browser" (i guess with rendering?) run locally seems like a better solution for local agents.
if you are just parsing the text you’ve lost a ton of information encoded in the layout/formatting.
that doesn’t even yet consider actual visual assets like graphs/images, etc
> On GAIA, a benchmark for general AI assistants, Open Deep Research achieves a score of 54%. That’s compared with OpenAI deep research’s score of 67.36%..Worth noting is that there are a number of OpenAI deep research “reproductions” on the web, some of which rely on open models and tooling. The crucial component they — and Open Deep Research — lack is o3, the model underpinning deep research.
1. running things in production/self hosting is more annoying than just paying like 20-200/month
2. openTHING makers often overhype their superficial repros ("I cloned Perplexity in a weekend! haha! these VCs are clowns!") and trivializing the last mile, most particularly in this case...
3. long horizon planning trained with RL in a tight loop that is not available in the open (yes, even with deepseek). the thing that makes OAI work as a product+research company is that products are never launched without first establishing a "prompted baseline" and then finetuning the model from there (we covered this process in https://latent.space/p/karina recently) - which becomes an evals/dataset suite that eventually gets merged in once performance impacts stabilize
4. that said, smolagents and HF are awesome and I like that they are always this on the ball. how does this make money for HF?
---
[1]: i think opendevin/allhands is a pretty decent competitor to devin now
Apple probably wouldn't accept a 3rd-party AI app store on MacOS and iOS, except possibly in the EU.
If antitrust regulation leads to Android becoming a standalone company, that could support AI competition.
OpenAI pretty much never acknowledge prior art in their marketing material because they want you to believe they are the true innovators, but you should not take their marketing claims for granted.
Completely agree that a real RL pipeline is needed here, not just some clever prompting in a loop.
That being said, it wouldn’t be impossible to create a “gym” for this task. You are essentially creating a simulated internet. And hiding a needle is a lot easier than finding it.
Working on complex problems tends to explode in to a web of things that needs done. You need to be able to separate these in to subtasks and work on them semi-independently. In addition when a subtask gets stuck in a loop, you need to work on another task or line of thought, and then come back and 're-run' your thinking to see if anything changed.
Instead, RL just rewards the model when it accomplishes some measurable goal (like winning the game). This works for certain types of problems but it’s pretty inefficient because the model wastes a lot of time doing stuff that doesn’t work.
This is an important point. As people rely more on AI/LLM tools, reliability will become even more critical.
In the last two weeks, I've heavily used Claude and DeepSeek Chat. ChatGPT is much more reliable compared to both.
Claude struggles with long-context chats and often shifts to concise responses. DeepSeek often has its "fail whale" moment.
Which reliability problems did you face? I heard about connection issues due to too much traffic with Deepseek, but those would go away if you self-host the model.
Also, having an open source alternative - even if slightly worse - gives projects an alternative if OpenAI decides to cut them off or sth.
What I tell startups is to start off with whatever OpenAI/Anthropic have to offer, if they can, and consider switching into Open Source only if it allows for something they can’t do (often in gen graphics), or when they reach product market fit and they can finetune/train a smaller model that handles their specific case better and cheaper.
What makes gen graphics stand out?
What style? Is it a texture? If so, I’ll need a model that can generate a large tiled image. Is it a logo? What kind? Is it badged, vintage, corporate, etc
When I first looked into stable diffusion I wanted to see what people were making with it and there was a few sites that showed stuff people had generated and it was 70% porn 29% high fantasy wallpapers (numbers are illustrative). Recently I've been looking at different text generation inference platforms like tgi and vllm on reddit, and 1 in 5 posts in the localllm subreddits are "What's the best model for erotic role play?"
My question was, porn (or other censored thing) could be in text, too. I don't understand why those who want censored content are primarily interested in graphics.
gemini for llm
brave/duckduckgo for search
jina reader for reading a webpageThen OpenAI announced theirs on the 2nd: https://openai.com/index/introducing-deep-research/
Ethan Mollick called Google's undergraduate level and OpenAI's graduate level on the 3rd: https://www.oneusefulthing.org/p/the-end-of-search-the-begin...
And now this. I can't stop thinking about The Onion Movie's "Bates 4000" clip:
Is this a new cottage industry? Are they making money?
- OS process
- virtual machine
- LLM inference
Could have longevity as PC master race meme template.Firecracker has changed the nature of “VMs” into something cheap and easy to spin up and throw away while maintaining isolation. There’s no reason not to use it (besides complexity, I guess).
Besides, the entire rest of this is a python notebook. With headless browsers. Using LLMs. This is entirely setting silicon on fire. The overhead from a VM the least of the compute efficiency problems. Just hit a quick cloud API and run your python or browser automation in isolation and move on.
So do humans, or can my friend with cerebral palsy not use the internet any longer?
That's irrelevant. Humans totally hate CAPTCHAs and they are an accessibility and cultural nightmare. Just forget about them. Forget about making better ones, forget about what AI can and can't do. We moved on from CAPTCHAs for all those reasons. Everyone else needs to.
Open source is already at the finish line.
Any public comparisons of OAI Deep Research report quality with Perplexity + DeepSeek-R1, on the same query?
How do cost and query limits compare?
I reran some of these searches and I've so far found OpenAI Deep Research to be superior for technical tasks. Here's one example:
https://chatgpt.com/share/67a10f6d-28cc-8012-bf98-05dcdb705c... vs https://www.genspark.ai/agents?id=c896d5bc-321b-46ca-9aaa-62...
I've been giving Deep Research a good workout, although I'm still mystified if switching between the different base model matters, besides o1 pro always seeming to fail to execute the Deep Research tool.
You mean when it says it's going to research and get back to you and then ... just doesn't?
Well, in that particular case the open source version was actually here first three month ago[1].
[1]: https://www.reddit.com/r/LocalLLaMA/comments/1gvlzug/i_creat...
Granted, this isn't an entrepreneurial venture, so maybe it has some value. but this still stinks to high heaven of saving money rather than producing value. One day we'll see AI produce stuff of value that humans can't already do better (aside from playing board games), but that day is still a long way off.
Good grief, AI researchers need to learn basic humility if they want to market their tech successfully. Unless they're toppling our states and liberating us from capital (extremely difficult to imagine) I have a difficult time imagining any value or threat AI could provide that necessitates this level of drama.
And for the big AI companies, like OpenAI, it has the very beneficial side-effect of establishing the narrative that lets them influence politics into regulating their potential competitors out of the market. Because they are, of course, the only reasonable and responsible builders of self-described doomsday devices.