Cradle: Empowering Foundation Agents Towards General Computer Control
baai-agents.github.io
baai-agents.github.io
https://github.com/BAAI-Agents/Cradle/pull/44/files#diff-3f3...
Similarly, when I'm typing this post my output isn't "keyboard interrupts" nor "TCP frames". Those things do happen, but only by the grace of others' work.
Still I like the basic framework idea & I suppose future work looks at pushing the provided layer lower & lower - so the system can learn those atomic commands, rather than having humans provide them - and pushing the logic higher & higher via AI-generated composite commands.
> So this approach can't be generalised, can it?
The paper definitely demonstrates that the _approach_ can be generalized because they use the same approach across a variety of different environments and tasks, but you can also see that they did have to specialize the prompts, tools (like how the incoming screenshot was decorated with object detection / segmentation, stuff like that), and set of skills for each environment.
Perhaps something like the eye tracking tech in modern vehicles to ensure you're paying attention if the lane assist is turned on.
Of course, that would be awful. But what other recourse is there?
Charge money to the content consumer? Of course, that will be unacceptable for companies that make money from ads…they will prefer the dystopia. It will also be bad for us humans when our income dwindles because the machines take our jobs…which this paper shows it’s just a matter of time.
Other than the obvious path of getting over this fear of bots and whatnot, I see 2 options forward in this regard: 1) government ID verification on all major platforms 2) end-to-end verification of all software being ran, and refusal to run other programs if unsigned code is present; I'm pretty sure there's plenty of efforts in this area already
To be clear, I agree that this is a losing battle, but I don't think that's going to stop some interested parties in pushing systems like what we're talking about.
Rewatching Jurassic Park?
Also, neither here nor there but I enjoyed the discussion in the paper about how the model had a surprisingly low performance on sending an email in Outlook because while it well-understood the task and how to send an email, Outlook's UI still managed to confuse it - can relate.
Models like this will be useful for RPA.
The authors have developed Cradle, a multimodal-LLM-powered agent framework with six modules: Information Gathering, Self-Reflection, Task Inference, Skill Curation, Action Planning, and Memory. Once Cradle has processed high-level instructions, its inputs are sequences of computer screenshots. Its output is executable code for low-level keyboard and mouse control, enabling Cradle to interact with any software and complete long-horizon complex tasks without relying on any built-in APIs:
Oversimplified Big-Picture Diagram
+------------+
| Cradle | executable code
screenshots -> |(high-level | -> for controlling
| planning) | keyboard & mouse
+------------+
The authors' experiments show what to me looks like impressive generalization and performance across software applications, successfully operating daily software like Chrome and Outlook, and across commercial video games: It is able to follow the main storyline and complete 40-minute-long missions in Red Dead Redemption 2, create a city of a thousand people in Cities: Skylines, farm and harvest parsnips in Stardew Valley, and trade and bargain to make a profit in Dealer's Life 2.There are of course many caveats -- the technology is still in its infant stage -- but still, I'm impressed at how quickly things are progressing.
We sure live in interesting times!
No one is prepared.