OS-Copilot: Towards Generalist Computer Agents with Self-Improvement
os-copilot.github.io
os-copilot.github.io
In this "OS" situation, aside from not even simulating a virtual body, it's not clear that the inputs available to the agent are especially tied to the agent's ongoing actions. By comparison, what you see is a function of how you move your body, head, eyes and eyelids, and "seeing" involves saccades that let you take in the multiple important parts of a scene. So I think this doesn't even have an instantiation which acts _analogously_ to being embodied.
Many of the enterprise products in this category have used computer vision to help reduce brittleness, but for all of the improvements, these tools have remained highly error prone, not to mention extremely expensive.
The explosion of interest in the category across a broader community of researchers and developers seems like both a boon for the RPA space, and a major threat of disruption.
I would never have considered RPA, but I would build my own layer via LLM for my automation needs.
One of the core features of RPA products is centralized visibility and management of the lifecycle of automations. This will remain even if the entire interaction layer changes.
Something like the OS-Copilot layer lowers the barrier to entry though and I could easily see there being multiple choices for orgs looking to do this kind of automation in the future, and this is where I see the opportunity for disruption.
Most likely the RPA products will add all of the necessary buzzwords to keep people in buying roles interested and more likely to keep renewing.
I hope we get more open source options like Llava or CogVLM.
The paper makes it seem like there may be some way to control the mouse and read screen captures, but doesn't give any details. There is a GitHub link though so maybe it's in there.
This kind of project baffles my mind. RPA should be used for situations when there is no API available, and from what I understand of the ServiceNow product, it has all kinds of APIs for automation use cases.
But yes, RPA is big at our org right now. In ServiceNow too.
Because if so, you’re right on. The question is whether Sam knows it
*At certain thresholds, even solving the captcha correctly will falsely claim it’s wrong. Particularly if you’re not logged into a Google account.
I want to do a structured efficiency study of programming tasks of human vs human-plus-AI from problem statement through production-ready code. But my org doesn’t have enough devs to make it statistically significant, nor the spare human capacity to invest on duplicative tasks. I assume there must be some studies out there anyone have a reference?
I am eagerly looking forward to getting the chance to use these once matured and offload all my gruntwork.