Notes on Anthropic's Computer Use Ability
composio.dev
composio.dev
On one hand, it has really helped me with prototyping incredibly fast.
On the other, it is prohibitively expensive today. Essentially you pay per click, in some cases per keystroke. I tried to get it to find a flight for me. So it opened the browser, navigated to Google Flights, entered the origin, destination etc. etc. By the time it saw a price, there had already been more than a dozen LLM calls. And then it crashed due to a rate limit issue. By the time I got a list of flight recommendations it was already $5.
But I think this is intended to be an early demo of what will be possible in the future. And they were very explicit that it's a beta: all of this feedback above will help them make it better. Very quickly it will get more efficient, less expensive, more reliable.
So overall I'm optimistic to see where this goes. There are SO many applications for this once it's working really well.
god the future is here haha
A real killer app would be something that is adaptive and smart enough to deal with all the SEO/walled gardens in the travel search space, actually understanding the airlines available and searching directly there as well as at aggregators. It could also be integrated with your Airline miles accounts and all suggested options to use miles/miles&cash/cash, etc.
All of that is far more complex than .. clicking around google flights on your behalf and crashing.
Further, the real killer app is that it is bullet proof enough that you entrust it to book said best flight for you. This requires getting the product to 99.99% rather than the perpetual 70-80% we are seeing all these LLM use cases hit.
ThePointsGuy blog tried to implement something that directly tied into airline accounts to track milage/points and redemption options, but I believe they got slapped down by several airlines for unauthorized scraping. Airlines do NOT like third parties having access to frequent flier accounts.
Hence the "adaptive" part of my comment.
It really needs to be a client side agent.
And then add the complexity that you might be willing to pay cash if the price is right ... so then you add dozens more searches to that on potentially many websites.
All of this is "easy" and a solved problem but it's incredibly monotonous. And almost none of these services offer an API, so it's difficult to automate without a browser-based approach. And a travel agent won't work this hard for you. So how amazing would it be instead to tell an AI agent what you want, have it pretend to be you for a few minutes, and get a CSV file in your inbox at the end.
Whether this could be commercialised is a different question but I'm certainly going to continue building out my prototype to save myself some time (I mean, to be fair, it will probably take more time to build something to do this on my behalf but I think it's time well spent).
We don't need PBs of training data, millions of compute, and hours upon hours of training.
You could sit down a moderately intelligent intern as a mechanical turk to perform this workflow with only a few minutes of instruction and get a reasonably good result.
There will be enormous push towards steering these software agents towards similarly shady practices instead of making them act in the true interest of the user. The ads will be built into the weights of the model or something.
Is the endgame some free/cheap tool that abstracts away the entire ad based web economy to the benefit of end users?
Imagine something closer to a super duper smart useful Siri/Alexa that feeds you product recommendations, paid placement, and other ads interspersed with your actual request response.
Hey Siri what temperature is it? It's 45 and going to be chilly today, a North Face jacket might be handy today.. can I recommend you a few models? What's your size?
I mean, what really are ads and dark tactics (like those observed on accommodation booking websites)? They are ways to influence purchasing decisions.
If the decision is offloaded to AI, then logically ways to sway the AI decision will be developed. Such as backroom deals, hidden prompts and rules governing the assessment of the AI in making choices.
Imagine the most dystopian outcomes and you'll probably be closer than "well I don't have to see ads anymore!"
Sounds great. But corporations will find a way to fuck over their users for inverstors' gains in no time.
None of these agents are smart.
And if purchases become agentic, fine print or other shady tricks hidden in business terms will be how businesses draw consumers in.
Also, none of this will be existential, earth-shattering or enourmous until compute power per watt comes to a degree where all of this is economical at scale.
I wonder why leveraging accessibility tools for this wouldn't have been a better option. Browsers and operating systems both have pretty comprehensive tooling for accessibility tools like screen readers, and the whole point of those tools is to act as a middle man to programmatically interpret and interact with what's on screen.
Last I checked on it, maybe a year ago, there were browser proposals for standardizing the accessibility tree APIs but they were very early discussions and seemed pretty well stuck.
That would be a good reason for Anthropic using image processing here though, short of forking open source a11y tools there may not have been a simple way to use accessibility data to interact.
The end goal here is clear, being able to interface with anything available in the screen.
That approach is a weird one to me, though only as long as its limited to the current use. If this is just another test bed for a much more broad tool that could rely on accessibility APIs that makes sense.
I think I'm a blind user of the late 2024 Internet.
I use a screen reader for everything.
Now I'm paranoid that I'm just a computer use model, testing this a11y API hypothesis, in training.
Agents won't get anywhere because any user process you want to automate is better done by creating APIs and creating a proper guaranteed interface. Any automated "computer use" will always be a one-off, absurdly expensive, and completely impractical.
The possibilities are quite exciting, in fact, even though the technology isn't quite there yet.
Ideally it would be given a persona and a list of use cases, try to accomplish each task and save the state where you/it failed.
Something like a Chrome lighthouse but for usability. Bonus point if it can highlight what part of my documentation is using mismatched terminology making it difficult for newcomers to understand what button I am referring to.
Implementing tests is not the hard part. You could make that an intern project or hire a consultant for 3 months. The hard part is the interpretation of results.
That is - making a thing that spits out tickets/alerts is easy. The signal/noise tuning and actual investigation workflows are the hard part and still very manual & human operated. I don't see LLM mouse/keyboard control changing that yet.
I don't really believe that what I am asking for is hard, yet I still can't buy it as far as I know.
> actual investigation workflows are the hard part and still very manual & human operated.
Sure but it would allow your QA worker to have pre-tested usecase-based path with some flag on whether or not they may be problematic with a screen-recording and some timestamp of where it went wrong.
These will always need human-in-the-loop to vet the findings before cutting a ticket to development team.
I come more from a "big data" background, and have dealt with CTOs who think "can't we just use AI?" is the answer to data quality checking multi-PB data lakes with 1000s of unique datasets from 100s of vendors. That is - they don't want to staff a data quality team, they think you can just magic it all away.
The answer was always - sure, but you are fixated on the easy part - anomaly detection. Actual data analysis on what broke, when, how, why, and escalating to data provider was always 95% of the work. Someone needs to look at the exhaust, and there will be exhaust every single day.. so you can kill your dev teams productivity or actually staff an operations team responsible for the tickets the thing spits out.
I'm thinking for a single particular application under test and a mostly-static group of SMEs who might be involved to respond/tune
1. Pixels and screenshots (video really) and keyboard/mouse events is definitely the purest and most proper way to get agents working in the long term, but it's not practical today. Cost and speed are big obvious issues, but accuracy is also low. I found that GTP4o (08-06) is just plain bad at coordinates and bounding boxes and naively feeding it screenshots just doesn't work. As a practical example, another comment mentions trying to get a list of flight recommendations from Claude computer use and it costing $5, if my agent is up for that task (haven't tested this), it would cost $0.10-$0.25.
2. "feature engineering" helps a lot right now. Explicitly highlighting things and giving the model extra context and instructions on how to use that context, how to augment the info it sees on screenshots etc. It's hard to understand things like hover text, show/hide buttons, etc from pure pixels.
3. You have to heavily constrain and prompt the model to get it to do the right thing now, but when it does it, it feels magic.
4. It makes naive, but quite understandable mistakes. The kinds of mistakes a novice user might make and it seems really hard to get this working. A mechanism to correct itself and learn is probably the better approach rather than trying to make it work right from the get-go in every situation. Again, when you see the agent fail, try again and succeed the second time based on the failure of the previous action, it's pretty magical. The first time it achieved its objective, I just started laughing out loud. I don't know if I've ever laughed at a program I've written before.
It's been very interesting working on this. If traditional software is like building legos, this one is more like training a puppy. Different, but still fun. I also wonder how temporary this type of work is, I'm clearly doing a lot of manual work to augment the model's many weaknesses, but also models will get substantially better. At the same time, I can definitely see useful, practical computer use from model improvements being 2-3 years away.
Does this already exist? If not, would the benefits be lower than I think, or would the costs be higher than I think?
This _seems_ more like a normal user so clearly could not do anything nefarious. /s
Webwright is a front-end shell that presents to me; I'm suggesting a back-end shell that presents to Claude.
It doesn't appear that Webwright enables tool-use. In other words, there's no task-oriented feedback loop between AI-provided shell commands and the results of those shell commands. Please correct me if that's not right.
Just the other day someone used Claude to write a script to configure a server. It left a port open and the server was hacked hours later and used to attack other servers. Hetzner almost banned the hosting account.
Why didn't this project start with https://huggingface.co/meta-llama/Llama-3.2-11B-Vision
For the folks who are more savvy on the Docker / Linux front...
1. Did Anthropic have to write its own "control" for the mouse and keyboard? I've tried using `xdotool` and related things in the past and they were very unreliable.
2. I don't want to dismiss the power and innovation going into this model, but...
(a) Why didn't Adept or someone else focused on RPA build this?
(b) How much of this is standard image recognition and fine-tuning a vision model to a screen, versus something more fundamental?
Either we're going to use these tools to augment our abilities or basically just become wiped out, at least our jobs will be, and there is no plan to provide support for anyone. Maybe the tech will make the transition to a post employment world so swift we don't even feel any negative economic effects at all, but let's see.
(unless the robots of the future are like Bender)
Maybe this is the cliff , but it feels unlikely.
I think this is what differentiates the speed at which AIs have gotten from ok -> good -> great -> better than humans at say chess, versus say driving a car, summarizing a paper, understanding human requests, recommending music, etc.
I think a lot of people are extrapolating the rate of progress & possible accuracy rates from chess bots to domains that do not compare.
But "you" don't, that is precisely the point. The speed at which the gap between rich and poor grows keeps increasing, after all -- the rest is commentary --, and people who right now send people to die and murder in wars for oil, and what not not, will not suddenly start sharing when they fully captured all means of production for good. That's like hoping the person who keeps stealing your shit every chance he has, leaving you in sickness or death without a thought will give you a billion dollars once all the locks on your house have rusted off completely and you no longer have means to call the police.
QA...
We had one that was a simple download from here and login and upload there. Having the accounting team be able to automate that versus devops is huge.
TBH, while I giggle at the thought of anybody being replaced, I dont think it's likely, it's just that the standards and expectations have shifted in some domains. I think if anything LLM's raised the tide for everyone (in relevant roles) and we're all able to move a little faster now, like when we went from abacus to calculator a while back, just a different scale of magnitude.
Our own personal digital squire.
Then eventually we become assistants to AI.
1) I tried using it for QA for my SaaS but agent failed multiple times to fill out a simple form, ending with it saying the task was successfully completed.
2) It couldn’t scrape contact information from a website where the details weren’t even that hidden.
3) I also tried sending a message on Discord, but it refused, saying it couldn’t do so on someone else's behalf.
I mean, I’m excited for what the future holds, but right now, it’s not even in beta.
If anything, the examples involving moving the mouse to the address bar or getting csv's of results are very poor examples, because we can already do that much better without "computer use".