I really, really like this new product/API offering. Still crashes quite a bit for me and obviously makes mistakes, but shows what's possible.
For the folks who are more savvy on the Docker / Linux front...
1. Did Anthropic have to write its own "control" for the mouse and keyboard? I've tried using `xdotool` and related things in the past and they were very unreliable.
2. I don't want to dismiss the power and innovation going into this model, but...
(a) Why didn't Adept or someone else focused on RPA build this?
(b) How much of this is standard image recognition and fine-tuning a vision model to a screen, versus something more fundamental?