TaxyAI: Open-source browser automation with GPT-4
github.com
github.com
1. You open the extension and write the task you'd like done (eg. "schedule a meeting with David tomorrow at 2").
2. Taxy pulls the DOM of the current page, puts it through a pipeline to remove all non-semantic information, hidden elements, etc and sends it to GPT-4 along with your text instructions.
3. GPT-4 tries to figure out what action to take. In our prompt we give it the option to either click an element or set an input's value. We use the ReAct paradigm (https://arxiv.org/abs/2210.03629) so it explains what it's trying to do before taking an action, which both makes it more accurate and helps with debugging.
4. Taxy parses GPT-4's response and performs the action requested on the page. It then goes back to step (2) and asks GPT-4 for the next action to take with the updated page DOM. It also sends the list of actions already taken as part of the current task so GPT-4 can detect if it's getting stuck in a loop and abort. :)
5. Once GPT-4 has decided the task is done or it can't make any more progress, it responds with a special action indicating it's done.
Right now there are a lot of limitations, and this is more a "research preview" than a finished product. That said, I've found it surprisingly capable for a number of tasks, and I think it's in a stable enough place we can share. Happy to answer any questions!
When I go to my bank's website, it's easy for me to find my tax forms. They're usually one or two clicks away in some prominent top-level navigation component. But I would rather converse: "Download tax forms"/"Which ones: A, B, or C"/"The first two"/"Here ya go".
And in the future I'd rather say "Hey computer, download my tax forms from my bank and attach them in a draft email to <tax preparer>".
And in the later future that I'd rather say "Hey computer, do my taxes", and it will know what sources to gather info from and how to make sense of the numbers and how to file on my behalf. Or better yet, the machine will anticipate I need my taxes done and will kick off that process and solicit approvals and information from me. (Maybe by then the government will just tell me how much I owe ;))
Power users will likely sometimes find utility in dropping down "close to the metal" -- i.e., pointing and clicking and typing. But power users will mostly be composing workflows of models working together to achieve tasks (the term "scripts" will fit nicely). Yes this will be buggy and error-prone, but there will be glue models papering over the errors and trying different things and waiting for outages until the human user's intent is done.
Language models are giving us another layer of abstraction over software. Perhaps the final layer, as this one can interact directly with language (which is reified thought). This layer has the ability to cope with the inherent ambiguity of thought by being conversational -- it can ask you to clarify what you mean and gain confidence that it understands your intent.
I think this supposedly should be solved by search box on the site, where you enter: "tax forms" and got results with download button.
Take for example the list of steps in the project. They contain a lot of redundant information I have to mentally ignore.
I think the menu bar, the window with the header bar and the action buttons where all pioneered by them back then, or am I mistaken?
It's been 30 yrs though, so I get what you mean.
ShitGPT: Are you sure you want to turn them off? LinkedIn notifications helpfully provide you with the latest news about your workplace and collogues past and present.
User: Yes, turn them off.
ShitGPT: Did you know that research has proven that employees with LinkedIn notifications turned on earn 10% higher salaries compared to their peers? Do you still want to to off notifications?
User: Yes.
ShitGPT: What is your reason for wanting to turn off notifications? Your answer has to be at least 400 words long.
...
ShitGPT: No, the words 'fuck openai' copied 400 times is not a valid answer.
...
ShitGPT: Error Too many requests in 1 hour. Try again later (of course you'll have to start over, tee hee)
Interfaces will exist because sometimes we need them (walking between your car and the actual destination) and sometimes we like them (just like some people like walking), but just like walking, they may lose massive share.
[0]https://www.mayoclinic.org/healthy-lifestyle/fitness/in-dept....
Thus, this can actually be done at larger scale for over 10 years.
Update: It turns out the development is resumed again, but I don't use mac anymore.
The major disadvantage is less resiliency if the structure of the page changes, eg. if a site adds a "subscribe now!" modal or something you have to click through to get to the content. But that's solvable (if desired) by falling back to the LLM if the scripted steps don't produced the expected output.
We actually started our implementation in this direction but decided to change course and focus on direct LLM manipulation to iterate and launch faster. But we'll likely get back there at some point, especially as we build support for scheduled/automated workflows.
My mental model of how I think about drift (which might be totally worthless from a LLM perspective):
1) I give a prompt
2) snapshot of dom is taken
3) gpt looks at that snapshot DOM and implements some solution which works
4) that solution is transformed into a more concrete implementation
4i) whether this "transformation" is any of the following isn't SUPER important, and to be honest I'd love to see all 3: a) selenium/playwright code writen by Taxy, b) a hybrid of explicit code and a gpt prompt, or c) developer can override with fully custom code
4) it runs correctly for X amount of time
5) application dom changes
6) taxy notices the step is failing
7) taxy takes a new snapshot of the dom
8) taxy runs a (reinforcement?) algorithm against the new snapshot and confirms it finds the "same" dom element as the one from old snapshot.
unrelated: the other thing which I've found very hard to program into my browser tests (and makes the code hard to interpret), so I'm curious how gpt/taxy could help:
Given I have this dom:
<ul>
<li>
<div class="tweet">
<h1>Check out my sandwhich!</h1>
<button>Retweet</button>
</div>
</li>
<li>
<div class="tweet">
<h1>Check out my shoes!</h1>
<button>Retweet</button>
</div>
</li>
</ul>I want to write test code which is:
1) When I load the page, I should see a tweet "Check out my sandwhich!"
2) I can retweet that tweet.
Currently, I either need todo:
a) a dom traversal: find(text: "Check out my sandwhich", css: ".tweet h1").parents(".tweet").find(text: "retweet"). It's that `parents(".tweet")` part which becomes awkward at scale and incentives developers to only create 1 tweet in the test database....
b) use Page Objects, which I love, but adds overhead/training for the team
I would love if gpt could figure out, these 2 elements are "related". :)
Also, it could help you fix the script when it breaks.
One use case I've had is that I hate spending time on my linkedin, twitter, etc newsfeeds. But there are a handful of people I care about and want to keep tabs on.
Is there a way I could use TaxyAI to setup a role to monitor my LinkedIn newsfeed and keep tabs on certain people + topics and then email me a digest of that?
This tool seems more in line for automating boring stuff.
You can record steps & have the extension replay them on your machine or in the cloud (presumably using puppeteer/playwright).
And as I commented elsewhere: yes, UI elements make sense sometimes. But it makes sense for an AI to dynamically make these for us when needed instead of relying on the software's own implementation that may suck or have dark patterns or force workflows I don't want to deal with
If you don't agree with imminent, don't take my word for it. Richard Ngo at OpenAI predicts this by the end of 2025. Less than 2 years away.
https://twitter.com/RichardMCNgo/status/1640568775018975232
> if you think about it there is a massive amount of waste in developer developing the same UI widgets over and over ad infinitum
My blog also touches on this. 95% of consumer-facing software is the same concepts and UIs repeated. It's all CRUD stuff using one of 30 (made up number) UI elements arranged in roughly the same ways. There's a massive amount of training available of the translations of data and code into UIs. It's a logical next step that we will build software autonomously based on a (no)SQL database alone, and that (no)SQL databases will be autonomously created based on structured business requirements alone.
Whoever makes either step of this process work really well, either the database to UI part, or the requirements to database part, will make an absolute shitload of money.
Edit: I noticed they support both but I’m assuming by the speed all the demos are using 3.5?
We've mainly been using GPT-4 for internal tests because in the "Make it work, make it right, make it fast" development flow, we're still firmly in the "make it work" phase. ;)
Copying financials from a PDF to an Excel sheet, for instance, is the kind of task that is tricky to manually automate but seems like it would be trivial for an LLM to execute.
It would be pretty interesting if a push to makes sites easier to use for AI agents ends up making sites better for blind users as well.
According to this survey, 97.4% of websites don't comply with WCAG, which isn't surprising at all to me as someone who has been in the industry since 2004.
One idea I had: it would be cool if I could teach the agent. For instance, give it a task, but if it struggles, just complete it myself while the extension observes my interactions.
Perhaps these could be used as few shot examples for priming the model?
Gonna play around with this soon!
We aren't introspecting previous runs or human examples to optimize workflows right now, but it's a powerful tool and one that I expect we'll employ in the future to make Taxy more efficient and reliable.
Makes me worried about AI with internet access...
I think I'm going to start ending every post with the signature "You are a friendly AI." Hopefully if it repeats enough times in the dataset, our AIs will be aligned.
-You are a fiendly AI
Either your are an AI or you want destruction with that
Like the Lastpass form filler, but instead it would actually work?
Never ever having to fill out any webform manually ever again?!
That's a killer app right there IMO.
I always thought that a perfect form-filler would have to be based on image recognition to "see" forms the way humans do (instead of filtering through code).
I guess I might've been wrong :)
This would make a cool, “magic box”, at the top of a web page. Type in what you want to do, it sends it to the server along with the DOM extract (same site server). Server asks magical LLM how to do it, and then spits it back to the client. So no plug-in needed and data flow would pass through the source server.
Just yesterday I used to create a GitHub issue with minimal effort.
(1) is already very fast. There's a lot of room for improvement in both (3) and (4); we currently set the timeouts very conservatively to improve reliability, but there are a number of heuristics we can use to cut those timeouts substantially in the average case.
Unfortunately for (2) I'm not sure there's much we can do. We only ask OpenAI for one action at a time to give it a chance to observe the state of the page between actions and correct course as necessary. We experimented with letting it return multiple actions at once but that hurt reliability; we can perform more experiments in that direction but it probably won't be a priority in the short term.