Act-1: Transformer for Actions
adept.ai
adept.ai
https://mobile.twitter.com/sergeykarayev/status/156937788144...
Also this is very very cool, I love copilot, I hope I get to use this thing very soon.
We are spending a lot of time thinking about reliability and it's true that existing models fall a little flat here. I think ultimately the key to making this work really well is some combination of
a) collecting and training on human feedback and b) doing intelligent things to samples from the model after the fact
PS - please don't let this me used at a way to prevent human interaction. Chatbots are a disaster and literally the worst possible application of ML, as a shitty interface to a menu system. I hope this will be used in a way that is not consumer-hostile and that the company actively resists ignorant business attempts to use it to avoid paying for customer support.
And we are not building a chatbot, we're building something collaborative that you can work with to accomplish the stuff you want to do!
And was the feedback data used to train the model with reinforcement learning? Or did you request users to "correct" the action and get a supervised signal?
Can I make ACT-1 Sybil a few thousand people on mechanical turk?
Can I submit CVs with ACT-1 for entry-level full remote jobs and have it work for legacy companies, if those companies cannot setup ACT-1 themselves but provide a traditional human jobs interface?
Can I put an interface that interracts with the real world through controls and text on a webpage and have ACT-1 take a physical presence?
Are the example given in the blog post considered zero-shot learning?
Was the model trained on the websites in the examples given (e.g. on the Redfin site)?
How much labeled data was used?
Some people self host it, so it can solve some subset of the questions, the image understanding part is important, both it can understand a question with image e.g. if a user drops a fridge to search by image, or multiple images (e.g. which of the 10 images is the nicest looking fridge in 1 API request) as well. Also supports getting shared embeddings for images/text/code, which can be important for the information retrieval/question answering example where it needs to first find the relevant context on wikipedia then feed to the reader model to read it out
Also do other custom stuff like retraining etc. Thanks, Lee https://leepenkman.appspot.com/
When I search for content, I use key terms to produce refined and better results. If you don't use such terms then what you're looking for may be difficult to find.
Here's my source: https://dl.acm.org/doi/pdf/10.1145/1753326.1753333?casa_toke...
Wouldn’t something similar apply here though, where after using it for some time you get an inherent understanding of what it works well with and does not?
There’s always some sort of mix of things to learn and make available to the end user though, so that they understand how to use the tool to be successful.
Disclaimer: I also work at Adept
I'm game to take the other side of that wager.
My instinct from the last 5 years of advances in ML language research is that we're right at the cusp of having radically better natural language interfaces.
I chose a half-decade horizon because "Attention Is All You Need", the paper that introduced the transformer model, was published in 2017. Two of its co-authors are co-founders of Adept.
At this point ... think of google. How many tasks start with some vague statement of intent in a google search and then refinement by reviewing pages or better search queries: "How do I ..." Already many (most) internet users are using natural language to interface with the internet. Average don't go straight to wikipedia - they type a question and get a response where wikipedia is the first first or the knowledge graph except is sufficient. Or "plane ticket to $location" and get linked to airline sites (or the search engines interface).
We are much further along this path than many might suspect.
We just didn't have the tech that was able to take our high level instructions and carry them out for us like a human can. I think that is the long term goal of human computer interaction. This product seems like a significant step towards that.
For code-technical decisions, design documents, or the code itself provide a good record for the ever increasing details and scope, but we certainly haven't found the best way to capture everything. Perennial debates about code documentation, or what should go in PRs, or literate-style programing and more show we haven't figured it out. But at the same time, I definitely find myself having non-productive slack conversations, or PR threads or whatever, when my colleague (or myself) prompts: "hangout?" or similar to have call and talk through the issue. Something about the real-time conversation is able to cut-through written mis-communication very well - which then can be captured back in text for the benefits of other colleagues.
However, on your point I think natural language is the gateway to all those things. If you have a spec you will need to describe it in natural language first then use it.
Visual though is I think another way we should be able to communicate with a computer in the future and it seems like we are heading in that direction.
I guess how I imagine all this working is you provide all this high level instruction to the computer through writing/speaking as well as any diagrams and it figures it all out as well as queries you for any further clarification it needs. Just like interacting with your colleagues in day to day.
Non-verbal (body language) communication is a bigger part of our interaction than verbal communication (https://www.entrepreneur.com/leadership/you-dont-say-body-la...)
The highest standard of UX is the genie that does as one wishes.
The best interface is no interface.
Certainly not - language models (including multi-lingual models) map words to some kind of (300ish?) multi-dimensional concept space. I wonder if we can translate that back to some sort of symbolic representation that is much more precise than human language. Some kind of IR where we could compile human into representing but also program against. I suspect the early prolog people were attempting something like this but were very wrong about framing reasoning as logical deduction rather than a stochastic process.
“I want to travel from Seville to Berlin next October, avoiding weekends, for a two or three nights stay in a hotel by the river. Direct flights preferred.”
>Anyone who can articulate their ideas in language can implement them
I'd be shocked if even 10% of the users who can't navigate a GUI could accurately describe what they want the software to do. To the user who doesn't know they can use Ctrl-Z to undo, the first half dozen times the AI mangles their inherited spreadsheet might be enough to put them off the idea.
I agree with you that it won't be basic users, however, use anything long enough and you will become an expert.
This vision would fundamentally change how people interact with computers.
What I find more concerning would be people operating under misconceptions, or being more precise than needed, thus not actually accomplishing their objective with the introduction of irrelevant detail.
- OK here's my email
- Please select all pictures of taxis to prove you are not a robot
ಥ_ಥ
Seriously though, the potential is good. I see several things they're doing right that have the potential to distinguish them from competing offerings.A few questions:
1. I'm curious if you're representing the task-operations using RL techniques (as many personal assistant systems seem to be) or if this is entirely a seq2seq transformer style model for predicting actions?
2. Assumption: Due to scaling of transformers, I assume that this is not directly working on the image data of a screen, and instead is working off of DOM trees; (2a) is this the case? and (2b) if so, are you using purely linear tokenization of the tree or are you using something closer to Evoformer (AlphaFold style) to combine graphs-neural nets and transformers?
3. Have you noticed that learning actions and representations of one application transfers well to new applications? or is the quality of the model heavily dependent on app domain?
I noticed multiple references to data applications (Excel, tableau, etc.). My challenge is that large language models and AI systems in general are about to hit a wall in the data domain because they fundamentally don't understand data [1] [2], which will ultimately limit the quality of these capabilities.
I am personally tackling this problem directly. I'm tying to prove more coherent data-aware operations in these systems by building a "foundation model" for tabular data that connects to LLMs (think RETRO style lookups of embeddings (representing columns of data)). I have been prototyping conversational AI systems (mostly Q/A oriented), and have recently been moving towards task oriented operations (right now, transparently, just SQL executors).
There seem to be good representations of DOM tree/visual-object models that you all are working with to take reasonable action, however I assume these are limited in scale (N^2 and all), and so I am wondering if you have any opinions on how to extend these systems for data (especially as the "windowed context grows" (eg. an excel with 100k+ rows))?
[1] https://arxiv.org/abs/2106.03253 "Tabular Data: Deep Learning is Not All You Need" [2] https://arxiv.org/abs/2110.01889 "In summary, we think that a fundamental reorientation of the domain may be necessary. For now, the question of whether the use of current deep learning techniques is beneficial for tabular data can generally be answered in the negative"
1. There's a spectrum (sort of) between using full on RL techniques and just doing sequence modeling. We're trying to pick a reasonable place on that spectrum that lets us model whether things have gone well without doing too much fiddling.
3. It really depends on how closely related the domains are. I think it's safe to say that you should expect more transfer of abstract/high-level capabilities than nitty-gritty things related to the specific domain - that's part of why we're excited about training one big model to use all software tools.
I feel like some of this could one day be built using a shared model that understands HTML and JavaScript code etc with a few example prompts. Or maybe something that understands intent+a browser automation language like Selenium, if not then some custom input output language+training as adept alludes to.
If interested in building something like this also checkout https://text-generator.io which already pulls down links and images to analyse to generate better text so has a lot of the required parts