Can I make ACT-1 Sybil a few thousand people on mechanical turk?
Can I submit CVs with ACT-1 for entry-level full remote jobs and have it work for legacy companies, if those companies cannot setup ACT-1 themselves but provide a traditional human jobs interface?
Can I put an interface that interracts with the real world through controls and text on a webpage and have ACT-1 take a physical presence?
Are the example given in the blog post considered zero-shot learning?
Was the model trained on the websites in the examples given (e.g. on the Redfin site)?
How much labeled data was used?
Some people self host it, so it can solve some subset of the questions, the image understanding part is important, both it can understand a question with image e.g. if a user drops a fridge to search by image, or multiple images (e.g. which of the 10 images is the nicest looking fridge in 1 API request) as well. Also supports getting shared embeddings for images/text/code, which can be important for the information retrieval/question answering example where it needs to first find the relevant context on wikipedia then feed to the reader model to read it out
Also do other custom stuff like retraining etc. Thanks, Lee https://leepenkman.appspot.com/
PS - please don't let this me used at a way to prevent human interaction. Chatbots are a disaster and literally the worst possible application of ML, as a shitty interface to a menu system. I hope this will be used in a way that is not consumer-hostile and that the company actively resists ignorant business attempts to use it to avoid paying for customer support.
And we are not building a chatbot, we're building something collaborative that you can work with to accomplish the stuff you want to do!
And was the feedback data used to train the model with reinforcement learning? Or did you request users to "correct" the action and get a supervised signal?