ScreenAI: A visual LLM for UI and visually-situated language understanding
research.google
research.google
Work-in-progress: https://github.com/OpenAdaptAI/OpenAdapt/pull/610
Makes you wonder if the goal of their captcha system was ever really to stop people from botting.
I can't wait for AI to become the ultimate ad-removal tool.
There might be an arms race, but the anti-ad side will win as long as there isn't a unilateral winner (strongest models, biggest platform).
There will be enough of a shake up to the current regime -- search, browsers, etc. -- that there is opportunity for new players to attack multiple fronts. Given choice, I don't think users will accept a hamstrung interface that forces a subpar experience.
We basically just need to make sure Google, Microsoft/OpenAI, or some other industry giant doesn't win or that we don't wind up living under a cabal of just a few players.
I'm already hopefully imagining AI agents working for us to not just remove advertising noise, but to actively route around all of the times and places we're taken advantage of. That would be an excellent future.
The competitor is Google.
And 90% of users spend 90% of their time in walled gardens like Instagram or TikTok anyway. They see built-in ads.
Do I need to say more?
I think either I'm crossing paths with you a lot, and am always stricken by so much trusting enthusiasm... - or there're many people who have the same dream you're mentioning, in which case, one of you should build this product :-)
I don't think that was the main goal, but rather for them to get a massive labeling dataset for training their models on the cheap.
0. https://github.com/THUDM/CogVLM?tab=readme-ov-file#gui-agent...
• Semantic change comparison between screenshots. Visual regression testing, where you prompt the model to ignore certain things instead of masking and where it labels the changes with a message, like "chart color changed" or "text shifted down by 3 pixels".
• Using plain English as test scenarios, instead of brittle WebDriver-like APIs.
• Autonomous agent fuzzing the application by free roaming the UI.
• RAG the design artifacts from Jira and Google Docs for more targeted feature exploration and test scenario generation.
• Automatic bug reports as the output of the above. Or even send a draft PR to fix an issue, while at it!
The more I think about use cases, the more it sounds like full software development automation. Late in game, we won't probably need software as it exists today at all. This feels like reading Accelerando again, but this time it's happening for real and to you.
P.S.: Didn't expect to see a cafe from Cyprus - Akakiko Limassol - used as the demo, I'll remember to visit next time I'm in the area :P
Congrats to all the folks involved :)
Unfortunately we can't compare it to ScreenAI directly since as far as I can tell it is not generally available. However ScreenAI does not appear to use a separate segmentation step, which we needed to implement in order to get good results.
https://private-user-images.githubusercontent.com/774615/320...
Regarding failure modes, we have yet to do extensive testing, but I've seen it confuse the divide and subtract buttons on the calculator before only once.
The results are not as good as GPT4-V or Gemini. I've posted the output for each of `gpt-4-vision-preview`, `gpt-4-turbo-2024-04-09`, `gemini-1.5-pro-latest`, and `claude-3-opus-20240229`. Claude is the only one who makes mistakes, at least on that test.
Looks useful to me for replicating some things. Good stuff!