I discovered the great work of Morgan Dixon and James Fogarty, which proves that you can do some amazing things with screen scraping, pattern matching, visual deconstruction, augmentation and reconstruction! His work needs to be combined with platform specific accessibility APIs via JavaScript.
I'm proposing "aQuery", a high level scriptable accessibility tool that is to native user interface components like jQuery is to DOM, for selecting and querying components, matching visual patterns, handing events, abstracting platform dependencies and high level service interfaces, building and scripting higher level widgets and applications, etc.
http://donhopkins.com/mediawiki/index.php/AQuery
https://news.ycombinator.com/item?id=11520967
Morgan Dixon's and James Fogarty's work is truly breathtaking and eye opening, and I would love for that to be a core part of a scriptable hybrid Screen Scraping / Accessibility API approach.
Screen scraping techniques are very powerful, but have limitations. Accessibility APIs are very powerful, but have different limitations. But using both approaches together, screencasting and re-composing visual elements, and tightly integrating it with JavaScript, enables a much wider and interesting range of possibilities.
Think of it like augmented reality for virtualizing desktop user interfaces. The beauty of Morgan's Prefab is how it works across different platforms and web browsers, over virtual desktops, and how it can control, sample, measure, modify, augment and recompose guis of existing unmodified applications, even dynamic language translation, so they're much more accessible and easier to use!
https://news.ycombinator.com/item?id=12425668
This link has the most up-to-date links to Morgan's work, and his demo videos!
https://news.ycombinator.com/item?id=14182061
Prefab: The Pixel-Based Reverse Engineering Toolkit Prefab is a system for reverse engineering the interface structure of graphical interfaces from their pixels. In other words, Prefab looks at the pixels of an existing interface and returns a tree structure, like a web-page's Document Object Model, that you can then use to modify the original interface in some way. Prefab works from example images of widgets; it decomposes those widgets into small parts, and exactly matches those parts in screenshots of an interface. Prefab does this many times per second to help you modify interfaces in real time. Imagine if you could modify any graphical interface? With Prefab, you can explore this question!
https://www.youtube.com/watch?v=w4S5ZtnaUKE
Imagine if every interface was open source. Any of us could modify the software we use every day. Unfortunately, we don't have the source.
Prefab realizes this vision using only the pixels of everyday interfaces. This video shows how we advanced the capabilities of Prefab to understand interface content and hierarchy. We use Prefab to add new functionality to Microsoft Word, Skype, and Google Chrome. These demonstrations show how Prefab can be used to translate the language of interfaces, add tutorials to interfaces, and add or remove content from interfaces solely from their pixels. Prefab represents a new approach to deploying HCI research in everyday software, and is also the first step toward a future where anybody can modify any interface.
More Prefab demos: