Automate any GUI using screenshots
sikuli.csail.mit.edu
sikuli.csail.mit.edu
One such example is to track real time images of a webcam pointed at a baby and using Sikuli to watch for a yellow dot placed on the baby's forehead. Another is to track movements of something across the screen; in this case a bus moving along Google maps.
I agree that there are better ways to do most of the things in their examples and that they should probably re-work their videos a bit, but just because this system doesn't solve your problems the way you want it to doesn't mean its useless.
I believe most people are warning about the demonstrated technology for this task:
> Sikuli is a visual technology to search and automate graphical user interfaces (GUI) using images (screenshots).
As I said below, for personal scripting of known applications I think this GUI automation is great. But upon seeing it, we had all hoped to see a technology which would get over our biggest hurdles in GUI development/testing. As it doesn't sound like it would do that (well) for a number of reasons (i.e. localization, themes, OS/app versioning, design changes during development, coloring, etc), it loses a lot of its practical application appeal.
It's still really cool technology and is a cool way to think about this kind of problem.
It sounds... scary. Like it will work well enough at first, and then explode when someone changes their desktop theme (especially icon theme), or wants to upgrade to a new version of whatever. Treating things as change-controlled APIs when they aren't just seems dangerous. Still I guess there's at least some amount of change control coming from platform conventions and human interface guidelines, and this comes closer to operating at the correct level of abstraction to benefit from that.
One problem with GUIs from the start has been that automating their behavior is fragile.
Automating actual keystrokes and mouse movements is the most fragile and kludgy of automations. Even drilling to the level of system messages is problematic.
Aside from it's other good qualities, the web is really nice for reducing possible user interactions to a very codified set of operations.
The article here points to ... the wrong way of doing things
That just leaves the other 90% of GUI changes that involve moving a setting to another screen, or changing the way an entire interaction works. ;-)
All the stuff you said about GUI automation being klunky is true. Still, this thing could come in handy as a kind of "GUI batch file," or perhaps as a tool to help produce screencasts.
I don't think any program guarantees that the GUI is a stable API! I do wonder about its ability to handle unexpected conditions, or for the scripter to know that these conditions exist in the first case. For example, if the GUI changes based on the day of week, the script will probably either get very confused, or do the wrong thing. I imagine there's some way to do conditionals (if you see this, then do thing 1, otherwise do thing 2) but it's hard for the scripter to discover all the possibilities. Still, might be useful for some things.
http://news.ycombinator.com/item?id=1067111
A snazzy title apparently makes all the difference.
'You don't sell the steak, you sell the sizzle'
Nonetheless, that's a cool paper and it's a shame it didn't make the front page first time round :-)
Autohotkey works, but matching by screenshots with computer vision would cut the amount of work required in half.
Bravo!
(Sometimes you should take a step back and ask yourself, "is looking for pictures on the screen really the best way to do this"? The example they show on the main page is a one-line "ifconfig" invocation, for example.)
Sometimes there's a long-term advantage to not doing things in a minimally-complicated way.
The actual program logic is tested by the usual integration/unit tests, not by clicking buttons in the GUI. (And if you are worried about "what if clicking button foo doesn't run function foo", then you need to write more tests for the GUI generator library, not for your application.)
It’s a Rube Goldberg machine, but playing with those is always fun and intuitive.
(By the way, that was all very explicit and non-automatic, would it be possible to do this in some way where you just hit record and then click your way through? You would just have to snap screenshots whenever the user is clicking. Might be a problem with things that change their appearance on mouse-over like drop down menus, but I guess that’s solvable.)
How many people fit that description, but will still find this easier than ifconfig? My guess is few.
That said, I think this is a neat hack.
I know people who are just as much afraid of the terminal as they are of the network settings. And stepping someone through opening the terminal and typing something is approximately the same number of actions (in clicks and keystrokes) as navigating to the settings panel and typing in the info.
And while I agree that the example of modifying the network settings is contrived and thus not a good example of the usefulness of this setup, the difference between modifying network settings via the system settings GUI and from the terminal with ifconfig is that one will be remembered after a reboot and one will not be despite both having similar surface complexity to change. This is mainly because the actual settings are buried in some obscure, OSX specific file somewhere that most likely isn't even editable without the GUI (same with Windows (buried in the registry) and with Gnome's Network Manager (buried in ~/.gconf/system/networking/connections)).
But that being said, the plist file format isn't all that great (there was a time when Apple had binary "compiled plist" files, but the pl command no longer supports that) -- the alternating "key" and value tags only associated by order in the file is kind of anti-XML. On one of my Xserves with two interfaces active, this file is 9k, massive for its purpose. This is not something that even an experienced administrator would want to edit using anything other than the GUI, doing so would be quite error prone. So if you want your settings to stick, you should be doing it from the GUI on OSX. The complexity of doing it from the GUI in a way that sticks and from the command line in a way that doesn't stick are roughly the same, but the results are different: to get it to stick by editing from the command line is quite a bit more complex.
All in all I love the way the idea works right now, although Java feels less than elegant on the Mac.
*(er, although for me it's got a killer bug - using the hotkey to make a screenshot does not work, gives no option of cancelling... hardcore crasher in my book)
The screenshot approach this tool takes is very unique. My only criticism is that, judging by the video, the image processing approach seems slow compared to an autohotkey's script.
What I'm really waiting for is a tool that can take this one step further and do OCR on any on screen text. This would make it easy to interact with gui's that present text that can't be read using system api's - imho that would be the holy grail of gui automation.
- Some tech support situations where you have to have a user do x amount of steps on their computer that are the same for all users. Sort of like an automated Geek Squad.
- Sell a prepackaged GTD style organization system that creates all the folders for you in the right places, downloads files (pre-made budget spreadsheet for example) into them, etc. (trivial, but it's a pain point for people)
- Make a bunch of different productivity apps that mimic the steps a professional programmer/ photographer/ marketer etc does when they first setup a new computer (bookmarks, preference settings, etc.)
Because it uses literal images, it seems like any change in OS theme, OS version, app version, localization (e.g. text or control shape), or colors (e.g. high contrast mode) would break the scripts.
It'd be neat to use for GUI automation during software development except for the fact that the GUI changes, the button wordings are tweaked, etc.
In all of these cases, back-end or OS GUI automation is probably better, but if you have an unchanging environment or want a quick on-the-fly test, the screenshot approach is novel and probably a bit cooler.
Little lesson in creating a good video demo....
Get to the point.
Then provide more videos for details.
(I guess you could say this should be expected from an MIT project website)
There is big money in tools like that, but I can tell you, its a real PITA to write test scripts using tools such as these. Given the option, you are better off exposing your app's object model to a scripting language, and letting testers script it like that.
Obviously that doesn't work for third-party or legacy apps. So it definitely has a market. And their computer vision algorithms have to be better than the godawful bitmap comparison tools that QTP used.
The demo (automatically setting an IP) is a one-time job. How many times do we have to do this task? So there is no need for me to automate those kind of jobs. But having said that, this could still be useful in some use cases. One example I could think of is testing desktop apps.
The demo video is proof of concept; make sure you read the paper.
Tnx
http://www.cs.washington.edu/homes/lsz/papers/slpz-cacm00.pd...
I think it's a more ambitious vision because it talks about programming by example systems, which automatically generalize from examples.
But Sikuli is very cool even if it's generalization is limited to ignoring minor differences in images.