Sikuli: Automate Anything You See on Screen
sikuli.org
sikuli.org
I was hoping that image based testing would eventually overtake Selenium/Appium for web/mobile testing. In principle there are several advantages:
+ It works the way human works. If your button id has change from #submit to #submit-application, your UI test should work just fine. However, if your button became one pixel at top-left corner, test should break. Right now most frameworks do the opposite.
+ If it's based on visual, it would be easier to maintain tests. E.g. Your button change color from light grey to dark grey, would you like to update your tests.
+ Tests would be much easier and faster to read and write. Even less technical folks like current manual QA testing can contribute to automate testing tools.
True, although for some things I do still like a visual test. Your users may be used to clicking a button that looks a certain way in a certain place, and suddenly shifting that might be worthy of requiring a change to a test. You wouldn't want all your behavioural tests done like this though.
> Maybe it makes sense to have one or two of these types of visual checks, but they should be automated so that regenerating them is easy.
I agree. Perhaps there's a nice integration between the two? If you were able to give, say, a css selector and a template image rather than just one or the other to look for you'd be able to report these errors:
* I can't find the button I was looking for, but there is something that matches the css selector [blah] that looks like this: [image] instead of [image]. Do you want to update the image in the tests? Y/N
* I can't find the css selector [blah] but the button appears here [image] with [class] and [id], do you want to update the css selector in the tests? Y/N
If the program is actually still fine then it'd be a quick auto-update, and if it's not then these bits of information are probably the first bit of debug you'd want to do anyway to figure out what broke the tests.
They both run in Java, so you can combine Sikuli with Selenium pretty easily [0], so long as you're not doing headless testing. I've played around with it a bit, using Selenium to get coordinates of a container element and then having SikuliX operate just on that region of the screen.
[0]: http://www.softwaretestinghelp.com/sikuli-tutorial-part-2/
A decent example of this would be that i've implemented frontend features before but did not notice an unintended pixel shift. I introduced a bug, albeit just visual, but a bug nonetheless - and i had no idea. Likewise, a minor visual issue like that can get through many stages of review and even deployment.
Pixel interaction may be terrible for early prototyping, but the idea has merit. Or, at the very least, perhaps it could identify by css/dom, and throw warnings for pixel alterations.
I like the idea of letting design people programmatically enforce standards (Though, no idea how to make them do that programmatically lol)
You do realize we standardize on testing to ID so that you can change everything else (location, parentage, etc.) because that's the stuff that benignly changes during development or for localization, right? The ID is supposed to be semi-permanent and not change in the normal case, so don't make trivial changes like that.
Otherwise, we're fully capable of finding buttons by coordinate, parentage, attributes, whatever. It's just crappy practice to do so. We intentionally lock down some things and not others because we know how software development works and what changes are likely to indicate greater chances of bugs and what aren't. We also code to accommodate different resolutions, adaptive interfaces, visual changes for different languages--your button location and boundaries change when your English 5 character caption becomes a German 25 character caption--etc.
As for the 1 pixel button, it's not just that it might be 1 pixel. If all you care about is static visual correctness, you can do simple bitmap comparison to get that, though that has its own issues. And we can build in basic implicit checks into our frameworks for sane boundaries, z-order occlusion, etc. It's common for control selection to also validate the control could actually be accessed by the user. The bigger problem is that the human interaction details might not work right, It might not visibly depress when you hit it, for example, or might have a hit target that doesn't correspond to its boundaries, or have a janky delay after tapping, or whatever.
Even injecting the most user-like of user events won't always catch those so there's no substitute for actually exercising the interface. And since you're in there anyway, the visual checks become trivial and mostly automatic.
On another subject, there's no way in hell I'll trust a system that takes guesses as to what it's looking at for a regression test suite, which is what nearly all GUI automation suites are. Even if I did use this tool for something less deterministic like model-based GUI fuzzing, I'd need to have another suite built that verified beyond a doubt that the button I expect to be there is there and at least superficially behaves how I expect it to behave. That requires discrete selection and determinism, not fuzzy identification and non-determinism.
The point of a regression suite is NOT to try to make the tests pass, it's to verify that absolutely none of the assumptions I coded into it changed and alert me to manually explore that functionality if they have by failing. Then I can bless the new assumptions with an automation code change or file a bug and xfail the test until the bug resolves. But the assumptions are supposed to be crystalized. You can't trust an alert system unless you know it'd make its decisions to alert the way you would. Unpredictability isn't welcome there.
In this particular case, an ID changing means you removed the control and added another one, conceptually (and possibly practically) speaking, so the automation is supposed to break to trigger a re-examination. Your own code might reference that ID too, you know, so just you having done that means you greatly increased the chances of there being new bugs.
What a pleasant surprise. Thanks you for sharing!
Here is a demo: http://i.imgur.com/1BpMKpe.gif Script is here: https://gist.github.com/eirikb/ac8196beb0b57577a8fc47eb18427... Images are here: http://imgur.com/a/Uv3NH
What it does: Opens Google translate in Chrome, inserts some Norwegian text, copies the translated text and inserts it into this comment. To make the images I used gnome-screenshot, it let me grab areas.
Sadly Windows only. It's one area where Windows has a leg up on Macs.
A cursory glance at macrocreator doesn't really give me the impression of what it does better, do you have any additional insight?
Another example, this is a macro to jump to the end of a file name in windows explorer and append today's date, then jump to the next item in the list (F3 is the trigger):
F3::
Macro1:
Send, {Alt}{f}{m}{End} ; shortcut to rename, jump to end of filename
SetKeyDelay, 0
SendRaw, _09-06-2016 ; this is what is sent
Send, {Enter}{Down} ; jump to next file
Return
This is extremely readable and easy to understand.Perhaps it's possible to do pixel searches in Automator (click on certain buttons for example), but AHK + Pulover's Macro creator make this extremely easy - this is why I mentioned both, as these combined make a perfect drop-in replacement for Sikuli - and they're far more stable than Sikuli, which was quite unstable when I tried it at year or two ago.
[1] more specifically, I found sikuli scripts pretty painful to debug, since neither java debuggers nor python debuggers seem to work quite right.
Yes, the original API was written in java, but Sikuli scripts are written in Jython. I did an entire automation project using Sikuli written in jython without any hiccups a couple of years ago.
Did you somehow manage to read my comment and interpret it as "Sikuli scripts are written in java"? Because that's not at all what I said (or meant) :P
Sikuli itself was (and still is) written in java - which I find a pain to work with when I want to modify the internals of the library; and then scripts are written in jython, which means I don't have access to my cpython-specific tools and extensions.
Robotic process automation (e.g. Blue Prism) is coming to the fore in commercial settings - a mixture of API and computer vision (probably a horrible mix) so there's a market out there for developing something commercially.
That led to me writing Sikuli4NET about a year later. https://sourceforge.net/projects/sikuli4net/
UIs also tend to change more than keyboard shortcuts.
Some time ago I was demoing a script to my boss and the second time I run it, it clicked on a different place with the same text (and fortunately with the same destination).
We were shocked because the link wasn't where we expected to be (Ciscosecure ACS 3.x).