Using Python and Selenium for automated visual regression testing
github.com
github.com
My two cents: I've never seen automated visual regression testing that wasn't terrible to work with, and where the most common result (by a large margin) of a test failing was someone would update the expected file/image so it would pass with the new visuals. It's a hard problem, and one that I've personally decided isn't worth doing for the customer facing software that I've been involved with.
1. Figuring out if some change is actually wanted/expected
We found that doing a d-hash + hamming distance (levenshtein) is a very good way to handle tolerance, as it will ignore most text and subtle layout changes, but unlike most examples the screenshot is scaled down to 200px wide (this number is arbitrary, we're still figuring out how much tolerance it does provides).
2. How to handle this changes
This is the hard part. For now we are only selecting the images that failed in the previous step and by using some jenkins plugins they are presented during the pipeline side by side (old, diff, new) and are reviewed manually, but the versioning process is automatized and the results are stored in the changelog.
Ideally the tests should be maintained by the developers themselves so they know how and where to quickly change them to accommodate code and behavior changes.
There is a cost to a brittle test. UI testing suffers from this more than other kinds because there are many 'plausible' UI arrangements, and as the product shifts and changes, you need to distinguish
1. "The dialog moved slightly to the left"/"We refactored the HTML, but it still looks the same" from
2. "The dialog is now underneath another element".
Suppose it takes 10 minutes to find, fix, get reviewed, and push the fix, deploy the fix, and validate the fix. If the number of failures that are more like the former are RADICALLY more than the latter type of failure, then it is easy enough to come to the conclusion that the test is not giving you a reasonable ROI. Perhaps the cost of releasing a latter-type-of-failure is not that bad, if you can fix it and get it to production quickly. It might be cheaper overall then the ongoing maintenance cost on a test that would prevent this failure.
Also, and this is culture and product dependent, but we're talking about this like it's a single test, when it's usually a suite of tests (or multiple suites). If they have a 5% failure rate, and 95% of the time its really an 'update the test, this is the new expected', people will stop trusting the tests, and will take shortcuts. So you may find that instead of spending 100% of the maintenance costs for the suite of brittle tests, you're spending 70%, but only getting 20% of the benefit because people become accustomed to the failures , and once a test is failing, no one will notice that the failure changed from a 'benign failure' to a "customer-can't use" failure. (Note: numbers are imaginary, but not crazy).
At one company I worked for, it was so bad, that when we tried to introduce testing/checkin rigor, multiple developers pulled me over to make me explain "Why this dumb test is failing on my checkin attempt", and we would look at the logs and other artifacts to uncover that "It's failing because you changed something without updating the relevant tests". It took quite a bit of time to re-train developers used to brittle tests, to respect and maintain non-brittle ones.
And that is why I am against automated UI testing in general. :)
At the end of the day, if you can't get stable tests using the WebDriver W3C standard [0], you are doing something weird and overly complex with your web application. End to end testing isn't going to give you an objective right or wrong every time, but it should make you ask the question "Why does this happen?" and the answer is usually "Oh, we are doing something weird".
I don't know what specific problems you were having, so I can't give more specific advice than that.
The underlying architecture is a bad design for heavy javascript apps in my opinion. The roundtrips between the test runner talking to selenium server talking to selenium driver in a browser and back the other way is slow and so much can change on the page in between steps in your tests. Cypress runs your test code in the javascript process of your browser so I believe theres no or minimal roundtrip lag.
We use `waitFor()` for the UI to stabilize but thats been a hard mental model for devs to follow and as a result we have tons of unnecessary waits in the tests which slows them down. Even things like waiting for a loading modal to disappear before trying to interact with the UI is hard since your code:
`waitFor('.loading-modal', false) // wait for it NOT to exist`
may run BEFORE its even appeared, then fail in the next step when you try clicking on a button and the modal is there now. You can't wait for it to appear first to prevent that, as your code may run after its already come and gone too.
Tons of annoyances or strange behavior like chromedriver doesn't support emojis in inputs, ieDriver had select boxes set value broken at one point, setValue('123') on a field actually does something like 'focus, clear, blur, APPEND 123' so blur logic to set a default value of '0' on your field will result in the final value being '0123' in your tests… just the worst.
While valid that's not typical for many sites- what's th e point of a very short lived pop up? and even if it is part of your page, you can skip the "risky" part of the test and verify it otherwise (logs ? side effects ?) or not at all.
We didn't care about testing the popup at all, it was just breaking our other tests in the following way.
In our UI you could click a 'save' button, then a 'saving…' popup appears, meanwhile the 'save' button goes away and an 'edit' button appears (behind the popup), and when the response comes back it says 'Saved.' in the ui.
A test for `$('div=Saved.').toExist()` in wdio works, it does a waitFor under the covers and polls the UI until that text appears. It doesn't care if theres a popup shown or not.
However moving on to the next step in the tests, `$('button=Edit').click()` throws an error 'element is not clickable' if the popup is visible when it happens to runs. Doing multi-command steps like 'check if popup is there, if not click' in general doesn't work as theres so much latency between commands. You can inject javascript in to the page that does both checks in the browsers js process as a hacky workaround.
We did upgrade our webdriver library partly to get waitForClickable() which based on the name at least sounded like it handle the above, but there were no volunteers to update the 168 instances where waitForLoader() had spread in the codebase :/
You need firstly a framework that understands these technologies and then on top of that you want something that offers a structure to capture the common actions and build an automatable view of a page.
For these requirements I have used Geb [1] with the Page Object Model approach with reasonable success - but you still have to approach writing your tests as actual software, not ad hoc random scripting.
I hear it's great for unit testing javascript though.
This was all before the W3C standardized browser automation, so I guess the situation may have improved a bit.
It sounds like you were approaching the end-to-end tests as an add-on and not part of your production code. If your problem is changing selectors, you need to make sure your developers know that changing selectors is going to make the tests fail and you need to equip them to handle it through tooling.
Humorously enough though, people had never bothered to debug why the tests were failing. One reason was that a backend service crashed. It turned out that quickly deactivating users after doing things in the UI caused stuck threads that ran infinite loops. No one had ever looked at the logs.
- spent days/weeks refactoring tests to work with new testing framework - got tests 100% passing locally, passing in CI env - tests fail randomly on other people's machine or in CI - pick new framework, rinse repeat.
Tried cypress, webdriver, jest, etc. Total waste of time.
Last place I worked had a team that built a similar system using Selenium and some image diff. It worked 90% of the time. It was even integrated into their CI pipeline and the system would email you when it failed. You could make a change to one area of code and get an email from the system for a completely separate area that for whatever random reason failed. I once submitted some code to that project and when I got the email I asked one of the project maintainers what I should do to fix it. He told me to just ignore it. Their process was to check the output of the tool on those failures to see if there was any legitimate problem. When I checked the output myself a large amount of the output was garbage (about 100 application states tested and at least 10 in a garbled state).
My own team even tried to integrate Selenium and we even found an external vendor where you can ship your Selenium tests along with your app and they will run it against a matrix of browsers of various versions. We barely got a Chrome version running - no hope of Safari, Firefox or IE/Edge. It was a blackhole of time just fighting to get an equivalent of hello world running consistently as a test in all browsers.
One day someone will prove me wrong and get this kind of UI testing working reliably. But after nearly 20 years seeing optimistic people die on this hill - I do not support wasting more time on it.
It's by far the leading player in this space and for good reason. Selenium is unreliable and too generic.
Selenium is great. It really is. But the level of flaky tests that generally get produced is just so painful. And as you say, getting it to run in Firefox - that's not been too bad, but any other browser as well was a pipe(line) dream. IEDriver last time I tried to use it with Grid only allowed one instance anyway, so couldn't run multiple on one machine.
UIPath, AutomationAnywhere and these other 'no code required but we have a place you can write code in case' tools - I'm looking forward to seeing how they deal with these sorts of problems. When you're scraping data off a dynamically loading page, or is time dependent, it's not trivial. Even with implicit waits, retries and other 'tricks' as someone described further down, it's frustratingly difficult at times to just do a single action in an app/page.
And then you hit shadow dom websites, or Windows apps that create new windows (not child popups) for every screen/dialog box. The joy :)
Still, I find it an enjoyable, if frustrating challenge.
I am working on an open source library that generates Playwright tests (uses Chrome Dev Tools protocol) and I hope we can prove you wrong about getting UI tests working reliably. https://github.com/qawolf/qawolf
Which is why (shameless plug alert) we at Rainforest [1] are working with crowdsourced testers - humans are still much, much better at visual diffing and judgement than machines are. Feel free to shoot me an email if you're interested in talking about it and figuring out how much to automate and how much to leave to humans.
[0] Why do Record/Replay Tests of Web Applications Break?, Mouna Hammoudi, Gregg Rothermel, Paolo Tonella
My approach was to capture DOM element attributes and store them in a database, then compare snapshots.
https://github.com/gulkily/Selenium-Utilities/tree/master/sr...
Being able to have a customer accept what the output looks like and then listening to when the page changes would be great for giving non-technical people control over passing tests.
Anyone remember that project ?