Web automation: Don't use Selenium, use Playwright
new.pythonforengineers.com
new.pythonforengineers.com
start_chrome('github.com/login')
write('user', into='Username')
write('pw', into='Password')
click('Sign in')
I am the author of Helium. It's a wrapper around Selenium. It's fully open source. Under the hood, it uses Selenium 3, which is old. But boy does it work beautifully.The readme mentions iframes but not that.
Yes, your copy might change, but that is also a change in spec (whereas a change in DOM is not always a change in spec).
A lot of the selenium code I have run up on people try and get the exact element they want on the first try and repeat the process several times instead of doing something like searching for a div, form, or some parent element that has all their elements then all they need is to search relative xpaths from that element.
Also the issue with Selenium isn't with finding elements on a page
Obviously such an automation runs much slower than any Selenium based automation, but it its sooo much easier to create and maintain, especially on complex websites with tons of Javascript and frameworks. I believe this is the future of causal (test) automation.
But the thing is, early on when your app is in flux, neither does writing Selenium code. There's a pretty big truism in UI automation that writing UI tests before a UI freeze is a recipe for a shit-ton of pain. Coding to ids or xpaths only gets you yay far if the UI flow itself fundamentally changes.
But re-recording might be easy.
Don't use stuff like this for long-standing tests. Unless you architect your app just right so that IDs are always stable, and you never change the interface, it'll break and break in ways that you can only fix by a complete re-record. Plus the tests tend to be extremely timing-fragile because recording uses delays instead of synchronization points, so they just won't work in CI.
But do use stuff like this at your desk during bring-up, when the cost of a re-record is lower than the cost of a test rewrite and it's ok to re-run the thing if the first try craps out due to a timing glitch.
And from there, keep an open mind.
I went to a GTAC (Google testing conference) where a presentation made a very good argument--with numbers and everything--that for smaller projects with simple and more or less static UIs, and where the tests were all fundamentally scripted, there was almost no advantage to coding the tests later. Record-replay was the best way to go.
But I definitely don't think a system like PlayWright fully replaces remote stubs like WebDriver and coding tests in the languages that can talk to it.
At some point you hit the issue that the login screen changed and now every single test is invalid. It's awfully nice if you used Page Object Model and you have a single-point-of-truth to fix it.
More to the point, test automation can do more than just execute a script repeatably. Randomized monkey testing is something you can really only do in code, ditto scenarios that need to talk to testing interfaces exposed from the app itself.
Glad you found a tool that resonates with you!
I tend to find in most apps that unless you directly change the front end templates to give elements you interact with better identifiers and then deliberately use those identifiers your browser tests will always end up an unreliable sucky mess.
Dynamically-created UIs are usually the hardest thing to deal with--if they get a different ID every time, record-replay is out the door, and even Selenium is a lot harder.
But personally, with a lot of automation architecture experience, I think you're exaggerating a little re: doesn't work 2 minutes later half the time. But it also depends on whether your app is driving some invisible external resource that makes the timing highly variable, etc. Even then you can usually "harden" the recording by putting in worst-case delays. It's just that then your test takes so long to run you can't CI it either.
It's really situational. What I'm arguing against more than anything is the knee-jerk reaction that it's never appropriate. It's just never enough usually. The fact that we tend to skip over the option entirely in testing is probably a blind spot and a mistake. Devs dabbling in testing as a side task definitely shouldn't ignore the option.
I like fewer sources of truth, but I dislike trying to force everything into a "page" analogy. Try maintaining a library of commonly-used actions instead, like:
logIn()
registerNewUser()
addItemToCart()
And maybe group them thematically.The underlying intent is the same... reduce duplication, improve organization. But you won't be stuck wondering "what page am I on now?" in a single page application, or rifling through folder upon folder of nonsensical "page" files.
And even in the worst case, Ctrl-Shift-H still works.
Use granular element-oriented functions (i.e. loginButton.click() or fillForm(name, pw) type stuff for the very small part of the very few tests that were specifically exercising that portion of the UI.
Those you probably define traditionally for POM, as methods of the page (or functions in the page module depending on the language).
Use result-oriented functions (logIn(), registerNewUser()) whenever it's "travel" i.e. things you do to get to the start of your scenario, or a setup/cleanup task.
Those you do not keep with your pages, they live in modules organized by task or result. Plus they have to work from anywhere. They can leave you somewhere, if that's their defined result, but they should be callable from any UI state. By the same token, tests shouldn't assume how they used the UI to get there, again, unless that was defined as their result.
In other words, they're functions: black boxes. The biggest point there was "you can't change this function without preserving that contract, and you can't assume anything but that contract."
The advantage is that if you could wave a magic wand for setup, travel, or cleanup and get the result, the test would still work and be valid. IOW, you can select the most robust and direct way to accomplish those things, even going completely around the UI with cookie injection or whatever.
What most other teams I've had visibility into tend to do differently than I advised is they'll use the granular element-oriented POM functions everywhere except fixture setup/cleanup. They don't have the "travel" concept of things you have to do on the way to your scenario start, and they include that in the test scenario itself with granular element calls.
And travel is really all setup. But some reason when it's "set up yourself via login, option selection, loading a file, etc" people's thought process goes out the window and they think that all needs to be strung together in UI like a user would do it. But intelligently separating out the very small bit of "specific UI manipulation that causes a state change + verification of that change" that is the test from everything else in the scenario that is setup/travel/cleanup gives you much more maintainable tests.
Or even when they do separate them out, they're not really "result-oriented" functions. Instead they're "flow-oriented" macros that you couldn't replace with a magic wand, because the meat of the test assumes intermediate UI flows they performed rather than just end state, and they're written to be strung together in some coupled (and usually undocumented) way.
Then you have the systems that try to use the same functions for setup/cleanup and testing, caught in between the need for granularity and robustness. Those tend to get extra "doItThisWay" flags on their functions and stuff really goes to hell.
Gotta keep 'em separated!
TL;DR I agree with you, and even a few steps further.
This has the concept of actors, abilities, interactions, questions, and tasks. This allows good separation of concerns as well as much more user-focused tests.
[1] https://serenity-js.org/handbook/design/screenplay-pattern.h...
I hadn't realized there was prior art to look at or potentially clone by mistake. I wonder how old this pattern is.
Thanks for showing me!
createUser
logIn
addCreditCard
addItemToCart
checkoutItem
Was pretty powerful, but ended up being used more as a CD platform for testing and less like automated regression testing.I agree with you. Most people here discuss recording vs not recording. I think most people really lack a basic understanding of testing. A tool is still a tool.
In my point of view most websites aka apps can be treated like that and maybe should. Usually there needs to be a testing strategy in place. Why and what are you testing? What do you wanna confirm, what kind of bugs are you looking for, what is your testing strategy? How much time and effort go into test?
We found out (finance), that one essential ingredient was missing in our tests: the happy path. Without a single test for the end to end happy path, testing for anything else becomes useless.
And the happy path can and maybe should in most cases be tested using a recorder, because it comes closer to what a real person does.
I was/am a pretty big fan of the Context-Driven Testing (https://context-driven-testing.com/) concept from James Bach and Cem Kaner, though I think Kaner later backed away from either it or Bach, not sure which. It was basically the testing version of the Agile Manifesto.
The idea in general of "THINK about what you're doing and the specific results you need and do what's OPTIMAL, not what's DOGMATIC" has guided my career for many years, both inside and outside of testing.
To your point, though, the advantage of the class is that other things in your stack probably don't have opinions about what class you tack on, whereas they might want to define your (actual) id for you. IIRC this was how I asked devs to get around dynamic-UI runtime issues where the ids were also auto-generated.
E2E tests should mimic how users interact with the page and they see no data attributes.
You can query by text, label, html semantics, position.
Hell even for clickable icons you should have alt text.
Nor should unimportant stylistic changes in class, or changes in element order/positioning, break our functional tests. Thus we give important elements data-attributes which are resilient to all forms of refactoring.
HTML has specific semantics, and I find that interacting with applications the same way a blind person would (aria labels, text, etc) leads to much more solid tests.
Copy changes are far from common on stuff like labels and any actionable item you aren't changing "submit" to "send" every other week or "pay" to "checkout" (moreover, such a button would have a meaningful aria role like "search" or "register").
And if you do, fixing the tests is generally very cheap and quick, so I see it as a non-issue.
Nothing will change my idea that abusing data attributes (which are sometimes, but rarely, a necessary evil) leads to the great amount of non accessible and semantically incorrect bloated HTML we see everywhere on the net.
Good tests will lead to better websites, and data-attributes do nothing to help in that direction.
I bet a team like that would be totally transformed by getting some early success under their belts, and giving valuable feedback early enough to actually be included in the whole process.
That's exactly what a QA team can use these tools for, especially early on--it can accelerate manual testing if nothing else, by people record-replaying at their desk for repeatability. Even if that won't work for verification, sometimes just recording the script that does all the setup you have to do for every test gets you really far--then you just take over manually from there.
In general, I think QA's sometimes-bad rap comes from putting in maximal effort and cost, in trade for a very limited and usually unquantifiable increase in actual quality confidence, and that usually comes down to dogma. Most of the "traditional" ways of doing QA come down to writing maximum docs without any single point of truth (and therefore become maintenance hogs and/or wrong) then spending a million bucks or more a year on a team whose whole job is to tell you 99 times out of 100 that things look the same as yesterday.
It makes the sector a thankless grind from the inside, and makes it an expensive spreadsheet generator that sometimes blocks your releases for unclear reasons from the outside. Fun times.
> for smaller projects with simple and more or less static UIs, and where the tests were all fundamentally scripted, there was almost no advantage to coding the tests later.
How many times in my career have I been asked to meaningfully design tests for basic static sites? Hardly ever.
Complex, multi-tenanted, many user-ed hydras with dynamic client side rendered content? All the time.
The scenario in which that Google presentation suggests it's most effective is the rarest ever, in my experience, especially if site is complex enough to have someone working on the team with browser automation coding experience/ability in the first place.
I also think (but am hesitant to assign to them because obviously I built this up in my head over the last ten years too) that part of the argument was that this enabled testing earlier and for smaller projects than people normally bothered testing--for example, internal-facing tooling. Whether or not they said that, I think that is one of the big advantages.
Also, left unsaid, this was for situations where a very small number of critical path UI tests are sufficient for the moment and you're not re-recording a huge suite if something breaks. Of course, if you're familiar with the testing pyramid, you probably know that only having a small handful of critical-path tests running E2E via UI is your ideal, period. Most organizations who do it at all heavily over-test via UI automation.
Your point is well-taken, though. By the time your app is that complex, I'd say you're probably in the position of creating those long-standing tests I wouldn't recommend recording.
Excuse me, but isn't selenium actually a browser robot and not a javascript parser? I assume this is just that DOM is mutating so much that they could or could not find that element or always have to wait for it and comes with unreliability and such.
I haven't used alternatives, but using selenium feels like solid "Engine" to make abstractions on.
Back in Windows-heavyweight app land, it made way more sense to interpose a low overhead shim to intercept and copy events.
I can't imagine the pain you'd encounter trying to paper over the other way around.
They do make some tasks tricky to handle, like tracking state, or avoiding race conditions, but in reality events are a good fit for interacting with browser DOM events, representing browser activity, and being able to react to it.
Consider the complexity of an average modern web site, that loads dozens of scripts, uses a frontend framework, is highly dynamic and interactive (likely a SPA), contains multiple frames, etc. All of this happens concurrently, and the only abstraction that makes sense to represent it is an event-based system. You can't achieve the same level of responsiveness using a traditional request/response protocol.
This is why modern browser automation protocols are event-based, and built on top of WebSocket. Selenium is pushing for WebDriver BiDi, while the browser standard is the Chrome DevTools Protocol.
Fully agree then. The only sane way to handle browser automation is by exposing all of them.
If it can only drive its internal browser engine, that'd possibly be a big argument against it for any test strategy I was architecting. I'd want the tests to be portable and usable cross-browser for acceptance--particularly UI-driven E2E tests, since that's really the only time they are critical.
The era of custom headless browsers like PhantomJS, or Selenium/WebDriver for that matter, is dead, IMO. The Chrome DevTools Protocol is the modern standard API to interact with a browser programmatically.
This does sound more promising then, if they've found new ways to add maintainability or QoL aspects to record/replay. As I said elsewhere, I think r/r gets a bad rap, and those sorts of improvements would increase the number of situations where it makes sense (at least for testing in incubation).
I do recommend giving it a try. The web browser automation landscape is much more reliable these days for writing E2E tests than back when Selenium was state of the art.
I think you're wrong there. See [1].
It doesn't use custom browser builds, but specific versions of official builds of mainstream unbranded browsers (you can also use branded ones if you wish). There are no "custom headless builds". All browsers support headless mode, and you can use it on your main browser by passing a CLI flag.
The reason it officially supports only specific versions is because there may be slight incompatibilities with the Chrome DevTools Protocol in each browser version, so they suggest running supported versions only. But like a sibling comment mentioned, you can also connect it to any existing browser instance, and, assuming there are no CDP incompatibilities, it will likely work.
> Playwright "records" your steps and even gives you a running Python script you can just directly use (or extract parts from). I'm serious about the running part- I can literally record a few steps and I will have a script I can then run with zero changes.
Does Selenium have that ?
However, regardless of the tech used, recording playback of HTML manipulation is pretty brittle - if you think adopting playwright because it has a record feature will save any kind of time/effort/money, you are in for a surprise. This is a feature everyone tries once and goes "wow, neat!" and then quickly realizes none of the suggested code is production ready, or even likely to work on a 2nd run for anything other than the most basic of static sites.
Also Playwright is a general browser automation tool and can be used for more things than just testing. For example, scraping website data or generating preview thumbnails for URLs.
You can open multiple tabs within the same browser window, or you can create multiple WebDriver instances for multiple windows.
From what I can tell here are the benefits:
* With Playwright you don't need to put manual waits in your program. It's async.
* You also don't need to keep updating the Selenium.Webdriver to stay in sync with your browser version. This is a 15-20 minute time sink each time it happens.
* Can run in headless mode which will run the test suite much faster than selenium.
* Also appears to have better command line options to be run from a CI/CD environment
That is just choosing between a blocking method or... a blocking method which suspends and resumes a thread.
> headless mode
That is really up to the browser driver to implement that, and you can do that perfectly fine in Selenium.
The webdriver issue is definitely annoying though. Not sure if there's a better way, but we're using a binary regex on the electron app in question to find a matching version and download it on the fly as part of test setup.
I'm pretty sure they have equivalent (or very close to that) functionality, in that anything you can do in Playwright can be expressed in Selenium and vice-versa.
But Playwright, from my experience, is much easier to use. The documentation is more comprehensive, and the API design is much more aligned with the kind of things I want to get done.
This is because Playwright has benefited from decades of experience in the space which started with Selenium. The engineers behind Playwright previously built Puppeteer, which was itself an evolution of ideas from Selenium.
Selenium sucks
I tend to see that negatively. Community driven projects (or non profit organisation behind them) like Python, Postgres, Firefox always are more trustworthy and likely to last longer in general.
My favorite example: VSCode is a great piece of software, but almost half of it is actually only available in a proprietary build from Microsoft's website that contains telemetry.
Which half?
VSCodium works great, so I'd go as far to say 99% of VSCode is actually open source (the name and icons being exceptions). You cannot say proprietary extensions are half of VSCode.
Not necessarly your case, but surely for most complainers.
For example, if Microsoft wanted to change CPython to automatically collect sloppily anonymized user data and send it to Microsoft without asking, as they did with .NET[3], I would expect the steering council reject that proposal.
[1] https://www.python.org/psf/
Playwright is Open Source - the worst thing that will happen if Microsoft will kill it - you’ll be able to use the latest release until you find some replacement.
GitHub was bought by Microsoft in 2018.
Recently Microsoft announced they were killing off GitHub Atom.
Being able to use the latest release of a dead project is no big reassurance.
This might be an anecdotal case, but it is indeed real world proof that a project being backed by a company the likes of Microsoft does increase the odds of being killed off.
A shame to see Cypress lumped in and dismissed like that. It really is a fantastic way to test.
> The killer feature of Playwright is: You can automatically generate tests by opening a web browser and manually running through the steps you want. It saves the hassle I faced with Selenium, where you were opening developer tools and finding the Xpath or similar. Yuck
This is absolutely not the primary hassle with testing. Recording steps can help kick start your testing, sure. But very quickly you start to realize that it's only saving you from the easiest work at best, and is creating flaky tests at worst.
For massive numbers of tests, now it's time to look into parallelism. Really hope you're not sharing your test data between test cases! Or holding open a single db session to your write instance!
How about multimedia artifacts management?
Point being, test automators do a shit ton of work nobody ever wants to dig into.
Tests are not there to only be written. They must be read.
I think Cypress is really cool when you're on their happy path - testing on Chromium browsers and in JavaScript.
I tried it some time ago so maybe this might be outdated, but they had little support for browsers other than Chromium or even rather Electron based and also it was a little difficult to wire up with Gherkin scenarios.
Keep in mind that the OP refers to Python tools, AFAIK Cypress does not offer Python bindings while both Selenium and Playwright do.
Cypress definitely has some limitations, like any other tool, such as only supporting JS-based test. Playwright is definitely a great tool, too. I like them both for different reasons :)
[0] https://www.cypress.io/blog/2022/09/13/cypress-10-8-experime...
Asking mainly because there’s a tool I have use without an API and if I could script something like that it would make my life a lot easier. Just never tried that type of thing before.
Sometimes you can even do these with pure curl. POST the login form, get back the necessary cookie or token, then request the download URL, if it's easily predictable.
Most testing tools make this kind of thing pretty easy. Cypress sits in the page's JavaScript and has access to a Node back end, so when you're not clicking on stuff you can be firing off requests or doing basically anything.
What you might find is that you don't need any fancy stuff though. Maybe look at the download request in Chrome Dev Tools and see if maybe you can just execute a POST command in your language of choice?
See that another way: You, as a human, can’t scale for a team of 50. e2e tests can.
This is absolutely true. Who cares if the JSON payload is correct if the page is wrong?
This is how I felt before I started using Cypress. You need a tool that helps you rather than being a drag on you, and you need to commit some time to it. It's helped me learn a lot of coding I didn't otherwise know. Writing the test itself is usually not "real coding" but writing helper methods and learning how to organize your code certainly is. The trick to making it all work is making your tests not flaky. They need to pass 99% of the time. Now we have a few hundred long end-to-end tests and maybe 2 of them are flaky at any given time. They are easy to fix when they break, and they uncover important bugs on a regular basis.
Granted, I'm in the business of testing. Anything that isn't automated I have to do manually. And I have to represent a real user, so simple API tests are not going to cut it, and I have an incentive to automate the entire end-to-end process with multiple variations. But if you want to do short API tests you can certainly do that, and Cypress is actually a nice tool for that too.
This was the first time I was happy writing ui tests, Playwright is great!
(Though I don't use the recording tool at all this article focuses on I rather write the tests manually)
That's because automatic code generators rarely work (anything more dynamically generated and it's a no-go).
If that's not it keep looking. If there's a thing you want to do, odds are Gleb has a video or blog post about it.
Anyway its subjective, I don't like Cypress.
I figure there must be a way to do that, but it's very possible I'm misunderstanding what you're trying to solve!
And Cypress gives a great API surface where almost everything can be done with just this declarative API, and it enables them to show the full plan in a UI even while it's executing.
But if you actually need custom control flow, with code you control... you need to "pause" your action-enqueueing to let Cypress catch up. Which means you're going to be nesting callbacks, or sprinkling awaits in all the places that aren't logically awaiting, because you need an explicit yield point out of your JS block. Our team finds this rare enough that it's acceptable. But... yeah, it's far from a perfect solution.
Both of these issues are well documented in git issues with hundreds of replies and work arounds that just don’t always work.
I’m ready to throw in the towel and rewrite everything in playwright. Maybe b I’ll have different issues there…
I have only seen this issue with JS based VDOM frameworks.
I really really like the cy.intercept() for intercepting and modifying requests.
- The Test Runner lets you retrace your actions with a visual UI and interact with the page as you rewind, with access to Dev Tools the whole time. It also has an excellent tool for targeting elements. (Although you'll eventually lean more on Dev Tools once you're fluent with CSS selectors.)
- There's very little boilerplate for anything. Several useful libraries are packed in and available globally.
- You can easily spy on and interact with fetch and XHR on the page
- You can easily execute Node stuff
- They have a great Dashboard (their only pay feature)
My learning from playing CDP and the test automation applications mentioned is that this is way harder (and slower) than test automation should be and it’s a huge pain in the ass for authoring tests. The experience is so bad that it only seems to appeal to professional testers whose jobs depend upon it.
To solve for that I wrote my own test automation solution that does not require remote access to the browser. In my personal application where I use this form of testing I am able to execute about 30 user events per second in the browser. That performance speed is completely dependent upon other performance conditions of your application, hardware, and transmission handling. The test cases are simple data structures defined as TypeScript interfaces in a Node application communicating to the browser over WebSockets. The events are executed by creating a new Event object in the browser and executed with the event’s dispatchEvent property.
When your automation gets really fast (less than 30 seconds for a major test suite) everybody will make use of it including the developers, product owners, and even the business leaders. It becomes a faster way to setup and access various features of a browser application than doing it manually.
Any more details? Are you still working with actual chrome?
You can’t know what’s faster unless you are measuring things in isolation and making incremental improvements to a bunch of different bottlenecks. For example people love to tell me how fast their framework, their application, or whatever is but it’s clear that these are almost always anecdotal observations that are not measured or compared to anything.
When looking at performance it does not matter how fast something is, which makes anecdotal observations worthless. The only thing that matters is the difference in speed between things compared, the performance gap.
I recommend running the site from an aliased subdomain. This will allow ownership of https certificates with a wildcard to your primary domain and that subdomain can point to the environment that contains your site database or services. You can also have a subdomain that uses production https certifications but resolves to a loopback IP like https://localhost.example.com pointing to 127.0.0.1 and/or a AAAA record pointing to ::1.
content = driver.find_element(By.CSS_SELECTOR, "p.content")
In Playwright Python: content = page.locator("p.content")
Playwright is just nicer to use. Selenium feels very Java-ish.You may consider trying this extension - https://chrome.google.com/webstore/detail/devraven-recorder/.... This extension can record the tests using your Chrome browser instead of launching a separate Playwright chromium browser. Full disclosure: I am the developer of this extension.
My company uses Selenium for some of our projects, but I evaluated the browser automation space some time ago and found that it was the best-in-class solution at that time.
It looks like things have since changed.
Best features for me:
- Auto-waiting - it is a BRILLIANT feature!
- Shared authentication;
- Supported browsers list;
- API testing;
- Tests generator (with recognition of data-test-id attributes);
- Flaky tests detection;
- Flexible config for recording failed tests;
- Ease of installation and updates;
- Great performance.
Thank you, Playwright team!
Eventually you'll have to write your own scripts. I will admit Playwright can spit out code that will at least serve as a starting point.
It’s a fantastic and versatile tool. I couldn’t agree more.
You can combine traditional RPA with modern ML and do very interesting things.
What all of these tools actually need is to be bundled with some AI-powered recognition tools - computer vision, semantic understanding, etc ... - that can sort of "understand" the web page served before interacting with it.
All people do right now is a combo of XPATH-based search for an element in the DOM, or if that fails some sort of regex based pattern matching on the HTML.
These are weak and brittle techniques that are pretty much guaranteed to fail at some point.
For example I keep seeing auto-wait being talked about for Playwright but doesn't cypress have that too? It seems playwright has better browser support and multi-window/concurrency story and cypress have better built-in tools (like a test-runner UI).
Basically looking for recommendation for test-automation for an old complex app mostly written using server-side tech but has potential for react and such moving forward.
describe(('the google home page') => {
it(('allows you to search for banana pictures') => {
cy.visit('https://google.com')
cy.get('[aria-label="Search"]').type('banana{enter}')
cy.contains('Images').click()
cy.get('[alt="Image result for banana"]').should('be.visible')
})
})
Nothing too weird, and a lot of the library functions you'd need are available globally without any importing. It even has jQuery and Lodash.Playwright is faster.
It's fully free and OS.
It's written by Microsoft to test their own thousands of services not to sell you anything like Cypress.
The authors of Playwright are the leads from Puppeteeer, Lighthouse, the first node debugger and Chrome Dev Tools, the engineering team has simply no comparisons in the browser automation field.
[0] https://github.com/chromedp/chromedp [1] https://github.com/chromedp/docker-headless-shell
There are no tests as valuable as E2E tests.