Theheadless.dev – open source Puppeteer and Playwright knowledge base
theheadless.dev
theheadless.dev
Works flawlessly, and has exactly the configuration knobs you would expect and want. Took a bit of plumbing to call into a Node.js script from my Scala build logic, but all in all ended up being like 20 lines of plumbing which was straightforward to write and understand. A welcome change after struggling with bugs in wkhtmltopdf!
I'd just started looking into the existing tools like those you mentioned for e-book authoring and it's quickly gotten overwhelming. From what you've described, I'm interested in looking into puppeteer as an alternative workflow.
One of my favorite tricks I've seen employed are detection measures that look to see if common detection bypass tricks have been implemented (like checking the toString output of commonly overridden native functions.)
https://theheadless.dev/posts/challenging-flows/#bot-detecti...
If you're good at spoofing all of that fingerprinting you'll blow straight past them - it's all client-side in-which you have control all the way down to the bits and bytes.
0: https://support.cloudflare.com/hc/en-us/articles/36002751945...
Something worth noting about toString is that it can now be undetectably modified (to fake “native code”) with the new ES6 Proxy object. There was a really interesting blog post written about this at https://adtechmadness.wordpress.com/2019/03/23/javascript-ta... (I also incorporated this into my project).
edit: Really like that repo! I use a lot of those techniques as well.
https://www.npmjs.com/package/puppeteer-extra-plugin-stealth
I've created a couple scripts to delete accounts, and even signing in can include randomization between scrolls and clicks to make each change entirely unique and to mimic real user interactions. Sort of scary to think about what is possible with this, especially given a large pool of residential IP addresses.
I believe with enough money and the right skills it has always been possible to manipulate the media - social or otherwise.
The question is if puppeteer means we can decrease the amount of money significantly for a minor increase in the amount of skills.
It was probably great for his engagement / viewership numbers till we shut him down.
E.g. if they were paying for the 1000 concurrent puppeteer sessions, would everything be in the clear with your SaaS?
Presumably, your service doesn't care what the users use it for. Sure, though, it's a violation of twitch's terms of service to fake viewers on the platform. I may be naive -- could twitch sue a service that is used to fake viewers?
https://www.checklyhq.com/docs/browser-checks/responsible-us...
Luckily, most users are completely ok. But fraud and abuse will always be a thing.
Beside risks like Twitch (or worse, all of Amazon) blocking your traffic wholesale, letting someone use your product to abuse another is just a crappy thing to do; Especially if you let it because "well we're making money".
Not even saying about smart paid trolls, who are good at using eristics.
Seeing what people put in their bios, what their network looks like, their posts and all the metadata, its incredibly what amount of user data is pretty much out there in public. Seeing that this is just the tip of the iceberg in terms of what companies like Facebook are collecting, all those stories about algorithms predicting pregnancies before the people themselves know about it seem much more realistic.
The main think holding people back from mass creating accounts to spam social media (even more) are chaptas, and those can be outsourced to chapta farms somewhere in the world.
It would make writing end-2-end integration tests much easier...
Not great. They're prone to timing problems and fragile selectors. They work well for rapidly recorded, rapidly changing sites (where you can throw them away) or very stable sites (where the DOM is basically fixed), but they rapidly become a hassle in the middle.
But they sure are easy. I'd suggest recording as a first pass, then going back and refactoring them to be a bit more stable; Change XPATH or structural CSS selectors to use classes and IDs, add waits and assert for page loads, etc.
I'm keeping the codebase open on GitHub: https://github.com/umaar/learn-browser-testing/ so anyone who wants to follow along can do so for free.
I've almost finished some cool content such as:
- An Amazon price checker which sends you a text message when the price decreases
- An Playwright script which gets to Wikipedia Philosophy (https://en.wikipedia.org/wiki/Wikipedia:Getting_to_Philosoph...)
- Having an automation script constantly running on a cheap Raspberry PI
But it would STILL suck, because the input is so dodgy.
Playwright interests me even more though. We've been getting a lot of requests for ArchiveBox.io to support other browsers as the rendering engine for web archives, and it's always seemed daunting to try and reimplement multi-browser support ourselves for puppeteer-style workflows, but Playwright seems to completely take care of that!
- https://checklyhq.com - synthetic monitoring
- https://microlink.io - automation
- https://www.browserless.io - testing
- https://www.scrapingbee.com/ - scraping
All are business leveraging these types of frameworks
My app is powered by Selenium but the concept is the same.
Its purpose in our case is front-end performance measurement. In short, have a script running checks in a Cron Job, published to static reports we can reference over a timeline.
Nothing overly sophisticated, but it suits our current budget ($0).
- generally higher speed
- higher reliability (specifically lower base false-positive rate)
These are mainly due to architectural choices (less moving parts between script and browser).
That being said, Selenium has been the open-source standard in cross-browser testing for a long time now, and is more polished and feature-rich. Also, multi-language support makes it an easier choice for non-JS teams. I would suggest a quick hands-on POC if you want to use these tools in a project.
Edit: I was kind of wrong already and it seems I will be completely wrong soon. Which is good in this case :-) As hlenke points out below Firefox support is on its way :-)
* The Playwright API auto-waits for the right conditions on every action on the page (click, fill). This ensures automation scripts are concise to write and maintain over time.[1]
* Unlike Selenium, Playwright uses an bi-directional channel between the browser and automation script. This channel is used to listen to events from the browser (like page "load" event, network requests). These events enable Playwright scripts to be precise about browser state and prevent the need to rely on sleeps/timeouts, which contribute to flakiness of Selenium scripts. This is also exposed in the API, for more powerful automation[2].
* Playwright also has a wider coverage for modern browser features, including device emulation, web workers, shadow DOM, geolocation, and permissions.
[1] https://playwright.dev/#version=v1.3.0&path=docs%2Factionabi... [2] https://playwright.dev/#version=v1.3.0&path=docs%2Fapi.md&q=...
Selenium uses a chatty HTTP interface, whereas puppeteer/playwright use WebSockets or pipes to communicate. Under-the-hood, however, Selenium is simply using chrome's devtools protocol to communicate with it. The way selenium does this is by another binary, generally a `driver`, that has the protocol "baked" into it and has the HTTP selenium API as its input interface.
This is all a long way of saying that puppeteer/playwright have a lot less moving parts, and are generally more approachable. Selenium _does_ have a lot more history behind it, better support across languages and frameworks, and is more stable but it's also much larger and "clunkier" feeling. It's also a lot harder to scale with load-balancers since, again, it's all over HTTP so you'll need some way to load-balance with sticky sessions.
Practically speaking they all do the same thing at some layer. Both are high-level APIs around the devtools protocol, it's just what higher-level interface you prefer and what your language/runtime is.
The newer automation tools benefit from being newer; They can take advantage of hardened, well designed interfaces (like the Dev Tool protocol). Selenium's been around for a bit longer, and was built when browsers didn't make it easy to control them. That's influenced the semantics of Selenium quite a lot, as well as explaining the extra moving parts (Drivers exist to map the Selenium Wire Protocol (or W3C protocol) to whatever they're driving because Selenium wasn't built with a specific browser in mind).
I feel like, at this point in time, the real difference is how much abstraction you want from the browser. Selenium is a set of knives, Puppeteer is a Die Cutter. You'll put in more work with Selenium, but maybe you need something do happen a REALLY specific way. Or, you might just need shapes cut, and Puppeteer will be more reliable and faster.
I've tried using the various networkidle events, wait for some DOM element, and I find myself just using 5 seconds or something like that as the most reliable solution.
Is there a foolproof way of doing this? I feel like it should be way easier and less hacky.
You can also look for a specific element.
Humans don't know either, they're just better at guessing based on visual cues.
One could design a page that visually loaded, then jumped to a redirect after 10 seconds on the page. But who would?
The primary approach is always event-based, because most pages do that sanely.
If not... the best approach I've found is looking for sentinel elements.
Essentially, something that only matches once the website is de facto loaded (regardless of events). Sometimes it's a "search results found" bit of text, sometimes a first element. But more or less, "How do I (a human) know when the page is ready?"
Affiliate networks (the shadier they are, the more likely they will), because they are weird and load third party tracking beacons in transitional pages and want to make extra sure that the beacons (who also can redirect multiple times) have been loaded. To add to the fun, they're also adding random new tracking domains (to avoid being blocked, I assume), so you can't even say whether you expect some domain to be transitional or final to increase your confidence in what you measure.
You're right though, looking for elements is a pretty good way if you know the page you're checking. If you're going in blind, you can still look for things they probably have (e.g. <nav>, <header>, <section> etc), but I haven't found any that are reliably on a "real" page and reliably not on a redirect page.
Most of my work is making known transitions (e.g. page1 to page2) work reliably, so I have the benefit of knowing the landing page structure.
If you're crawling pathological, client-side redirect chains, maybe do pattern-matching scans on loaded code for the full set of redirect methods? There's only so many, and includes / doesn't-include seems a fair way to bucket pages.
A simple self.location.href = ...* is still doable (-ish, because I've seen conditional changes that were essentially if(false)... to disable a redirect, which we obviously didn't consider when pattern matching), but once they include e.g. jQuery (and some do on a simple redirect page) it got far too complicated.
Loaded is the usually the DOM Complete or window.onload event. Fully Loaded is an algorithmic metric that is looking for a time window after the load event when there isn’t any network activity (usually 1.5 to 2 seconds depending on the tool)
Best advice is to just follow what the open source tools do. Same from other computed metrics like TBT or SpeedIndex. (I’m the CTO of a company in this space and that’s what we do)
So when using a crawler I make some assumptions.
I wait for either more than 10 links or 20 seconds, the few pages that have 10 links or less get a longer wait.
I also do stuff like check the links once I have gone into page processing, if a page was 20 seconds and has no links to any other page on the domain, probably something wrong, goes into problem checking queue. etc. etc.
It has to be hacky because there is no way for the browser to tell you with any surety - everything rendered, no events in event queue that will effect display etc.
When you say "wait for some DOM element", not sure what that refers to exactly, but how about using: `page.waitForSelector` and in addition, using the `visible: true` option?
It really depends on the page you're testing, feels like when I'm automating Wikipedia, I very rarely have to wait for stuff, whereas with JavaScript heavy sites, I use a custom wait function, or the waitForSelector with visible: true.
The default waitUntil fires when the onload event is triggered, networkidle2 waits for less than 2 network requests in the past 500ms, and networkidle0 waits for no requests in the past 500ms. This + waitForSelector should help you out.
For a foolproof way, you could also just wait until the specific DOM node you want exists. Either by:
- Polling the DOM until it exists via; setTimeout/setInterval
- Using MutationObservers to check the DOM on every node change, until it's been added to the DOM
When on example.com , a request is sent to definitely-not-tracking.example.com . But it just going to tracking.ad-server.com . uBO blocks those on Firefox
https://medium.com/nextdns/cname-cloaking-the-dangerous-disg...
What is surprising is the linked site is quoting the FAQ of playwright that has been removed a couple of months ago: https://github.com/microsoft/playwright/pull/1930/files
That FAQ seems to be giving a lot of perspective on the subject. Wondered why it was removed.
Haven't worked extensively with them, but it seems to me playwright is the no-brainer now because of the guys behind it (the creators of puppeteer), their experience and lessons learned with puppeteer and also the support for Firefox and webkit.