HNHacker News
TopNewBestAskShowJobs

thalissonvs

43 karma · joined March 12, 2025

submissionscomments
thalissonvs··on Beyond User-Agent: A Guide to TLS, HTTP/2, Canvas, and Behavioral Fingerprinting
Author here. I've been working on pydoll, an open-source (Python/async) web automation library. While building it, I kept hitting a wall against sophisticated anti-bot systems.

This sent me down a deep rabbit hole to understand how they actually work. It turns out detection isn't about one thing, but about consistency across multiple layers: from the OS-level (TCP/IP, TLS/JA3), to the browser (HTTP/2, Canvas/WebGL), and finally to human behavior (mouse physics, typing cadence).

I decided to write down everything I learned in this guide. It covers the theory of how each layer is fingerprinted and the practical techniques to evade it (focusing on consistency, not randomness).

Hope you find it useful. Happy to answer any questions.

thalissonvs··on Show HN: PyDoll – Async Python scraping engine with native CAPTCHA bypass
I'm not actually bypassing the captcha with reverse engineering or anything like that, much less integrating with external services. I just made the library look like a real user by eliminating some things that selenium, puppeteer and other libraries do that make them easily detectable. You can still do different types of blocking, such as blocking based on IP address, rate limiting, or even using a captcha that requires a challenge, such as recaptchav2
thalissonvs··on Show HN: PyDoll – Async Python scraping engine with native CAPTCHA bypass
CDP itself is not detectable. It turns out that other libraries like puppeteer and playwright often leave obvious traces, like create contexts with common prefixes, defining attributes in the navigator property.

I did a clean implementation on top of the CDP, without many signals for tracking. I added realistic interactions, among other measures.

thalissonvs··on Show HN: PyDoll – Async Python scraping engine with native CAPTCHA bypass
you can check the official documentation, there's a section 'Deep Dive'
thalissonvs··on Show HN: PyDoll – Async Python scraping engine with native CAPTCHA bypass
cool, left a star :)
thalissonvs··on Show HN: PyDoll – Async Python scraping engine with native CAPTCHA bypass
I don't think it's similar. The library has many other features that Selenium doesn't have. It has few dependencies, which makes installation faster, allows scraping multiple tabs simultaneously because it’s async, and has a much simpler syntax and element searching, without all the verbosity of Selenium. Even for cases that don’t involve captchas, I still believe it’s definitely worth using.
thalissonvs··on Show HN: PyDoll – Async Python scraping engine with native CAPTCHA bypass
Well, it really depends on the user; there are many cases where this can be useful. Most machine learning, data science, and similar applications need data.
thalissonvs··on Pydoll: Async Web Automation in Python
Pydoll is an innovative Python library that's redefining Chromium browser automation! Unlike other solutions, Pydoll completely eliminates the need for webdrivers, providing a much more fluid and reliable automation experience. Zero Webdrivers! Say goodbye to webdriver compatibility and configuration headaches Native Captcha Bypass! Naturally passes through Cloudflare Turnstile and reCAPTCHA v3 * Performance thanks to native asynchronous programming Realistic Interactions that simulate human behavior Advanced Event System for complex and reactive automations