Deepest level possible, harder but required for some workflows
35 karma · joined October 13, 2023
Deepest level possible, harder but required for some workflows
But stealth, as we see it, isn’t about deception for abuse. It’s about making automated access behave closer to how real users interact with the web, in an ecosystem where most anti-bot systems default to blocking everything that isn’t explicitly whitelisted.
Right now, the model is broken. Unless you have a direct partnership, you're often locked out, even for legitimate use cases like research, monitoring, or building user-facing tools on top of public data.
We’re not supporting harmful behavior (credential stuffing, DDoS, piracy, etc.). The goal is to enable responsible access to publicly available information without forcing every use case into closed-door agreements.
There’s also a real tradeoff happening. Increasingly aggressive anti-bot measures (like harder CAPTCHAs) degrade the experience for actual users, while not necessarily stopping sophisticated automation, robots solve CAPTCHAs better than humans.
So the question isn’t “bots vs no bots” — it’s what kinds of automated access should exist, and under what norms. Right now, that line is blurry, and we think there’s room for better balance.
Happy to engage on where that line should be drawn.
In our case, we prepare the environment, load files that we need later and then we create the state. Once we start, we instantly start Chromium with the config requested by the customer.
Keeping the browser open and warm is also a problem, not all customers require the same features. The same engineering required to fix that (modifying values with Chromium open), also fixes the post-chromium snapshot
VM takes 20ms to start, browser around 300ms. Post-Chromium snapshot is at 50ms end-to-end, defeating the benefits of the warm pool you suggest, that will be our next step.
But yeah, in one server we can fit hundreds of browsers, or even thousands if we use bigger servers. And each one of them with dozens of tabs, no issue
So there's no benefit on reusing the VM but not the browser. VM isolation is also important, customers can leave downloads and other files that should not be accessible for freshly created browsers on that same VM.
Startups are absurdly slow, isolation is harder, etc...
Android bloat is insane, you need to run the entire Java VM to start the browser... It's also harder to fingerprint, and at scale that's something that we need for Browser Use
Cool experiment but not yet production ready, at least for us
Warm pools are nice but at the end they also consume resources, And you need to always keep the pool warm, starting browsers to balance, etc...
With the upcoming changes we will keep Chromium startup and the VM will be ready in 50ms, defeating warm pools at all
Also some customers need special parameters and features, increasing warm pools complexity. The happy path will be fast but the edge case will be extremely slow , and we want to guarantee fast speeds to matter which features you need on the requested browser.
Few issues we had with lambas: - Limited running time (15 min), we support up to 4 hours (we can run longer if needed) - Price - Lack of snapshotting mechanisms - Lack of low-level control over the running host
But yeah, lambda is way more than enough for most common use cases automating the web
Main blockers right now is fingerprint injection and profile injection, solved already.
It's always a balance of engineering effort & gains. Post-Chromium snapshot let's us save 200ms, which is not that important for 99% of use-cases, but that will come soon since it brings some other benefits (like CPU footprint)
Profiling and tools used are already included with Chromium, they provide nice debugging tools
Browsers like LightPanda lack stealth at all, they are trivial to detect. There are ways to make Chromium more performant, by removing everything that you don't need.
We believe that Chromium can reach that performance without starting an entire engine from scratch, and without losing stealth, a top priority for us.
The language is not the problem, C++ is as performant as Zig, but Chromium bloat is huge, agree on that.
That's why we moved to a fully in-house solution with Firecracker and auto-scaling on EC2
That's not true. Bots can still automate the web and there's demand for products that allow it. It's harder than years ago, but not impossible.
Defenders are always in favor, but the demand for automating the web exists, so research keeps going. There are ways to hide everything, including residential proxies.
For reference, I'm the blog author, and I have another one talking about this topic: https://browser-use.com/posts/bot-detection
We support GPU via software tho
Our focus is on staying ahead by eliminating signals that antibots aren't even checking yet, that's where the real research challenge is.
Benchmarks comparing competitors across high-security sites are coming soon, thanks for reading!