The /unblock API from Browserless: dodging bot detection as a service
browserless.io
browserless.io
On the other hand, there are a lot of people right now that want to keep stuff accessible to humans and not have it scraped for models. I know the lid's completely off the box on that front so it's probably useless to fight it, but seeing products explicitly designed to circumvent bot prevention feels kinda bad.
Source: Founder@browserless.io
A site I worked on changed our captcha, and in the 24 hours after, we had a 90% reduction in new accounts. We had a full API including webhooks, the only private data was on the account page, there was no reason to scrape, but even if you did, there was no reason to be logged in... these accounts were used to harass, intimidate and scam legitimate users.
Businesses use data from the user. The Business does additional crunching on that data to derive new interesting data for the user. Who owns the data? The user or the app?
At the very least the user partially owns the data and as such, I'd argue that the user should have the right to share the data between different applications however they see fit. However, businesses tend to think that they somehow have the legal (moral even?) right to keep that data in their walled gardens. For as long as this (imo unfair) stance is common, I think that data extraction by use of these anti-bot-bypassing technologies is fair game.
And most scraping isn't done by users. It's done by companies. For profit. Often for less than enlightened reasons.
LinkedIn is a good example: I want my data displayed to people on that job site. I don't want it harvested by every recruiter under the sun who will then spam me. I certainly don't want that data sold between those recruiters long after I deleted my account on LinkedIn. Tinder and sites like that are also an obvious example: yes it's (semi-)public, but I also wouldn't want it to be scraped and harvested by some company – I just want it to be shown temporarily to a limited set of people.
And, in general, I take the fact that you published something on the Internet as a tacit moral consentment for the rest of the world to use it how they want.
This comes with a couple of big asterisks, because (1) Copyright law exists, and I generally try to not break the law, even if I don't agree with it. But the discussion in this thread is mostly separate from copyright: for instance, I don't think a court would see someone scraping and redistributing data from someone's LinkedIn profile as a copyright infringement case.
And (2) because I think that in some specific cases, using published data can be morally wrong, but not as a general rule.
i feel the 'antibot' stuff is more related to the adtech industry vs site-scrapers - remember getting a dedicated server and having friends click on links just to pay for it? Geocities and all these free websites, the biggest costs were bandwidth and storage (not that its not now)
since the AI Boom, there's just more hype over people wanting 'credit' (or money) for something they posted on a forum X-units of time ago.
its called the World Wide Web for a reason.. keep it open, even if it is to 'a bot' - never know when somebody's 'bot software' is reading your webpage for somebody who has some disadvantage and needs assistance
The best you can do is try to form your own little community in a corner of the internet - which is actually a great idea. Mass social media has so many problems - trying to make a little nook for yourself and a small group of enthusiasts/friends seems like the best way to go.
Quite a few things would be vastly better when the information that is on websites is not just publicly readable but publicised in the full sense of the word. Several industries are paying 20~40% revenue just to be findable.
Now maybe scrapers aren't the best solution, but making stuff less accessible is usually not doing much good. Especially not for accessibility.
I suspect that over the next 100 years, as AI and biotech come together, the boundaries between human and bot will start to blur. At some point it will be unethical to differentiate between human, bot, and everything in-between.
I argue that that time is now, if at the very least for the sake of AI-based assistive technologies to be developed to turn a disabled {100% human} into a highly functional {80% human + 20% bot}.
Bot detection should, for this reason, be doomed as an ADA violation.
What if all users of the www were treated equally. There is in practice a conflict of interest between www users calling themselves "developers" and the rest of www users. Yet these "developers" purport to act 100% on behalf of the rest of www users, not themselves. The failure to acknowledge any self-interest is dishonest.
If you really really really want a technical mitigation in addition to a legal one, then paywalls are vastly better than incredibly messy and problematic heuristic bot-detection. There's a huge payment problem for content anyway, so paywalls would mitigate or solve many other problems in addition to the AI-training-data one.
An approach I recently saw gated the phishing page behind a Google Account login page. User that was logged in wouldn't even notice the brief redirect. Scanners would just get stuck. I hope they've patched it by now though...
If smaller indie hackers are going to build useful competitors, they’re going to need to scrape.
I understand that a lot of scraping is from spammers, but not all.
Any useful competitor will have to scale well beyond the point of indie hacker, where someone is (eventually) getting rich. That doesn't make for a great argument for bypassing consent.
Edit: As mmcclure points out Kagi actually doesn't index much. So that's a bad example that might invalidate my point.
> Our data includes anonymized API calls to traditional search indexes like Google, Yandex, Mojeek and Brave, specialized search engines like Marginalia, and sources of vertical information like Apple, Wikipedia, Open Meteo, Yelp, TripAdvisor and other APIs. Typically every search query on Kagi will call a number of different sources at the same time, all with the purpose of bringing the best possible search results to the user.
source: https://help.kagi.com/kagi/search-details/search-sources.htm...
both OpenAI and Google rely on scraping as their foundation for…everything.
OpenAI and Google also (claim to) respect the most basic bot deterrent, robots.txt, with their crawlers. If someone doesn't want their content scraped to be included by search indexes or LLMs/models, then that certainly includes "smaller indie hackers."What we really need is a browser that absolutely cannot access sites "protected" by Cloudflare, and this browser needs to become standard. Sites that put themselves behind Cloudflare need to pay a price for their hostility, and that price is the loss of our business.
If I want to program a bot to get around the horrendous lossage of the terrible UI and multi-megabyte bloat you have created by farming out your web server programming to the lowest bidder, I have the right to do so. In turn my bot will respect robots.txt and will throttle itself appropriately.
This is the deal you signed up for when you built a web server. If you don't like it, GTF off the web. And take your attestation and WAF with you.
I agree with another comment that called this "Abuse as a Service". It seems to me this product's design goal is nothing more than to circumvent measures site owners take to prevent abuse of their site and run a sustainable business.
I've used their product many times actually, and I'm shocked on Hacker News of all places no one's thinking of anything besides abuse. How often is it useful to get information from a webpage and apply it in a new context? Then think of how often said webpage is behind a Cloudflare bot detector.
They are completely in the right to block you though, you're not the owner of that data, you might be breaking their TOS.
In Europe, if the company is actually following the law, in theory yes.
> They are completely in the right to block you though, you're not the owner of that data, you might be breaking their TOS.
IANAL, but AIUI that's definitely not true in the United States and I suspect similar ideas hold elsewhere: https://en.wikipedia.org/wiki/HiQ_Labs_v._LinkedIn
- GDPR doesn't require it be a convenient export. Users want to paste a link on my site, a click a button, and have it magically appear. Not fill out a form, dump their entire account and sift through that.
- I never opined on the validity of blocking bots
- I never opined on if it's breaking their TOS
Abuse implies a harmfulness. Giving users a quick import option from already public data isn't harmful.
They’re not necessarily in the right to block you, if you’re the data subject or acting on their behalf.
Like if I want to programmatically unsubscribe from a subscription, why should I have to do it myself?
(and for that 1% of the cases where the address is not a spammer and user knows it, they can just hit "unsubscribe" manually)
How will well-behaved scrapers undermine the sustainability of a business? I guess adblocking is one, but we can already do that with uBlock and that's legal. Or adversarial bridging, but that only serves to boost competition.
In other words, the question is flipped; why would well-behaved (i.e. non-DDoSing) scrapers be illegal?
The question was literally,
> What are the legitimate (i.e. legal) use cases for a product such as this?
Nitter is an example of a service that explicitly disrupts Twitter/X's way to make money. If they can't make money then they can't provide the service, there would be no Twitter/X, and hence no Nitter. Of course they would try to prevent that kind of behavior and it should be obvious why. Resorting to using a service like this in order to continue using Nitter should raise some alarm bells. Sure you can still do it and rationalize it however you want, but you have to acknowledge you're trying to get the value of the service without paying for it.
Perhaps there are cases where there is a dissonance between a website's TOS and how they are blocking bot traffic? That sounds like a valid gripe. Otherwise, I don't buy the argument.
Ps: context why I need automation for such thing: those lessons are really popular and are announced at unpredictable time / there might be another spot when someone resigns
• I want to automate (or at least semi-automate) downloading bank statements. I've got ~14 accounts (checking, savings, credit card, IRA, investment, HSA) across 7 financial institutions.
It's tedious to go download statements from all of them manually.
• I want to save stories from FanFiction.net (FFN) for offline reading. FFN's terms allow automation as long as it doesn't operate faster than a human [1].
[1] From their TOS:
> You agree not to use or launch any automated system, including without limitation, "robots," "spiders," or "offline readers," that accesses the Service in a manner that sends more request messages to the FanFiction.Net servers in a given period of time than a human can reasonably produce in the same period by using a conventional on-line web browser.
Could you not shoot an email to those institutions asking for a copy of the documents?
They’ll respond within a few days, asking me to log into some web portal to prove that I am me, and then we’re back where we started
Scraping is important for example, to monitor competitors' prices to see the opportunity to raise your own prices.
And let's not forget that Google does a lot more scraping than anyone else and has ridiculous profits from it.
For years I've used my own terminal UI player (di-tui) for di.fm. At some point in the not-so-recent past, di.fm added Cloudflare's WAF, which prevents me from using one of my app's features: managing channel favorites within the app.
To be clear, I'm a paying di.fm customer, and my app only works for paying customers. But now my preferred method of listening to di.fm is slightly hamstrung because Cloudflare's WAF sits between me and little string token available to every browser that accesses di.fm (even non-paying customers).
If there was an accessible API to do what I need, I wouldn't do this because scraping sucks. I have to write 100 JavaScript edge cases to handle all the times the host's servers fail in very weird ways. Plus, walking DOMs on these shitty sites with 10,000 nested divs is not fun. GPT helps with this.
It's net-positive for the host though, as I upload a lot of valuable content that their users genuinely like, but it sucks that I have to be sneaky to get the data I need.
Data portability! Tools like this can be used to allow individuals to export their data from hostile web services trying to hold it hostage.
Legal in the EU, with GDPR.
Let's say I have a few 100 Gbps connections that is mostly idle, is it fine if I direct them at browserless? No? Exactly, that kind of traffic is not wanted.
You can't even block the amount of subnets that's coming for you in a DDoS attack, thinking a human is able to keep up with something like this is pretty blindsighted and naive. The differentiation of network protocols and relay attacks alone is way too slow to be mitigated in most systems.
Trying to defend against malicious bots is tedious at best, impossible at the worst. I don't really see how lowering the bar would be a net positive. People will start using methods that will cause more collateral damage and just reduce user freedoms.
My guess is that this only increases the push towards attestation and attestation-like approaches. Login walls and PAT/Privacy Pass are just the start.
If you can't bypass bot detection now because in response they'll make the bot detection harder then there's really no winning is there?
In addition to that any access restrictions would probably end up being less precise and thus cause more collateral damage.
Which would also mean an increased amount of false negatives, leading to an increase in spammy content.
It just annoys browsers and OS with a smaller marketshare to the point that I'm wondering if it's even legal with antitrust legislation.
Devices are way too cheap for an attestation system to work to counter bots
Betting on device attestation really means betting that computing devices will become more and more expensive in the future since the device cost is the only blocker created by the attestation.
And if there's one thing I'm expecting, it's that's not going to happen, device prices will continue to decline.
Not really. If you start a new server from an IP range with known bad actors, sure. Many have tried and failed to run their own mail servers from Digitalocean, vultr, or even more dedicated-esque server hosts and pretty quickly see how much of a hassle doing so is.
But if you buy a v4 /24 from some reputable or old company and get it assigned to your own AS, you won't have negative reputation and will be fairly successful as long as you have spf/dkim/dmarc set up properly.
Doesn't that have an extremely high barrier to entry? A legal entity, significant ongoing fees, and some infrastructure?
I'm under the impression that this is entirely out of reach for anyone just trying to run an email server at home for personal use. I would love to be wrong!
until you give yourself negative reputation by running abuse bots from that range.
Normally, you can just block non-consumer ISPs, but this site offers "residential proxy" services (basically a botnet-as-a-service), which means that now consumer IP ranges need to be selectively blocked as well.
I think PAT/Privacy Pass will solve this problem as far as normal users are concerned ("normal" meaning "running Windows, macOS, iOS, or Android, on devices with hardware attestation capabilities"), but soon enough we'll arrive in an age where you can only visit so many web pages before you've exceeded your daily internet allowance.
The only solution is actual paid services, or real-user verification (and deduplication, which means little privacy, which means legal problems in much of the world) for free accounts.
If you publish something on the internet for free, you _must_ accept that people can read it automatically. You can make it difficult, make it a bit more expensive, but at the end of the day paid data means paywalls.
I'm not saying I like this, I'm just saying I'm surprised it hasn't happened yet.
Which we're talking about running a real browser downloading all the bloat of a target website, 25GB doesn't go very far. Is this just absurdly expensive?
https://antoinevastel.com/bots/datadome
The first link shows the tests and is the most educational.
Does anyone know of more recent tools that show why bot detection was triggered? Versus just testing against Cloudflare bot detection and getting pass/fail.
Feel free to shoot me an email if you're interested in trying it out! paul@browserbase.com
That's actually one of the reasons why I started https://browserbase.com/. Maintaining headless browser infrastructure can be such a pain. I've spent a lot of time managing headless chrome fleets at scale, so happy to answer any questions.
Does all my crawling, it goes very slow, it's never trigger the bot detectors.
The idea was to have a script use Selenium to launch non-headless Chrome and then wait:
driver = Chrome()
driver.get("https://example.org")
input("Press enter when ready")
I could then manually deal with logging in, answering any CAPTCHA that came up, and navigate to the page I wanted to run my automation. Then I could press "enter" in my terminal and my script would continue.That used to work fine, but then on sites using Cloudflare's CAPTCHA it stopped working. Solving the CAPTCHA would just result in another CAPTCHA.
I tried an alternative Selenium Chrome driver that was supposed to be more stealthy, and tried setting various flags that were supposed to make it so JavaScript could not tell that Selenium was there, and those worked for a while, but then they stopped working.
The results were similar using Selenium with Firefox.
I also tried Puppeteer, with Chromium and Firefox, and they too could not get past the CAPTCHA loops.
I then tried Playwright. With Chromium and Webkit that got the CAPTCHA loops. With Firefox it actually worked. I didn't even see the CAPTCHA. The non-interactive check for not being a bot passed.
Still, the whole approach seems fragile. I don't know if Firefox/Playwright working was due to some fundamental difference between Firefox and the others or just Cloudflare having not yet gotten around to dealing with it.
I've only dealt with scraping on a small scale and I quickly realized that running "browsers as a service" is a pain in the ass, they're not exactly lightweight, they like to get "stuck", balloon in memory or some such.
I imagine your business will be quite successful if reliability is good and the price is right!
We also need to detect and react to UI layout changes, and the headless browser is the only real way to do that.
edit: Firefox Nightly because it gives you much deeper access to Firefox internals.
That will give you a good start on where to even begin looking. web-ext is insanely handy for launching Firefox with a WebExtension already installed (but I did a lot of other work around this to productize the offering).
[0] https://firefox-source-docs.mozilla.org/toolkit/components/e... [1] https://firefox-source-docs.mozilla.org/overview/gecko.html
Let's say you are Wal-Mart, and you'd like to know which of the products you sell are available cheaper at Target. Or which neighbourhoods they deliver to that you don't, or which stores have longer opening hours than yours, or whatever.
You can't legally exchange data with Target directly, that would risk making you an illegal price-fixing cartel. You can legally visit their website, but you don't feel like matching up 10,000+ products.
Instead, you call up a business intelligence firm who has already scraped your site and theirs and matched the products up. They'll send you a CSV, for a price.
"we just provide the tool" defence no longer works in court
like there's even a 7 day trial opening this up to anyone with bad intentions
This will get you past some very mundane bot detections, but really this is like, the very first baby step of a long rabbit hole.
The people who are taking this game seriously are 5-10 years ahead of this step. Good luck ¯\_(ツ)_/¯
There are lots of signals like timings, user tapping and scrolling behaviour, signed sessions cookies that represent browsing flows which may be legitimate or not. And that’s all assuming you’re on a good looking IP. To do this you need a large supply of residential IPs which then leads to the dodgy underworld of botnets.
I’d be surprised if this works for anything but the most basic bot protection, this is an advanced space.
If it does work for those cases, they should be either keeping it quiet and making bank, or boasting about having a secret sauce, not basic stuff like this.
Edit: for apps, Akamai provides an SDK that uses things like your motion data to create a signature that suggests that you're a real user. This signature is either injected into API requests or into a webview session. I'm sure it's crackable if you dedicate significant reverse engineering resources to it, but then you've got to crack every version, crack every other implementation from other companies, etc. Non-starter.
A device that stays rock-steady throughout an entire browsing session isn't necessarily suspicious on its own—for example, you could have your phone laid flat on the table while your browse with your pointer finger—but it can be a useful tell in combination with other suspicious factors.
For web-only, I believe they have a JS only bundle that your site can include which I would imagine does different things, but which would also bring a higher risk profile associated with it. Sites use these risk profiles to determine things like whether to offer specific services, whether to ask for more authentication, etc.
That said, since I wrote that comment, I found out that on Android, both Firefox and Chrome grant access to gyroscope data without a permission dialog, which is extremely surprising. I don't have an iOS device to verify, but apparently Safari gates the API behind a permission dialog.
Why do we even need an actual device? We can emulate if we even need to and set our headers to look like we're coming from a device browser.
> Why do we even need an actual device? We can emulate if we even need to and set our headers to look like we're coming from a device browser.
This one is much harder, your browser, OS, and hardware leave a uniquely identifiable fingerprint (with Javascript enabled). A website can render some graphical pattern on a <canvas> or audio in an audio context, and the resulting output will have minute differences that originate from your rendering and audio pipelines.
Check out: https://amiunique.org/fingerprint https://browserleaks.com/ https://fingerprint.com/ https://coveryourtracks.eff.org/
You can try to fake these, but it all depends on the sophistication of the target website. You can quickly end up in really deep rabbit holes: https://www.nullpt.rs/devirtualizing-nike-vm-1
Humans move their devices even when typing.
https://sensor-js.xyz/demo.html
I think (though am uncertain) that it’s similar for App Store apps too.
My assumption about an SDK like that is that it's getting that access through some assumed mean(s) that the user just clicks past without thinking.
Source: Founder@browserless.io
For example, Zillow goes to great lengths to restrict bots, but I want to get updates if a property changes somehow on Zillow. It's publicly accessible data if you're not using a bot. "IFTTT" is just an example of a service that automates these things. I can write my own automation service, but if I'm using a bot, Zillow will block me (I've actually done this before).