Cloudproxy – hide your scrapers IP behind the cloud
github.com
github.com
I've had some success in scraping lately with a similar project called FlareSolverr(1).
It's purpose it to get you access to sites which won't let you crawl unless you are using a real browser (e.g amazon, instagram). It doesn't hide your IP but uses puppeteer with stealth mode to get you access to otherwise restricted urls.
The main benefit of using the API is that you can request a LOT of data without hitting their rate limit. Unless you need to get dozens of results per second, you are usually better off with a spider (or just use Huginn). And if you are hit with a 503 and a captcha, gluing a free captcha solver with some middleware is a trivial task.
Also, in my experience Amazon's "ban" comes down to solving a captcha on every request, so it's more like some mild throttling than a real ban.
I haven't looked into this space, but what is a free captcha solver? And what is the drawback? Wouldn't this defeat the purpose of a captcha?
The reason Amazon have a reputation of being good at blocking IPs is that their responses are (purposefully?) obscure. The way it works, it filters out script kiddies and lets engineers through, which probably are a small minority of the people scraping Amazon.
https://developers.google.com/search/docs/advanced/crawling/...
Or is that always done via manual review?
I think sites like Zillow can detect after your third or fourth interaction that your actions aren't very human like and will prompt you a captcha.
From what I've seen the cloud IPs are the first to get blocked by anti-scraping tech.
e.g. Built an amazon scraping thing at home. Worked fine. Deployed to cloud. Literally first request Amazon goes nope you're a bot
It’s such a weird field to be in. It’s not illegal by definition of law, but you’re definitely in shady territory, with most of the customers being of the “get rich quick” persuasion. Or at the very least trying to cut corners. One way or another, they were not playing by the rules ;-)
They sell these "residential" IPs to Amazon, other ecommerce retailers and shady people for an extreme price.
In the e-commerce world, scraping is necessary to stay in business. Amazon has armies of scrapers constantly monitoring their competitors, and in some cases automatically undercutting price updates.
I'd expect cybercriminals and fraudsters also find a pool of disposable residential IPs to be very useful.
Just like everything else in this industry, the retort will be “but it’s the users fault! They didn’t scroll through the 2,375 page privacy policy, user agreement, hold harmless indemnification agreement, and terms of use when they agreed to an ad free experience for their mahjong app! What a stupid user! Ha!”
To that I say, enjoy the coming oppressive regulations. It’s already started. As a kid, I always wondered why we required stupid laws to regulate common sense. Now I know why.
If you're foisting an adhesion contract on more than 1,000 people, they are not deemed to have "agreed" unless a majority of a random sampling (say, ten) of them actually read and understood the entire document. Otherwise it's void. "Read and understood" is decided by a jury as part of any litigation involving the contract. "Random sampling" is made by court evidentiary procedures.
If the contract is negotiated or it was presented to less than 1,000 people the rules stay the way they currently are, since those are the kinds of contracts that English common law was developed for.
things like this are ... iffy at best: https://www.vice.com/en/article/pga9yk/your-tool-to-access-n...
https://www.scotusblog.com/case-files/cases/linkedin-corp-v-...
There is a good discussion of the more nuanced situation here:
The IPs for the different cloud providers are super well known, and the big guys either put you behind captcha hell, or just flat out block you if you're coming from one of them.
They follow links that are explicitly marked as do not follow, they do not even try to limit their rate, they spoof their user agent strings etc etc. These bots cause real problems and cost real money. I do not think that kind of misuse is ethical. In fact, using this tool to circumvent protections can turn your scraping into a DDOS attack, which I do not feel are ethical.
If your bot behaves itself though, public information is public imo. Just don't take down websites, respect rate limits and do not follow 'no-follow' links.
To give an idea of the size of the issue, we have websites for customers that have maybe 5 hits per minute from actual users. Then _suddenly_ you go to 500 hits/minute for a couple of hours, because some bot is trying to scrape a calendar and is now looking for events in 1850 or whatever. (Not the greatest software that these links are still there tbh, but that is out of my control.)
Or another situation, not entirely related, but interesting i think: A few years back for days on end 80% of our total traffic came from random IPs across china & request could be traced through HTTP referrers where the 'user' had apparently opened a page in one province, then traveled to the other side of China and clicked a link 2 hours later.
All these things are relatively easy to mitigate, but that doesn't make it ethical.
The issue with bots hosted on AWS or any cloud for that matter is that as a web host you can't just block the IPs because legitimate traffic comes from them in the form of CMS plugins, backups, etc.
The only thing that rel="nofollow" does is tell search engines not to use that link in their PageRank computation.
If you do want to block well-behaved crawlers from crawling parts of your site, the proper way to do that is to use robots.txt rules.
Been there, done that - at least on the side of fixing it. Anyone who implements a calendar, don't make pervious and next links that let someone travel time forever.
I've always wondered how much bit traffic costs us - but never actually tried to figure it out. It is a good portion of our traffic - even when we block a lot of it.
No, but many sites only allow Google to index them now.
I don't think that was the original plan for the web.
[1] https://techcrunch.com/2021/06/14/supreme-court-revives-link...
That's simply because you're in multi-family housing, it would be a different story if the neighbors smoked and put up a fan to blow the smoke into your yard/window/etc, and that's probably a more apt comparison to bots that literally fork bomb themselves to crawl your site at 500 requests per second.
We are allowed to talk about ethics separate from the law.
What does concern me is the other uses that scraping tools have. For example, what's to stop me from writing a bot specifically to fuck with a competitor's analytics and a/b testing?
Advanced scrapers mostly use residential IPs nowadays, but there are some services to detect those too, e.g. https://focsec.com/
Of course, the websites can raise an abuse complaint with the ISP, and they may take further action.
Create some preemptible instances on google cloud first, then connect them with commands like "gcloud compute ssh instances-name -- -D localhost:port".
And the last step is to connect scrapper to those proxy ports over localhost.
Apple uses Cloudflare for their private relay so there’s a chance that these IPs won’t be blocked.
So, you will not get many IPs this way.
For Christ sakes people, develop your own products!
If you have to develop and use a cloud based tool to "get around" businesses blocking your scraping app, maybe reevaluate your business and life choices.
If you feel that the owner of data or services maintains an illegal monopoly over that data, work to lobby government to resolve that issue.
It's not your individual place to unilaterally decide that a business owes you anything other than what is described in their terms and conditions of use. Period.
I mean, they're the ones who gave it to me, I'm just using it. Not my problem if they don't like it. I never saw any terms and conditions when my browser loaded the page, no reason I'm going to go looking for them before I scrape it. I have just as much right to access it as my web browser does or my phone does.
I'm not saying they owe me anything, they're offering it to me, I'm just taking it.
No you are not. If you are deploying proxy containered services on cloud computing hosts orchestrated by yourself to circumvent mechanisms to prevent such behavior you are delibertly and unethically acting in a manner directly not intended by the owner of such data.
But then I realized that’s just a cop out. A lot of structures and systems are in place specifically to make it difficult for people to change things and specifically to maintain monopolies.
Nobody’s going to lobby government to resolve this issue. That’s the point. Real estate companies are super happy to keep the status quo.
Sometimes you have to break some rules to innovate. There is a line of course (where you draw it depends on your own ethical code). But I certainly wouldn’t put “scraping real estate data” behind the line.
Would breaking into their premises to steal or copy the cards be on the same side of the line as the digital variant?
So the question is: would you consider that stealing?
Stop trying to lawyer this through analigies people, think about the actual situation at hand, instead of drawing broken parallels to some hypothetical
"Proof of work" style countermeasures (where data has to be decoded browser-side in expensive ways that are bearable for a regular user but onerous for a mass scraper) is another externality everybody pays for (the total amount of CPU wasted by regular users is at the system level a total waste).
Login walls...
of course no option if your 'target' blocks tor ips
If you are in a position were you need a rotating list of proxy servers in the cloud to scrape, Tor is probably way too slow.