How to scrape anything on the web and not get caught
tinyendian.com
tinyendian.com
This is a hobbyists' guide to scraping under the radar. Fine at that scale but quite incomplete for anything remotely mature or wide reaching.
A few times our website went down due to the load going >30, eventually I discovered Google was doing something funky, adding the dynamic domains to the "robot.txt" file fixed the issue. Then some other search engines / scrapers seemed to run into the same issue and started requesting hundreds of thousands of URLs per day (these pages were dynamically generated and took a moderate amount of compute power).
We eventually did have to implement basic anti-scraper rules because it was degrading the user experience.
Do you know any professionals with exceptional experience that is being shared on their blogs? For example, if someone is interested in .NET I can recommend this one: https://www.wiktorzychla.com
This is my favorite consensual alternative: https://github.com/mattes/rotating-proxy
I decided to run scrapper to fetch all the data about available apartments in my city. Thanks to that I was able to browse offer at speed of tinder. It took me a few hours to write all the stuff, it saved me probably weeks.
To avoid getting caught I decided to setup TOR on my raspberry pi and use it as a proxy. It was extremely easy and reliable. Sites were so slow I didn't notice significant performance drop. I didn't care about changing proxies because TOR made it for me.
Except that it is good idea to change User-Agents and add some random delays between calls. Luckily for this case it was enough.
Yeah, same experience. Right now I use luminati.io datacenter IPs that work ok, anyone know of a cheaper option that works well? Scraping tens of millions of pages a month.
Problem with VPN is it's shared and hard to get a lot of IPs, any specific ones I could get say 100 dedicated US IPs for a reasonable price?
The thing is, a lot of scraping goes unnoticed. Maybe you get an extra thousand hits here and there. But every spam campaign gets noticed and results in some percentage of spam complaints from users.
Using a list of proxies and hope that is enough to scrape _anything_ on web?
These companies tend to have armies of lawyers who can swat away even the likes of eBay when it comes to justifying web scraping. Nevertheless, the work required is still the same tedium that others deal with: CAPTCHAs; throttling; IP bans; etc.
I know the typical "travel site" or "comparison shop" use case.
There is also the "darn it, I want this" use case.
However, automated, periodic web scraping that mutates (I.e. ticket or reservation grabbing bots) has always felt a bit squicky to me in the same way DNS squatting does.
Or perhaps you want to monitor mentions of your name, and join the discussion.
Or you want to not lose your past thoughts, because discussions were deep and some may be interesting to re-read in the future.
Or a website is known to allow users delete their comments, or website itself bans users and hides all their content from others for flimsy reasons like "ISIS" in the title, no matter whether it's pro/against/neutral/irrelevant.
Or you want to organize the information differently than the site does.
Or the website is bloated and slow as hell, and you want to use it over gprs, so you create a lightweight/fast/better organized/filtered mirror.
Or the website doesn't have search or it is crap/slooow, and is not indexed in google, like some IRC logs out there.
Or you want to have content available offline.
...
Back in 2013, a guy scraped the results of about 150,000 students giving their 10th grade finals of a particular examination board in India. He showed that not only was there no privacy of student's marks because the roll numbers were all linearly incremented, but there was also mass-scale manipulation of marks going on.
The concept is simple but it's a very interesting read.
https://deedy.quora.com/Hacking-into-the-Indian-Education-Sy...
I was one of the 150,000 kids that gave those exams back in 2013 :)
This could be interesting for people that do scrape sites to know too, what basic reasonable measures can one take beyond looking at logs and doing IP bans?
Establish a baseline as to what constitutes "normal" browsing.
Any user exceeding a threshold of page requests gets rate limited/banned.
Stopping people from recording public information is unrealistic, but you can certainly make it more of a pain in the ass.
Google would be the leader with re-captcha, but as a human, I fail a whole number of them. They are a very annoying experience to your users.
IP bans alone are probably not enough to stop a motivated scraper, or work longterm, so checking whether or not traffic is coming from a cloud provider and throwing up a captcha would be a significant hurdle, while allowing humans on cloud-hosted VPNs to pass through without much trouble.
Is there a reason for using only gatherproxy and sockslist? There are more lists [0] available.
[0] https://github.com/chill117/proxy-lists/blob/4bb8064703b09ee...
So, for instance, they have a pool of servers that have 1000 IPs available. Your account allows connections to go out over 2 of those at a time. If something happens (like one gets banned by whatever service you're scraping), you can get a different set of 2 IPs and keep moving.
While you're still paying a relatively high price for what you're consuming (predominantly bandwidth in this case), you're paying for the flexibility.
How would an average joe even know?