I've been doing scraping for many years and at the end it's always the same, you build a lot of stuff to bypass site restrictions and finally, once you are done, you can start scraping. It all goes fine until the site you are scrapping bans you...
So what I do now?
- proxycrawl https://proxycrawl.com
- node http://nodejs.org
With proxycrawl I don't need to worry about bans or blocks and I can crawl sites like amazon, google and facebook without problems by just calling their API.
With node I can do all my work async and with low memory footprint using simple http.get calls and some logic.
So no framework, no tools, nothing other than those two things