Mastering Web Scraping in Python: Crawling from Scratch
zenrows.com
zenrows.com
My library, Skyscraper [0], attempts to help with these. It’s written in Clojure (based on Enlive or Reaver, both counterparts to Beautiful Soup), but the principles should be readily transferable everywhere.
My most extensive use of Skyscraper to date has been to produce a structured dataset of proceedings, including individual voting results, of Central European parliaments (~500K total pages scraped, ~100M entries). I’ll do a full writeup at some point.
Extraction => https://www.zenrows.com/blog/mastering-web-scraping-in-pytho...
Avoid blocking => https://www.zenrows.com/blog/stealth-web-scraping-in-python-...
Imagine an async-loaded list, that continues loading more content as it comes in, until it displays all of the content available to the backend.
When would you know such a list is finished loading?
This sounds insane, but it's pretty easy and common for an ambitious UXer to key in on, and is something I've seen in production pages.
(In the event you are a UXer, please include some sort of status update! Even an overlaid spinner that disappears solves the problem.)
Usually the response is in JSON and you can ignore the original page. You might have to auth/grab session cookies first, but thats still easier than working with the HTML.
e.g. good luck trying to get much out of youtube.com (or any other video site) without executing JS.
- session persistence
- dealing with cdns
- dealing with regional proxies
- dealing with captchas
- dealing with websocket data
- dealing with custom session handshake sequences
list goes on and on and on, but probably just edge cases haha
DoM emulation and selectors are pretty much equivalent between nodejs and python, you can use css or xpath selectors on html/xml content on either of them. Either way you need to emulate something like a DoM, as neither language/execution-environment has a "native" DoM.
You don't want to execute random javascript code from the web inside your scraper, and just being able to parse the scripts doesn't do you much good. So you're not getting the main advantage I think you're suggesting, being able to emulate page javascript, being able to actually run that code.
Generally if you want to interact with javascript you need to do it in another process (I guess a sufficiently advanced sandbox could work too, an interpreter in your interpreter, but so far that doesn't exist). If you're already going to be running that javascript in a different process for security reasons that different process might as well just be a "remote controlled" web browser.
Historically that was done using selenium, which has good python bindings.
Now days it's being done more with playwright, which started out as a nodejs binding but is moving towards python....
Ultimately I think the reason is that there's no real advantage to using javascript and python is a nicer language with a healthier ecosystem, but your mileage my vary.
Personally I use this method with Puppeteer for advanced pages such as Single Page Apps (SPA) and other pages that depend on JavaScript, CSS, or other features in the page. Another example of an advanced page would be a site where you have to psychical scroll and wait for content to load from a web service. In these cases a headless browser with JavaScript makes the most sense to me.
I've found where it gets tricky with JavaScript is if you have a single missing `async/await` you can introduce bugs in your code that take extra time to solve.
For simple pages I do like Python and that you don't need `async/await`.
I see your point though. Also when I do playwright scripting I normally use async/await, so I guess the grass is always greener ;p
In python I find a missing async/await is apparently very early on and doesn't really take extra time to solve. Maybe just better tracebacks in python?
Colly[1] is a all batteries included scrapping library, But could be bit intimidating for someone new to HTML scrapping.
rod [2] is a browser automation library based on DevTools Protocol which adheres to Go'ish way of doing things and so it's very intuitive but comes with overhead of running chromium browser even if its headless.
As a result, js generated content cannot be scraped, and python scrapers also get blocked very fast as they don't execute fingerprinting scripts.
That opens up massive security problems.
Lxml doesn't work well with broken html, but is an or two orders of magnitude faster for parsing, and same for querying with xpath.
A part from that, there is also Scrapy which is used a lot, but same it is also very slow, it is just horizontally scalable easily.
There are a lot of times in which scrapping doesn't use html parsing, when you are scrapping pages which change a lot of structure, it might be better to go with full text search, and in this case, the faster the better. And in that area Python is far from the best, except when .split() and .join() are enough. Even re.match is slow because of the algorithm they use is slow
And to finish, Requests is also super slow, if you want something fast you have to use pycurl.
Not sure there's anything faster on the javascript side of the fence?
Django and Flask are also very popular libraries, so the language and culture gap isn’t as large as it may seem.
If you need a an quick way to scrape javascript generated content, you can just open your browser console and use `document.evaluate` with an XPath query.
Are you thinking more about the performance, or code maintainability?
One thing that is difficult is updating the BeautifulSoup code every time a website changes design/layout/etc.
That can be a lot of work though, use selenium or the more modern playwright to run entire web pages in a remote-controlled browser.
In particular, it doesn't matter why the browser did those http requests. It could be because the user submitted a form, or clicked a link, or javascript did some AJAX request, it's done by a web worker or browser plugin, or god help us calls some function some Active-X component. Provided the scraper emulates http request perfectly, there is no way server can tell if the request came from the component it expects or a scraper.
It is both a benefit and a curse. It's a benefit because all the complexity of javascript libraries, DOM's and what not goes away. For example, back in the day I've scraped the satellite imagery from maps.google.com. Maps is a giant horridly complex javascript application - you really want avoid understanding how it does what it does. The http requests it makes on the other hand are pretty simple.
However Google didn't want you scrapping it, so they included authentication in there. Authentication always boils down to taking some data they sent in a previous request, mangling it with javascript then sending it back as a cookie or hidden field in a form. You have to replicate that mangling perfectly, which involves reading and understand the minimised javascript. That's the curse. Such reverse engineering can take a while, but it's mercifully rare.
The payoff is speed, and reduced fragility. The speed comes arises because most of the crap a browser downloads is only useful to human eyes, and the scraper doesn't have to download it. Fragility is reduced because GUI's, even web GUI's and especially javascript laden SPA's often want mouse clicks and keystrokes in a certain order, and while particular parts of the screen have focus. For some reason web designers love tweaking their UI's which breaks that order. The data they send back with their forms and AJAX requests is far more stable.
I have given up on BeautifulSoup and Scrapy since so many modern websites use obfuscated JS to hide the underlying data they are serving up, so I feel like its better to just act like a user and slowly walk through whatever site actions need to be done to get to the data you want to ingest.
Needless to say, as many have touched on in this post's comments, scraping reliably, and selectively retrying based on the many tens if not hundreds of different potential errors that can occur (either server side / API limitations, or client side based on the interaction that your browser, i.e. shit crashing, etc) is really almost an optimization problem of its own.
Definitely a boon to have scraping as an option, but as always, licensure of data especially if you want to resell it becomes a major concern that you should be thinking about up-front even if you kinda just want to hack things together in the beginning.
What you're talking about is very sensible and I was equally surprised that Airflow didn't support long running tasks but you can layer over the workflow orchestration system a kind of ad-hoc higher order system that enables what you speak of. It kind of feels ugly but can get a lot done.
There are definitely ways to accomplish what you are saying using a combination of DockerOperators + ephemeral WebSocket servers running within containers as semi-long running tasks, and basically just have a dumb/heavy Redis container that persists to run streaming between the coordination architecture across these data flow jobs.
"Work in progress" lol!
EDIT: updating Airflow from 1.10.10 to 2.1.2 recently was a huge pain in the ass for what it's worth, good luck to all our fellow protagonists that are dealing with multi tens of thousands of task DAG setups... big ooooff
What most (uninitiated) developers do not realize is that web crawling is not for mere mortals.
1. We are at the mercy of the webpage authors.
HTML is a great lanuage to encode information. But most developers (usually webpage authors) see it as a kind of tool for presentation only. Infomation can go anywhere in the document. And they are prone to changes.
2. The internet society frowns on web crawling
You look into any site's TnC, you might come across a clause which prevents you from crawling. The specific word may not be in the legalese, however, it implies any kind of crawling is denied. There is good reason for this - it is mostly done to encourage fair use of the service.
3. No body designs services to be crawlable.
Most big name companies do have some alternative. Like Facebook had "graphs". (Now obsolete). They allowed end users to extract data using simple queries - like " list friends of X who live in city Y, and who is not your friend". But "graphs" feature came a lot later after Facebook launch. Not at beginning.
Usually at the beginning stages of any services we are at the mercy of #1 and #2
For #1 no one ever designs page to have have information always at a standard location. It changes.
4. The tech isn't ripe yet.
This is my personal view. I had been experimenting with puppeteer and selenium behind a corporate environment. I wasn't that happy with the "net" developer experience. I found things like taking a screenshot or pdf buggy. For e.g. to get the latter I have to run my browser in non-headless mode. In headless mode my laptop system policy disabled some extensions important for the webpage to load correctly.
So stuff like extensions don't work at all.
It's not that useful without being able to do so.
Personally, I use Scrapy and it works fine. For best practice, I wouldn't use the Pipeline concept Scrapy provides -don't do data transformation inside scrapy. Simply save the responses and perform the validation and transformations outside of Scrapy. The Pipeline concept is flawed because you cannot create DAGs with it -only serially linked pipelines.
(Disclaimer: My project)
Here is the code to do some of that: https://github.com/kordless/grub-2.0
Link extraction is ignored, but could be done with BS on the rendering of the DOM.
I'm curious if folks here had any other recommendations for scraping SPAs, ie React or Angular applications.
I tried using selenium to get around this but was never successful. The issue has really handicapped my ability to scrape.
https://www.zenrows.com/blog/stealth-web-scraping-in-python-...