Web Scraping with Python
scrapingbee.com
scrapingbee.com
Re: which headless library to use this post mentions Selenium which was one of the first headless libs but from what I've heard probably not the best (in terms of developer experience or reliability or robustness) but that's only hearsay... I've never used Selenium myself. Playwright seems like the best option in town if you want to use Python. Built by the former Chrome DevTools team (meaning those people really know how browser internals work). https://playwright.dev/python/docs/intro
On the memory side, headless you end up with far more memory leaks, having to manage stopping and starting new browser instances while maintaining scraping state. The devops overhead is probably 10x more with headless.
In the past I've built scraper infrastructure (headless pools, credential stores, proxy managers, agent profiles) and managed to get a pretty efficient service by tracking cpu,memory,network usage per each job and writing specialized versions. I got pretty far trying to automatically generate specialized scrapers from previous requests but I moved to other projects.
Scraping becomes boring really fast if you don't use the data in meaningful ways.
There's also a software marketplace where you can order custom scrapers. 98% of the projects ran thru it have a 5-star rating (Disclaimer: I moderate that marketplace). Pro-tip: submit your project with a Gmail address to skip sales and reach me directly.
Let me know how I can reach out to you!
Case study: YouTube.js
Obfuscated APIs like Pokemon GO and Netflix are in the tiny minority.
I did run into Cloudflare DDoS protection and Incapsula, which I will say is pretty irritating and IMO antithetical to the web. Incapsula is so bad I get captcha'd just browsing around in a Firefox private window. If I were polling every few seconds or something I'd get it, but denylisting all AWS IPs or looking for "headless" in the User Agent (or looking at navigator params, testing TLS fingerprints, etc.) is bonkers. It's the laziest kind of upselling from web developers where you're making the site harder to use, but not actually keeping real scrapers out, because they're doing even more JavaScript interventions ahead of the HTTP request and using residential IP proxies.
DDoS protection does throw a wrench into the mix, though I don't blame anyone for using it. DDoS protection might seem antithetical to the web, but... so is DDoS and abuse.
Kind of like how being an asshole is antithetical to getting along as a society but you still have to address the reality that there will always be abusers and bad actors. I also think being able to do what you want with your service is a fundamental part of the web incl putting it behind a captcha. It's just part of the beautiful chaos.
For what it's worth, it didn't even work. Headless Chrome and some editing of the JavaScript environment was all it took. So it's definitely a ripoff.
I wish that were true. Maybe most new web projects do. Unfortunately, most web projects are not new.
For example I've been able to reimplement xmltv scrapers for several sources in less than a 100 lines with Scrapy. It's not hard, just requires a little discretion.
That is, making a scraper that can be pointed at an arbitrary site not known at the time of development.
What does "rendered client-side" mean.
Assuming that "rendered client-side" means interpretation and execution of Javascript is necessary to read a site's textual content, then how much of the web is rendered client-side.
If the focus is on textual content, e.g.,, someone is primarily "scraping" text as opposed to images and video, I would guess that only a minority of the web is "rendered client-side". How would we prove otherwise.
This guess I am making would not be an uneducated one. I have been accessing the web without using Javascript for over 30 years. Today, I still use a text-only browser to render HTML as text/hypertext. This allows me to read the site's textual content, quickly and easily. I initiate most HTTP requests with TCP clients, not the browser. All requests, whether from TCP client, browser, or otherwise, are made through a localhost forward proxy. If most websites were truly dependent on Javascript, it stands to reason I would not be able to read much of the web. In other words, another web user who reads the web with a Javascript-enabled browser should be able to read websites that I could not read. This has not been the case. In fact, I often see commenters on HN complaining that they cannot read a site that I am having no trouble reading. The culprit is often Javascript.
The truth is that I rarely encounter a site that cannot be read with the text-only browser. For example, I can read the content of almost every site submitted to HN. A very small minority of sites I find are, more or less, empty shells with links to some Javascripts but no textual content for the visitor to read. These "landing pages" expect a Javascript-enabled browser that automatically follows links in the page (e.g., to remote Javascript files), and that retrieves, interprets and executes Javascript automatically and indiscriminately.[FN1] In what some might see as a Rube Goldberg design pattern, the scripts then make HTTP requests to the "real" site. In such cases it generally only takes me a few minutes to find the "real" site, often what some refer to as a "JSON endpoint".[FN2] However this process has not lead me to rely on a "headless" browser to read websites.
Honestly, if a majority of sites adopted the "JSON endpoint" approach to serving textual content it would make reading websites even easier for me. I could just retrieve JSON and reformat it to a uniform brand of simple HTML that I prefer, as I already do for some sites. I could make the format of all websites 100% identical. IME, a web of uniformly-formatted content is much easier and faster to digest. I would imagine it would easier for machines to digest as well. The text-only browser I use currently makes the format of all sites look almost the same, since it only uses a single font and so many websites use similar designs. Because it does not automatically follow links or execute Javascript, it also tends to make the "load" time of all sites very similar. For me, this uniformity speeds up the ability to digest web content as compared to using a graphical browser for the same purpose.
FN1. Today we see "modern" browsers incorporating an ever-changing array of "features" and options to try to mitigate the risks of this behaviour.
FN2. Generally, IME, these "endpoints" serve the textual content with minimal markup or sometimes no mark up at all. Thus, the end user is free to format the text into whatever design suits their personal tastes. As a website visitor, this is relatively more efficient IMO than trying to read an infinite number of possible "web designs" which is the approach we currently see on today's www. It is more predictable. With the later approach, visiting a new website with a Javascript-enabled, graphical browser is always a "surprise". It might be easy to read or it might not. Visiting "endpoints" generally does not suffer from this problem.
I'm hurt :(
PS: we spend tens of hours writing those piece of content and even pay a technical editor to spot the typo and make it more readable since we're not native English. You might not like this post, but I can assure that genuine care was put into writing this!
Metaphor would be "Everything you need to know about fixing cars" and the article shows you how to check the engine light, change oil, rotate tires, and replace spark plugs. There's just no way to make a promise that large and have your article be considered high quality.
Would recommend you ignore passing comments with no constructive criticism. The title is going to be a point of contention as it’s a big claim and probably being misinterpreted as not “everything you need to know [to get started]” but rather “everything you need to know [ever is in this one article and you’ll need not read anything else]”.
While I’m glad it’s not GPT-3 level spam, or outsource to third world country for copy level spam, in my opinion the article fails in several fundamental ways, noted above. Putting “genuine care” into something is commendable, but is not a substitute for quality, relevant content.
OTOH you’re getting lots of clicks and views for whatever product you’re selling, and even my comments help the “traction” HN gives it, so it doesn’t actually matter what I think.
most of negative comments will go away
But thanks for your contribution I guess..
I'd add that it's often worth spending some time looking at the website for alternate ways than the obvious one of getting the data you're after. sitemap.xml sometimes give useful hints.
Another golden trick is to learn reverse engineering mobile app APIs with mitmproxy or something like it. Nowadays it's kind of a pain to do since Android has been locking things down more and more, but it's still quite possible. Apps very often provide endpoints that give you structured data when the web version is server-rendered HTML only, have fewer anti-scraping measures and rate limiting, and even provide data that isn't available at all for the web version.
This person would like a word with you -- https://stackoverflow.com/a/1732454
:D
To be clear, I'm not talking about building a syntax tree or a way to generically extract elements based on a CSS path selector. I'm saying if you're only interested in a couple of data points in a 3 MB HTML document, and you're sure they're always between some other specific text or even tags, then it's more efficient to use a simple regex than it is to parse the entire thing, which is computationally expensive when running over a large number of large files.
> using regex to parse data when the data you're scraping has a constant enough structure
Regex is fine, just don't parse the HTML itself.
> I think it's time for me to quit the post of Assistant Don't Parse HTML With Regex Officer. No matter how many times we say it, they won't stop coming every day... every hour even. It is a lost cause, which someone else can fight for a bit. So go on, parse HTML with regex, if you must. It's only broken code, not life and death
I guess the fact that it’s currently very high on the front page of HN kind of confirms this type of post works, though, which is unfortunate.
Always to improve the content we're writing here. What else would you have expected to read in such an article?
Is your guide everything someone needs to know? No. But anyone literate in the ways of modern English understands what you mean.
It is an excellent guide and I think you should consider expanding it & perhaps creating a book.
Please don't be discouraged by the people on here who don't have the skill or courage to write or submit anything.
Then, you don't actually do anything with selenium, click a button / link, or anything interesting.
Genuinely in favor of lesser submission vs. increased noise in submissions. Beginner articles are not taboo, but goes against having high quality insights in general.
PS: Flagging is mechanism to filter by community efforts. Guidelines set some general preconditions to the quality of articles for larger dissemination.
When I had to do it I ended up duplicating each page request twice. Once for scrapy and once again with selenium.
Relying on the page structure only is not a robust alternative.
For sites that are hard to scrape (usually bigger sites that get scraped a lot), I pivot towards buying a data feed. Economies of scale incentivize these data companies towards putting someone on maintaining the feed full-time.
All I get is “your browser is not safe” etc which blocks me completely.
Looking around a bit with the requirement that the online tutorial mention this rather important fact, I found this alternative option, which helpfully notes:
We want to run all our scraping projects in a virtual environment, so we will set that up first.
https://python-adv-web-apps.readthedocs.io/en/latest/scrapin...
Compare and contrast that discussion with the one presented in this post - the above is far superior. Also, I don't understand why one would suggest PostGreSQL to a beginner when sqlite3 is included already in Python, and is going to be easier to use for small databases. Towardsdatascience seems to have a nice intro-to-sqlite3 tutorial.