Web Scraping in Python – The Complete Guide
proxiesapi.com
proxiesapi.com
More than once, I wrote a scraper that did both of these steps together. Only later I realized that I forgot to extract some information that I need and had to do the costly task of re-crawling and scraping everything.
If you do this in two steps, you can always go back, change the scraper and quickly rerun it on historical data instead of re-crawling everything from scratch.
One of those fundamentals was separating out the steps of landing the data vs subsequent normalisation and transformation steps.
I.e. first preserve your raw upstream via a 1:1 copy, then transform/materialize as makes sense for you, before consuming
Which makes sense, as ELT models are essentially agile for data... (solution for not knowing what we don't yet know)
But the ETL functionality should itself lives in a (sub)system that has its own logical datastore (which may or may not be physically separate from the destination store), and things should be ELT where the L is with respect to that store. So, its E(LTE)L, in a sense.
Assumptions -- We're talking about two separate systems (source and destination) with non-neglible transfer time (although perhaps "quick")
ETL -- Performing the transform before/during the load, such that fields in the destination are not guaranteed to have existed in the source (i.e. 2 db model)
ELT -- Performing a 1:1 copy of source into an intermediary table/db (albeit perhaps with filtering), then performing a transform on the intermediary table/db to generate the destination table/db (either realized or materialized at query time), with the intermediary table/database history retained (i.e. 3 table/db model)
In short distinction, if regeneration or altering the destination is required, ETL relies on history being available in the upstream source.
ELT pulls control of that to the destination-owner, as they're retaining the raw data on their side.
Keeping raw data when possible has been huge. We keep some in our codebase for quick tests during development and then we keep raws from production runs that we can evaluate with each change, giving us an idea of the production impact of the change.
It's honestly not that difficult to build ETL pipelines from scratch. We're using a ton of different sources with different data formats as well. Using the Serverless framework to set up all the Lambda functions and cron jobs also makes things a lot easier.
i have figured out these questions by seeing how more experienced devs do it and on my own, but i want to learn from a book or video series because you can only figure out so much yourself, eventually you need to seek out experts and sometimes the experts around you also figured it out themselves and you need to find an expert outside of your circle. unfortunately a lot of the "ETL experts" teaching stuff online are trying to sell me on prefect or airflow or snowflake etc
Haven't heard much beyond "ask the Old Ones", but "Murphy's law strikes again", "eventually someone will want that data even though they swore it was unnecessary", "eventually someone will ask for a backfill/replay", "eventually someone will give you a duplicate file", "eventually someone will want to slice-and-dice the data a different way" and "eventually someone will change the schema without telling you" have been some things I have noticed.
Even de-duplicating data is, in a sense, deletion (or someone will eventually want to get at the data with duplicates -- e.g. for detecting errors or repeats or fraud or some other analysis that mirrors looking at the bullet holes in World War 2 bombers)
Store the data as close to the original form as you can. Keep a timestamp of when you landed the data. Create a UUID for the record. Create a hash of the record if you can. Create a batch_id if you load multiple things at once (e.g. multiple CSVs). Don't truncate and reload a table - rather, append to it. If you still need something that looks like atomic table changes, I've gotten away with something close: "a view that shows only the most recent valid batch". (Yes this is re-inventing the database wheel, but sometimes you make do with the tools you are forced to use.)
Someone, somewhere, will hand you a file that does not conform to the established agreement. You want to log that schema change, with a timestamp, so you can complain to them with evidence that they ain't sending you what they used to, and they didn't bother sending you an email beforehand...
They're not going to fix it on your timeline, so you're probably going to end up hacking your code to work a different way... Until, you know, they switch it back...
So, yeah. Log it. Timestamp it. Hash it. UUID it. Don't trust the source system to do it right, because they will eventually change the script on you. Keep notes, and plan in such a way that you have audit logs and can move with agility.
I find, in data engineering ,the goal is not to prevent everything, it's to be flexible and prepared to handle lots of change, even silly changes, and be able to audit it, observe it, maneuver around it, and keep the mean-time-to-resolution low.
JS/puppeter seems a bit easier for things like rotating user agents, from article:
> "Websites often block scrapers via blocked IP ranges or blocking characteristic bot activity through heuristics. Solutions: Slow down requests, properly mimic browsers, rotate user agents and proxies."
For this I've used the requests-cache lib.
import requests_cache
requests_cache.install_cache('dog_breed_scraping')
and responses will be stored into a local sqlite file.I send the URLs I want scraped to Urlbox[0] it renders the pages saves HTML (and screenshot and metadata) to my S3 bucket[1]. I get a webhook[2] when it's ready for me to process.
I prefer to use Ruby so Nokogiri[3] is the tool I use for scraping step.
This has been particularly useful when I've want to scrape some pages live from a web app and don't want to manage running Puppeteer or Playwright in production.
Disclosure: I work on Urlbox now but I also did this in the five years I was a customer before joining the team.
[0]: https://urlbox.com [1]: https://urlbox.com/s3 [2]: https://urlbox.com/webhooks [3]: https://nokogiri.org
It's primarily purpose is to render screenshots full-page or limited to viewport or an element. To do that well as it does the HTML has to be rendered perfectly first.
It's not as cheap as other solutions but we have customers who render millions of pages per month with us. They value the accuracy and reliability that's come from over a decade of refinements to the service.
Larger projects can request preferential pricing based on the specifics of the kinds of pages they are rendering.
My Clojure scraping framework [0] facilitates that kind of workflow, and I’ve been using it to scrape/restructure massive sites (millions of pages). I guess I’m going to write a blog post about scraping with it at scale. Although it doesn’t really scale much above that – it’s meant for single-machine loads at the moment – it could be enhanced to support that kind of workflow rather easily.
I use it for my shot-scraper CLI tool: https://shot-scraper.datasette.io/ - which lets you scrape web pages directly from the command line by running JavaScript against pages to extract JSON data: https://shot-scraper.datasette.io/en/stable/javascript.html
Some caveats though:
- It does fail on some sites. I think the value of scrapy is you get more fine grained control. Although I guess if you can use any JS with shot-scraper you could also get that fine grained control.
- It's slow and uses up a lot of CPU (because Playwright is slow and uses up a lot of CPU). I recently used shot-scraper to extract the text of about 90K sites (long story). Ran it on 22 cores, and the room got very hot. I suspect Scrapy would use an order of magnitude less power.
On the plus side, of course, is the fact that it actually executes JS, so you can get past a lot of JS walls.
When you ran it against 90,000 sites were you running the "shot-scraper" command 90,000 times? If so, my guess is that most of that CPU time is spent starting and stopping the process - shot-scraper wasn't designed for efficient start/stop times.
I wonder if that could be fixed? For the moment I'd suggest writing Playwright code for 90,000 site scraping directly in Python or JavaScript, to avoid that startup overhead.
I didn't realize starting/stopping was that expensive. I thought it was mostly the fact that you're practically running a whole browser engine (along with a JS engine).
If I do this again, I'll look into writing the playwright code directly (I've never used it).
Perhaps it's already being made obsolete by LLM technologies? I'd be curious to hear from anyone who's used a locally running LLM to extract written content, especially if it's been built specifically for that task.
What improvements are you looking for? For me, it works over 95% of the time, so I'm happy. Occasionally it excises a section (e.g. "too short" heuristic), and I wish it was smarter about it. But like you, I haven't found better alternatives. I also need something I can run in a script.
> Perhaps it's already being made obsolete by LLM technologies? I'd be curious to hear from anyone who's used a locally running LLM to extract written content, especially if it's been built specifically for that task.
It would be good to benchmark this across, say, 50 sites and see which one performs better. At the moment, I don't know if I'd trust an LLM more than Readability - especially for longer content. Also, I wouldn't use it to scrape 90K sites. Both slow and expensive!
Had a lot of fun building and automating the scraping, especially in order to get around some bot catching rules that they have. For example one of them blocks all the requests originating from non-residential IPs, so I had to use tailscale to route the scraper's traffic through my home connection and take advantage of my ISP's CGNAT.
You can take a look here: https://pricewatcher.gr/en/
* I'm not doing any deduplication or price comparisons between the supermarkets, I only show historical prices of the same product, to showcase the changes.
Seems kinda suspicious to me.
1. You choose your general location or 2. You don't choose
In the first option you get one less general category to choose from (for example they may not have fresh fish)
As far as I can tell, in both cases the supermarket closest to the delivery address is responsible for filling out your order and usually what happens is that they call you to let you know that they don't have something and suggest substitutions.
What Ive long wanted was the the ability to map prices to SCUs by having folks simply take a pic of the UPC + price, just like gasbuddy or what not - in addition to scraping from grocery posting their coupon sheets online for scraping, in addition to people just scanning (non-PII) portions of receipts.
Can you share what you've made thus far?
* could it be used as an automated "price matching" finder? (for those companies that do a "we price match!" thing
I understand that using Playwright in tests is probably the most common use case (it's even in their tagline) but ultimately the introduction section of a lib should be about the lib itself, not certain scenario to use it with a 3rd-party lib B (`pytest`). Especially when it may cause side effect (I wasn't "bitten" by it but surely was confusing: when I was learning it before, I created test_example.py as said in a minefield folder which has batch of other test_xxxx.py files. And running `pytest` causes all of them to run, and gives confusing outputs. And it's not obvious to me at all, since I've never used pytest before and this is not a documentation about pytest, so no additional context was given.)
> tagline
https://playwright.dev/python/docs/intro is actually the documentation for pytest-playwright - their pytest plugin.
https://playwright.dev/python/docs/library is the documentation for their automation library.
I just filed an issue pointing out that this is confusing. https://github.com/microsoft/playwright/issues/29579
https://htmlunit.sourceforge.io/
to crawl Javascript-based sites from Java. I think it was originally intended for integration tests but it sure works well for webcrawlers.
I just wrote a Python-based webcrawler this weekend for a small set of sites that is connected to a bookmark manager (you bookmark a page, it crawls related pages, builds database records, copies images, etc.) and had a very easy time picking out relevant links, text and images w/ CSS selectors and beautifulsoup. This time I used a database to manage the frontier because the system is interactive (you add a new link and it ought to get crawled quickly) but for a long time my habit was writing crawlers that read the frontier for pass N from a text file which is one URL per line and then write the frontier for pass N+1 to another text file because this kind of crawler is not only simple to write but it doesn't get stuck in web traps.
I have a few of these systems that do very heterogenous processing of mostly scraped content and something think about setting up a celery server to break work up into tasks .
Agree that Playwright is great. It's super easy to run on Modal.[2]
1. https://modal.com/docs/guide/workspaces#dashboard
2. https://modal.com/docs/examples/web-scraper#a-simple-web-scr...
I have a relatively sophisticated scraping operation going, but I haven’t found a great way to test methods that are dependent on JavaScript interaction behind a login.
I’ve used Playwright’s har recording to great effect for writing tests that don’t require login, but I’ve found that har recording doesn’t get me there for post-login because the har playback keeps serving the content from pre-login (even though it includes the relevant assets from both pre and post login.)
1. <domain>/robots.txt can sometimes have useful info for scraping a website. It will often include links to sitemaps that let you enumerate all pages on a site. This is a useful library for fetching/parsing a sitemap (https://github.com/mediacloud/ultimate-sitemap-parser)
2. Instead of parsing HTML tags, sometimes you can extract the data you need through structured metadata. This is a useful library for extracting it into JSON (https://github.com/scrapinghub/extruct)
APIs (for SPAs), OpenGraph/LD+JSON data in <head>, and data- attributes with proper data in them (e.g. a timestamp vs "just now" in the text for the human).
Scraping is a lot easier than it used to be.
I've been using Perl and Python for 30 years, and JS for a few weeks scattered across those same years.
I also suspect that DOM-like APIs are somewhat overrated here with regards to web scraping. JS/Node would only have an emulation of DOM APIs, or you're running a full web browser (which is a much bigger ask in terms of resources, deployment, performance, etc), and to be honest, lxml in Python is nice and fast. I generally found XPath much better for X(HT)ML parsing than CSS selectors, and XPath support is pretty available across a lot of different ecosystems.
One scraper is often not hugely valuable, most companies I've seen with scrapers have many scrapers. This means that the time investment available for each one is low. Some companies outsource this, and that can work ok. Then scrapers also break. Frequently. Website redesigns, platform moves, bot protection (yes, even if you have a contract allowing you to scrape, IT and BizDev don't talk to each other), the site moving to needing JavaScript to render anything on the page... they can all cause you to go back to the drawing board.
The concept of "tech debt" kinda goes out of the window when you rewrite the code every 6 months. Instead the value comes from how quickly you can write a scraper and get it back in production. The code can in fact be terrible because you don't really need to read it again, automated testing is often pointless because you're not going to edit the scraper without re-testing manually anyway. Instead having a library of tested utility functions, a good manual feedback loop, and quick deployments, were much more useful for us.
Because Perl excels at text processing.
This stuff is much easier to do in Python.
1. the async nature of JS is surprisingly detrimental when writing scraping script. It's hard to describe, but it makes have a mental image of the whole code base or workflow harder. Writing mostly sync code and only use things like ThreadPoolExecutor (not even Threading directly) when necessary has been much easier for me to write clean, easy-to-maintain code.
2. I really don't like the syntax of loops or iterations in JS, and there are a lot of them in web scraping.
3. String processing and/or data re-shaping feels harder in JS. The built-in functions often feel unintuitive.
4. Having the scraped data in Python-land makes it sometimes way easier to dump it into an analysis landscape, which is probably Python, too.
Hey don't worry there's probably a library that does it for you! It only pulls down a half gigabyte of dependencies to left-pad strings!
I hate the JavaScript ecosystem so fucking much.
Maybe I lack experience but I don't see JS being a barrier.
I would like to know what kind of string work are you doing. I can't imagine being dependent on parsing strings and such, that looks very easy to break, even easier that css selector dance.
As a basic example, `lstrip()` and `rstrip()` trim whitespace by default, but can also remove any characters you'd like from the end(s) of your string.
You'd need to code that up yourself in JS. Happens quite a lot.
Take this with a grain of salt, since I am fully cognizant that I'm the outlier in most of these conversations, but Scrapy is A++ the no-kidding best framework for this activity that has been created thus far. So, if there was scrapyjs maybe I'd look into it, but there's not (that I'm aware of) so here we are. This conversation often comes up in any such "well, I just use requests & ..." conversation and if one is happy with main.py and a bunch of requests invocations, I'm glad for you, but I don't want to try and cobble together all the side-band stuff that Scrapy and its ecosystem provide for me in a reusable and predictable way
Also, often those conversations conflate the server side language with the "scrape using headed browser" language which happens to be the same one. So, if one is using cheerio <https://github.com/cheeriojs/cheerio> then sure node can be a fine thing - if the blog post is all "fire up puppeteer, what can go wrong?!" then there is the road to ruin of doing battle with all kinds of detection problems since it's kind of a browser but kind of not
I, under no circumstances, want the target site running their JS during my crawl runs. I fully accept responsibility for reproducing any XHR or auth or whatever to find the 3 URLs that I care about, without downloading every thumbnail and marketing JS and beacon and and and. I'm also cognizant that my traffic will thus stand out since it uniquely does not make the beacon and marketing calls, but my experience has been that I get the ban hammer less often with my target fetches than trying to pretend to be a browser with a human on the keyboard/mouse but is not
PLEASE PLEASE PLEASE establish and use a consistent useragent string.
This lets us load balance and steer traffic appropriately.
Thank you.
One example is www.reuters.com. It makes no sense because the site works fine without Javascript but Javascript is required as a result of the use of DataDome. See below for example demonstration.
For anyone who is doing the scraping that causes these websites to use hacks like DataDome: Does your scraping solution get blocked by DataDome. I suspect many will answer no, indicating to me that DataDome is not effective at anything more than blocking non-popular clients. To be more specific, there seems to be a blurring of the line between blocking non-popular clients and preventing "scraping". If scraping can be accomplished with the gigantic, complex popular clients, then why block the smaller, simpler non-popular clients that make a single HTTP request.
To browse www.reuters.com text-only in a gigantic, complex, popular browser
1. Clear all cookies
2. Allow Javascript in Settings for the site
ct.captcha-delivery.com
3. Block Javascript for the site
www.reuters.com
4. Block images for the site
www.reuters.com
First try browsing www.reuters.com with these settings. Two cookies will be stored. One from reuters.com. This one is the DataDome cookie. And another one from www.reuters.com. This second cookie can be deleted with no effect on browsing.
NB. No ad blocker is needed.
Then clear the cookies, remove the above settings and try browsing www.reuters.com with Javascript enabled for all sites and again without an ad blocker. This is what DataDome and Reuters ask web users to do:
"Please enable JS and disable any ad blocker."
Following this instruction from some anonymous web developer totally locks up the computer I am using. The user experience is unbearable.
Whereas with the above settings I used for the demonstration, browsing and reading is fast.
The CAPTCHAs and walls are more of a desperate, doomed retreat.
Do you have any piece of advice for me?
1. use mobile phone proxies. Because of how mobile phone networks do NAT, basically it means that thousands of people share IPs and are much less like to get blocked.
2. Reverse engineer APIs if the data you want is returned in an ajax call.
3. Use a captcha solving service to defeat captchas. There's many and they are cheap.
4. Use an actual phone or get really good at convincing the server you are a mobile phone.
5. Buy 1000s of fake emails to simulate multiple accounts.
6. Experiment. Experiment. Experiment. Get some burner accounts. Figure out if they have request per min/hour/day throttling. See what behavior triggers a cloudflare captchas. Check if different variables such as email domain, useragent, voip vs non-voip sms based 2fa. your goal is to simulate a human. So if you sequentially enumerate through every document - that might be what get's you flagged.
Best of luck and happy scraping!
Flyscrape[0] eliminates a lot of boilerplate code that is otherwise necessary when building a scraper from scratch, while still giving you the flexibility to extract data that perfectly fit your needs.
It comes as a single binary executable and runs small JavaScript files without having to deal with npm or node (or python).
You can have a collection of small and isolated scraping scripts, rather than full on node (or python) projects.
But as more than enough websites there days are just an empty shell I am working on adding browser rendering support.
Really helps if you need to tweak your script and you're being rated limited by the sites you're scraping.
Only step you missed was embeddings to avoid all the privacy pages, and a cookie banner blocker (which arguably the AI could navigate if I cared).
import pandas as pd
tables = pd.read_html('https://commons.wikimedia.org/wiki/List_of_dog_breeds', extract_links="all")
tables[-1]
Highly recommend this approach, it allows you to separate infrastructure code, that gets highly complex as you need more requests, from actual spider/parser code that is usually pretty straightforward and project specific.
Couldn't recommend them more.
ScrapingBee really struggles on so many domains - ScraperAPI is almost as good as Brightdata when it comes to hard to beat sites.
Would you mind sending me the account you used as I wasn't able to find anything under Thomas Isaac or Tillypa and couldn't see what was going wrong then.
I'm sure your comment has nothing to do with the fact that you share the same investor as ScraperAPI but I just wanted be sure.
It's been lost to a bad bookmark setup of mine, and if anyone has a lead on that resource, please link, thank you and unlimited e-karma heading your way.
They are both powerful yet pragmatic dependencies to add to a project. I really like that both packages contain wheels or scriptable methods to install their underlying platform-specific binary dependencies. This means you don't need to ask end users to figure out some complex, platform-specific package manager to install playwright and pandoc.
Playwright let's you scrape pages that rely on js. Pandoc is great at turning HTML into sensible markdown.
For example, below is an excerpt of the openai pricing docs [3] that have been scraped to markdown [4] in this manner.
[0] https://playwright.dev/python/docs/intro
[1] https://github.com/JessicaTegner/pypandoc
[2] https://github.com/paul-gauthier/aider
[3] https://platform.openai.com/docs/models/gpt-4-and-gpt-4-turb...
[4] https://gist.githubusercontent.com/paul-gauthier/95a1434a28d...
## GPT-4 and GPT-4 Turbo
GPT-4 is a large multimodal model (accepting text or image inputs and
outputting text) that can solve difficult problems with greater accuracy
than any of our previous models, thanks to its broader general knowledge
and advanced reasoning capabilities. GPT-4 is available in the OpenAI
API to [paying
customers](https://help.openai.com/en/articles/7102672-how-can-i-access-gpt-4).
Like `gpt-3.5-turbo`, GPT-4 is optimized for chat but works well for
traditional completions tasks using the [Chat Completions
API](/docs/api-reference/chat). Learn how to use GPT-4 in our [text
generation guide](/docs/guides/text-generation).
+-----------------+-----------------+-----------------+-----------------+
| Model | Description | Context window | Training data |
+=================+=================+=================+=================+
| gpt | | 128,000 tokens | Up to Dec 2023 |
| -4-0125-preview | | | |
| | New | | |
| | | | |
| | | | |
| | | | |
| | **GPT-4 | | |
| | Turbo**\ | | |
| | The latest | | |
| | GPT-4 model | | |
| | intended to | | |
| | reduce cases of | | |
| | "laziness" | | |
| | where the model | | |
| | doesn't | | |
| | complete a | | |
| | task. Returns a | | |
| | maximum of | | |
| | 4,096 output | | |
| | tokens. [Learn | | |
| | more](ht | | |
| | tps://openai.co | | |
| | m/blog/new-embe | | |
| | dding-models-an | | |
| | d-api-updates). | | |
+-----------------+-----------------+-----------------+-----------------+
| gpt- | Currently | 128,000 tokens | Up to Dec 2023 |
| 4-turbo-preview | points to | | |
| | `gpt-4 | | |
| | -0125-preview`. | | |
+-----------------+-----------------+-----------------+-----------------+
...> Features: Excellent HTML/XML parser, easy web scraping interface, flexible navigation and search.
It does not feature any parser. It’s basically a wrapper over lxml.
>lxml
> Features: Very fast XML and HTML parser.
It’s fast, but there are alternatives that are literally 5x faster.
This article is just another rewrite of a basic introduction. It’s not a guide, since it does mot describe any issues that you face in practice.
lxml even has a module for using beautifulsoup's parser.
> lxml can make use of BeautifulSoup as a parser backend
https://lxml.de/elementsoup.html
> A very nice feature of BeautifulSoup is its excellent support for encoding detection which can provide better results for real-world HTML pages that do not (correctly) declare their encoding.
Most of the time it won't even register on the scale, compared to the time spent sending/receiving requests and data.
What alternatives are 5x faster?
It's a boring but challenging problem.
I've started using LLMs to generate web scrapers and data processing steps on the fly that adapt to website changes. Using an LLM for every data extraction, would be expensive and slow, but using LLMs to generate the scraper code and subsequently adapt it to website modifications is highly efficient.
The service is using many small AI agents that basically just pick the right strategy for a specific sub-task in our workflows. In our case, an agent is a medium-sized LLM prompt that has a) context and b) a set of functions available to call. Tasks involve automatically deciding how to access a website (proxy, browser), naviage through pages, analyze network calls, and transform the data into the same structure.
The main challenge:
We quickly realized that doing this for a few data sources with low complexity is one thing, doing it for thousands of websites in a reliable, scalable, and cost-efficient way is a whole different beast.
The integration of tightly constrained agents with traditional engineering methods effectively solved this issue.
Feel free to give it a try: https://www.kadoa.com/add
Also, do you support scraping private sites, ie. sites that require a login/password to access the data to scrape?
Thank you!