Web Scraping 101 with Python
scrapingbee.com
scrapingbee.com
import pandas as pd
dfs = pd.read_html(url)
Where ‘dfs’ is an array of dataframes - one item for each html table on the page.https://pandas.pydata.org/pandas-docs/stable/reference/api/p...
https://chrome.google.com/webstore/detail/sqanything/naejbcf...
You can export the results to Google Sheets too. One advantage of the extension is it works with JS rendered tables.
It's worked unbelievably well. It's been running for roughly 5 years at this point. I send a command at a random time between 11pm and 4am to wake up an ec2 instance. It checks its tags to see if it should execute the script. If so, it does so. When it's done with its scraping for the day, it turns itself off.
This is a tiny snapshot of why it's been so difficult for me to go from python2 to python3. I'm strongly in the camp of "if it ain't broke, don't fix it".
Any chance you could tell me your setup for this?
* Set an autoscaling group with your instance template, max instances 1, min instances 0, desired instances 0 (nothing is running).
* Set up a Lambda function that sets the autoscaling group desired instances to 1.
* Link that function to an API Gateway call, give it an auth key, etc.
* From any machine you have, set up your cron with a random sleep and a curl call to the API.
And that should do the trick, I think.
You might as well just call the ASG API directly.
My only recent change is that we no longer use Items and ItemLoaders from scrapy - we've replaced it with a custom pipeline of Pydantic schemas and objects
Rendering the page in Puppeteer / Selenium and then scraping it from there sounds like a lot easier than somehow trying to replicate that in your scraper?
If they're generated server-side like you would expect, and sent to the client, you'd get them the same way you get anything else, by asking for them.
Doing that for web scraping purposes where everything is changing all the time and you have more than one target website is just not feasible if you have to reverse engineer some custom JS for every site. Using some kind of headless browser for modern websites will be way easier and more reliable.
If it's a static website that has consistently structured HTML and is easy to enumerate through all the webpages I'm looking for, then simple python requests code will work.
The less clear case is when to use a headless browser vs reverse engineering JS/server side APIs. Typically, I will do like a 10 minute dive into the client side js and monitor ajax requests to see if it would be super easy to hit some API that returns JSON to get my data. If reverse engineering seems to hairy, then I will just do headless browser.
I have a really strong preference for hitting JSON apis directly because, well, you get JSON! Also you usually get more data then you even knew existed.
Then again, if I was creating a spider to recursively crawl a non-static website, then I think Headless is the path of least resistance. But usually, I'm trying to get data in the HTML, and not the whole document.
what??
Page loads -> Javascript sends request to backend -> it returns data -> javascript does stuff with it and renders it.
Then I can craft my regex/selectors/etc., once I have the data stored locally.
This helps if you get caught and shut down - it won't turn off your development effort, and you can create a separate task to proxy requests.
I'd say 99% of the time you can get by without a browser.
I ended up building like 8 versions of it, literally using every PHP and Python library and resource I could find.
I tried httrack, php-ultimate-web-scraper (from github), headless chromium. headless selenium, and a few others
By far the biggest problem was dealing with JS links...you wouldn't think from the start it would be such a big deal but yet..it was.
Selenium with python turned out to be the winning combination, and of course, it was the last one I tried. Also, this is an ideal project to implement recursion altho you have to be careful about exit conditions.
One thing that was VERY important for performance was not visiting any page more then once because, obviously, certain links in headers and footers are duped sometimes 100s of times.
JS links often made it very difficult to discover the linked page, are certain library calls that were supposed to get this info for you often didn't work.
It was a super fun project, and in the end considering I only worked for 2 months, I shipped some decent code that was getting like 98.6% of the pages perfectly.
The final presentation was interesting...for some reason my client I think got in his head that I wasn't very good programmer or something, and as we ran thru his list of sample sites expecting my program to error out or incorrectly mirror the site, but it handled all 10 of the sites about perfectly and he was rather flabbergasted because he told me it would have taken him a week hand clicking the site for the mirror but instead the program did them all in under an hour.
My favorite part was having a nice working system, then throwing it in the cloud and finding out a socking number of sites tell you to go away if you come at them from a cloud-based IP.
Shouldn't be surprising, but it was still annoying.
Preferably, only give money to one that tells their users this is how it works.
For our purposes, and the websites we needed to track, a traditional VPN was good enough.
Other than that, I'd strongly caution anyone looking into making parallel requests. Always keep in mind the sysadmin and engineers behind the site you are targeting. It's can be tempting to value your own time by making a ton of parallel requests to reduce the overall time of your script, but you can potentially cause massive server load for the site you're targeting. If that isn't enough motivation to cause you pause, keep in mind that the site owner is more likely to make the site hostile to scrapers if there are too many bad actors hitting the site heavily.
The services I found to be reasonably priced for small jobs, but at scale they quickly become vastly more expensive than setting this up yourself. Especially when you need to run these jobs every month or so. Even if you have to write some code to make the open source solutions actually work.
I use their Python-based chalice framework (https://github.com/aws/chalice) which allows you to add a decorator to a method for a schedule,
@app.schedule(Rate(30, unit=Rate.MINUTES))
It's also a breeze to deploy. chalice deployIf authors did name websites they wanted to scrape, or show tests on actual websites, then we might see others come forward with different solutions. Some of them might beat the ones being put forward by the pre-packaged software libraries/frameworks and commercial scraping services built on them, e.g., less brittle, faster, less code, easier to repair.
We will never know.
It is honestly almost never worth it unless you have constraints on what packages you can use and you MUST use regular expressions. Just do your future-self a favor and use BeautifulSoup or some other package designed to parse the tree-like structure of these documents.
One way it can be used appropriately is just finding a pattern in the document- without caring where it is w.r.t. the rest of the document. But even then, do you really want to match: <!-- <div> --> ?
Having a good variety of tests helps.
> tree structure
You'll need a complete language to parse a tree.
I'd love to see more innovation/developer-UX research on the interactions between regexes, document parse trees, and NLP. For instance, "match every verb phrase where the verb has similar meaning to 'call' within the context of a specific CSS selector, and be able to capture any data along that path in capturing groups, and do something with it" right now takes significant amounts of coding.
https://spacy.io/usage/rule-based-matching does a lot, but (a) it's not particularly concise, (b) there's not a standardized syntax for e.g. replacement strings once you detect something, and (c) there's no real facilities to bake in a knowledge of hierarchy within a larger markup-language document.
Services like Scrapingbee and ScraperAPI are serving quite good for such problems. I personally liked ScraperAPI for rendering dynamic websites due to the better response time.
Shameless Plug: In case if anyone is interested, long time back, I had written about it on my blog which you can read here[2]. Now you do not need to setup remote Chrome instance or anything. What all is required is to hit an API endpoint to fetch content from a dyanmic JS rendered websites.
[1] http://blog.adnansiddiqi.me/tag/scraping/
[2] http://blog.adnansiddiqi.me/scraping-dynamic-websites-using-...
Still, I would love to learn more about your approach if you would be willing to share.
That said, there are quite a few services which battle these systems for you nowadays (such as scraperapi - not affiliated, not a user). They are not always successful, but they have an advantage of maaany residential proxies (no doubt totally ethically obtained /s, but that's another story).
Load the page on it and it shows you all the request being made along with the payload as it happens.
Find the one you need, copy the data, endpoint and HTTP verb and recreate it in your language of choice :D
Can you scrape a webasm site?
It'll change, but who knows how much. At least currently, most scraping professionals are not even using headless browsers as their targets are statically rendered.
The problem with web scraping is that you really don't know the ethical point of scraping ends. These days I will reverse engineer a website to minimize the request load and only target specific API endpoints. But then again I am breaching some security measures they have while doing that.
https://www.kashifaziz.me/web-scraping-python-beautifulsoup....
In other words: Is there a drop in library to solve all the big common issues people run into scraping websites in the wild?
At least, that's how I read it.
A couple years ago, I discovered browserless.io which does this job for you and it's amazing. I really don't know how they made this but it just scales without any limit.
It’s been a blessing. Not only can it handle difficult sites, but it’s super quick to write another spider for the easy sites that provide the JSON blob in a handy single API call.
Only problem I had was getting around cloudflare, tried a few things like puppeteer but no luck.
I do wish there was a Go version of it, mostly because I much prefer working with Go, but also because single binary is extremely useful.
As a beginner it makes a lot of sense to iterate on a local copy with jupyter rather than fetching resources over and over until you get it right. I wish more tutorials focused on this workflow.
Most sites IME are pretty easy.
Yes, unequivocally.
> frequently breaks
It can definitely depend on what you're scraping, but in the last few years or so the only project I had trouble with was one where they changed the units for the unpublished API (the real UI made two requests which mattered, one to grab the units, and I missed that in my initial inspection -- it bit me awhile later when they changed the default behavior for both locations).
A few tips:
As much as possible, try to find the original source for the data. E.g., are there any hidden APIs, or is the data maybe just sitting around in a script being used to populate the HTML? Selenium is great when you need it, but in my experience UI details change much more frequently than the raw data.
When choosing data selectors you'll get a feel for those which might not be robust. E.g., the nth item in a list is prone to breakage as minor UI tweaks are made.
If robustness is important, consider selecting the same data multiple ways and validating your assumptions about the page. E.g., you might want the data with a particular ID, combination of classes, preceding title, or which is the only text element formatted like a version number. When all of those methods agree you're much more likely to have found the right thing, and if they don't then you still have options for graceful degradation; use a majority vote to guess at a value, use the last known value, record N/A or some indication that we're not sure right now, etc. Critically though, your monitoring can instantly report that something is amiss so that you can inspect the problem in more detail while the service still operates in a hopefully acceptable degraded state.
Although it might just be easier to scrape their api endpoints directly instead of mucking with html if its a dynamic page. The data is structured that way, and easier to query.
In my experience with large scale scraping you're much better off using something like Java where you can more easily have a thread pool with thousands of threads (or better yet, Kotlin coroutines) handling the crawling itself and a *NUM CORES thread pool handling CPU bound tasks like parsing.
As mentioned elsewhere, using anything other than headless isn't useful beyond a fairly narrow scope these days.
Even Google SERP can be scraped with a simple HTTP client.
I imagine, for example, building on the SERP example might hit a wall if you added logged in vs not logged SERPS, iterating over carousel data, reading advertisement data etc.
From what I can observe, 2/3 websites can be scraped without using a headless browser.