Web Scraping with JavaScript
qoob.cc
qoob.cc
From there, you just find the component you want to copy data from and you copy the state or props. Very little string parsing or data formatting required, no malformed data, etc. There's a library floating around on GitHub somewhere that makes loading a simplified version of React Developer Tools inside Puppeteer just a script you eval with a jQuery-like API for selecting React components, but I can't remember the name right now.
Someone could probably do this without needing a headless web browser (via jsdom)
I did this with an investment website, where I was able to retrieve all data using simple python. It _should_ be more robust than parsing react components/html.
Content-heavy websites using React often generate static versions of pages at build time (using e.g. https://nextjs.org/docs/advanced-features/automatic-static-o...). In those cases, there might not be a public API endpoint to fetch the data you want
For applications though, it's definitely easier to just make an HTTP request if you can. However, you're more likely to run into issues like APIs blocking datacenter IPs, rate limiting etc than when it appears you're just loading the website like a human
The library is: https://github.com/baruchvlz/resq
Example code:
// resq is the stringified source of the library
// page is a Puppeteer page
// this line injects resq into the page
await page.evaluate(resq);
// This finds a React component with a prop "country" set to "us"
const usProps = await page.evaluate(
`window["resq"].resq$("*", document.querySelector("#__next")).byProps({country: "us"}).props`
);
// This finds a React component with a prop "expandRowByClick" set to true
const news = await page.evaluate(
`window["resq"].resq$("*", document.querySelector("#__next")).byProps({expandRowByClick: true}).props.dataSource`
);If your whole stack is JS and you need a little bit of web scraping, this makes sense. If you're starting a new scraping project from scratch, I think you'll get far further, faster, with Python or Ruby.
I think from es6 and up this is handled pretty well.
Python/Ruby are far more expressive for these sorts of data manipulation tasks.
I scrape for a living and I work with JS, because currently, it has the better tools.
I prefer JS + jsdom/cheerio as it's closer to the in-browser experience for scraping.
I'm currently working to turn my hobby scraper into something profitable. "Working with strings" is already the least of my concern. I've spent most of the time with finding an architecture / file structure that allows me to
- easily handle markup changes on source-pages and
- quickly integrate new sources
I've feared it would be impossible to handle unexpected structural changes from a multitude of sources. Turns out that rarely happens. Like, once every x years per source's page-type.
String manipulation and collections on Python are not an after thought, the syntax and API make them convenient and easy to use.
Text are DOM nodes. If I were making a business of this I would automate the shit out of it by:
1) Gather all text nodes directly
2) Eliminate all text nodes that only contain white space
3) Add context. Since text nodes are DOM nodes you can get information about the containing element directly from the node itself.
Hands down walking the DOM will be programmatically faster to write and execute than anything else you can write in any language.
Here is some tiny code that does just that: https://github.com/prettydiff/semanticText
My current favourite stack for this is Selenium + Python - it lets me write most of my scraper in JavaScript that I run inside of the browser, but having Python to control it means I can really easily write the results to a SQLite database while the scraper is running.
I wrote a bit about this here: https://simonwillison.net/2020/Oct/16/weeknotes-evernote-dat...
Crawling the list pages and then each edit page in turn allowed for dumping the name and value from each input field to the log as key:value pairs for processing offline.
Navigating paging was probably the biggest challenge.
I totally don't recommend doing this. But it worked for this case.
"I hate to advocate drugs, alcohol, violence, or insanity to anyone, but they've always worked for me." -- Hunter S Thompson
Good practice anyway so you don't overload the site and find your logs empty or full of gaps.
Randomising field names, seeding hidden bogus data and messing with element order was more what I would look at once a persistent scraper was using enough IPs to get around rate limits.
One comment though, once you've past that "I can't do this without a real browser" line in the sand a few times, you end up with a collection of snippets and skills that moves that line much closer. Sure, I'll load the page and watch in browser tools to see what's in the html and what's coming back to XHR calls, but when I've got a directory full of previously used example code to fire up that uses Python/Selenium and deals with "boilerplate" parts, it's a much easier decision to jump that way than the first time I stared at the BeautifySoup documentation.
(When the only tool you have is a nailgun, every problem looks like a messiah...)
Like if you do ad-hoc web scrapping then it's fine to spend time looking for the most efficient way, but if your web scrapping framework is part of a data pipeline that scrapes all sort of website then a browser is the most development time-saving route.
So for example if I buy some electronics module on aliexpress, my scrapper automatically saves all the product description and images to the database right from the browser as I'm making the order.
These details usually contain vital info to use the module, so it's important to me to have an easily searchable reference for all this information. I really don't trust myself to collect all the necessary info manually.
I've also used cheerio when I want to save a functioning local cache of a webpage since I can have it transform all the various multi-server references for <img>, <a>, <script>, etc on the page to locally valid URLs and then fetch those URLs.
If I recall correctly, what was really helpful about it that I could write whatever code I would need to query and parse the DOM in the browser console and the copy and paste it into a script with almost no changes.
It made it really simple to go from a proof of concept into pipeline for scraping material and feeding it into a database.
var links = stew.select(dom,'a[href]');
extended with support for embeded regular expressions (for tags, classes, IDs, attributes or attribute values). E.g.:
var metadata = stew.select(dom,'head meta[name=/^dc\.|:/i]');
It's on npm as `stew-select`
[1] https://github.com/rodw/stew/
[2] there's an optional peer-dependency-ish relationship to htmlparser or htmlparser2 or similar to generate a DOM tree from raw HTML but anything that creates a basic DOM tree (`{type:, name:, children:[] }`) will suffice
(I'm a maintainer of jsdom.)
It doesn't mention puppeteer or why you may need to use something like that. It doesn't mention cookies or sessions or anything like that. And it doesn't mention using proxies or any web scraping countermeasures. It's very easy to make crawling difficult, and only very basic sites are easy to crawl with the methods described in the article.
Our infrastructure actually does procedure for some of our scraping needs: we scrape puppeteer's GH documentation page to build out our debugger's autocomplete tool. To do this, we "goto" the page, extract the page's content, and then hand it off to nodejs libraries for parsing. This has two benefits: it cuts down the time you have the browser open and running, and let's you "offload" some of that work to your back-end with more sophisticated libraries. You get the best of both worlds with this approach, and it's one we generally recommend to folks everywhere. Also a great way that we "dogfood" our own product as well :)
In short: we wanted to dogfood the product at the cost of some time and machine resources
I have myself deployed a Scrapy web scraper as AWS Lambda function and it has worked quite nicely. Every day for the last year now I guess, it has been scraping some websites to make my life a little easier.
It gives you tools to work with both HTTP requests and headless browsers, storages to save data without having to fiddle with databases and automatic scaling based on available system resources. We use it every day in our web scraping business, but 90% of the features are available for free in the libary itself.
Try it out and tell us what you think: https://github.com/apify/apify-js
Brute or Generic Scraping - you need to be able to scrape any site and get the data into your organization to serve to your customers, therefore you probably don't care about manipulating things on a string level and you do care about having something that can handle a JS based site. Here you do not make money from the individual scrapes but being able to have everything for everyone, and thus you cannot afford to spend much extra development effort for a site because scraping that site in itself probably isn't worth much money for you.
Bespoke scraping, here you care about being able to extract data at a very atomic level and you need string manipulation and everything else. Probably you make money on each individual site scraped because the sites have been strategically chosen to enhance a product - for example you have a product serving the legal needs of everyone in the EU but you want to expand into all EEA / EFTA countries, each legal info site you adopt your scraper for is worth lots of money and you put developer effort into getting things at a granular data level matching your data model of legal information.
on edit: changed minimal to atomic
It's fine, and probably faster, to parse HTML with a regex for a wide variety of use cases. You won't release zalgo.
Engineers often love to say you can't do this because regular expressions parse regular languages, and HTML is context-sensitive, not regular, and therefore it's impossible to parse.
What they often miss is that the language actually being scraped may only be regular. If you want to parse a page to see if it has the word Banana on it, then your language may defined as .?Banana.?, and that's regular, it doesn't matter that it's HTML. This even applies to questions like "does this contain <element> in the <head>?", or "is there a table in the body".
HTML is not regular, but you're not implementing a browser, you're implementing the language of what you're scraping, and that may well be regular.
Starting with a real HTML parser is a good way to future-proof your code for when someone asks you to add just one more thing.
For me this just highlights why it's important that engineers understand at some basic what these different things all mean, and what limitations you may have with your solutions, or even those you may want.
<h1 class='Rocks Mineral Banana Poison'>Things I won't eat!</h1>There are definitely cases like this where you have to be careful, but my point still stands that it's important to understand the language you are parsing, and the fact that it might be a regular language. Hell, it could even be Turing complete and then you're out of luck!
In my experience, until you've made the mistakes that not properly parsing html leads to - you mostly jump to naive regex/substring solutions too quickly where you should learn/use well tested html parsing libraries instead. Those mo4re advanced techniques aren't always required, but they're worth knowing and once you know them it's smarter to "over solve" the problem sometimes than "cowboy it" with a regex just because it looks like it'll do the job.
Many HTML documents will have the same data included multiple times, so a lot of the limitations can be avoided by targeting the places that appear the most consistently. Most of the reason why a web scraper would break would be because only one place was being targeted for data, and often very loosely. That place would get changed. Suddenly, you wind up with either a lot of wrong data or none at all.
It's great if each of these processes can be invoked separately, so that after the HTML is saved, you don't need to redownload it, unless the source has changed.
By dividing scraping into; rendering, caching and parsing you save your self a lot of web requests. This also helps prevent the website from triggering IP-blocking, DDOS protection and Rate-limiting.
That way one can build the scraper with the ui in the browser from the extension https://chrome.google.com/webstore/detail/web-scraper-free-w... and scrape on the server.
I just started with this theme and I'm having a lot of unexpected "fun" :)
- get to the webpage (selenium)
- do some clicks to expand certain information (selenium)
- save the html (selenium)
- and parse (selectorlib)
For me, almost everything can be done by css selectors or xpath. Selectorlib allows you to write just a tree of css selectors. The css selectors in the children only apply to currently selected objects.
The nice thing is the magical browser tool of the same name, which makes the first iteration much easier. However, the browser tool output and the python code does not always match, that causes some headaches. Overall, it cut down like 90% of the code and move it into a configuration.
I'd say to anyone learning how to code it's a good exercise in learning. Think a scraper for most sites can be built in an hour or two, depending on navigation and how data is sent to the client.
Definitely noticed a trend towards XHR and JSON responses typically using a numeric ID. Probably the easiest type of site to scrape where you don't need to crawl navigation, simply iterate over a number range and the scraped data is already pretty much structured.
Some sites simply don't work without JS...
as mentioned here apify is the scrapy-in-python version with javascript.
I’ve been thinking about building a web app that scrapes specific subreddits.