Web Scraping via JavaScript Runtime Heap Snapshots (2022)
adriancooney.ie
adriancooney.ie
Dunno, a lot of the time it actually makes scraping easier because the content that's not in the original source tends to be served up as structured data via XHR- JSON usually- you just need to take a look at the data you're interested in and if it's not in 'view-source', it's coming from somewhere else.
Browser based scraping makes sense when that data is heavily mangled or obfuscated, laden with captchas and other anti-scraping methods. Or if you're interested in if text is hidden, what position it's on the page etc.
That said, it's surprising how many high-traffic sites still use "send an HTML snippet over an AJAX endpoint" - or worse yet, ASP.NET forms with stateful servers where you have to dance with __VIEWSTATE across multiple network hops. Part of the art of scraping is knowing when it's worthwhile to go down these rabbit holes, and when it's not!
It's still generally easier, because you don't have to worry about zeroing in on the right section of the page before you start pulling the data out of the HTML, but not quite as easy as getting a JSON structure.
But perhaps if there's xx endpoints for some inefficient reason, the browser DOM would be better.
What I was saying is that method is not quite as easy as pure JSON to get data from, but still easier to parse and find the specific data for the specific item you're looking for, as it's a very small amount of markup all related to the entry in question.
My interpretation of btown's comment is along the same lines, that it's surprising how many sites still serve HTML snippets for dynamic pages.
But also, some more modern sites with JSON API endpoints will have extremely bespoke session/auth/state management systems that make it difficult to create a request payload that will work without calculations done deep in the bowels of their client-side JS code. It can be much easier, if slower and more costly, to mimic a browser and listen to the equivalent of the Network tab, than to find out how to create valid payloads directly for the API endpoints.
It's either in the DOM or in one or two other payloads.
Isn't this sort-of-why people hide themselves behind Cloudflare, to remove the lowest common denominators of scraping.
The old SSI includes of Apache would probably just be as efficient.
Reading the OP and comments it seems like a generational difference with a younger gen not appreciating server side generation in the same way.
The challenge there is automating it, though - usually the rest endpoints require some complex combination of temporary auth token headers that are (intentionally) difficult to generate outside the context of the app itself and expire pretty quickly.
puppeteer: https://pptr.dev/api/puppeteer.page.setrequestinterception
playwright: https://playwright.dev/docs/network#network-events
If there's ephemeral cookies, they tend to follow a predictable pattern.
Whatever method you were using before SPAs to authenticate your scraper (HTTP requests, browser automation), you can use that same method now.
It doesn't matter if you scrape the DOM or get some Json.
I just need to scrape some public account posts and, I may be dumb, but I dunno how to do that with the official APIs (developers.facebook is hard to understand for me).
But they're sitting behind Cloudflare and aggressively blocking attempts to fetch data programmatically, which is a huge problem for us with 6000+ players worth of data to fetch multiple times every 3 months.
So... I built a Chrome Extension to grab the data at a speed that is usually under their detection rate. Basically created a distributed scraper and passed it out to as many people in the league as I could.
For big jobs when we want to do giant batches, it was a simple matter of doing the pulls and when we start getting 429 errors (rate limit blocking code they use), switch to a new IP on the VPN.
The only way they can block us now is if they stop having a website.
Give one of the commercial VPN providers a try. They're usually pretty cheap and have tons of IPs all over the place. Adding a "VPN Disconnect / Reconnect" step to the process only added about 10 seconds per request every so often.
You'd have to disable CSP manually in your browser config to make it work, but that leaves you with an insecure browser and a lot of friction for casual users. Not sure if you can tie about:config options to a user profile for this use case. Distributing a working extension/script is getting harder all the time.
I do recall intercepting requests when I used a chrome extension to change CSP values though and not needing to when doing something similar later in tampermonkey, but it may not have been quite the same issue as you're describing, so I can't definitively say whether I had a problem with it or not.
the best of the best are 4G rotating proxies
the fingerprint needs to change also
For Instagram, here is a link that maybe helps: https://apify.com/apify/instagram-profile-scraper
I'll just say that firefox still runs tampermonkey, and that includes firefox mobile, so depending on how often you need a different IP and how much data you're getting, you might be able to do away with the whole idea of proxies and just have a few mobile phones that can be configured as workers that take requests through a tampermonkey script. Or that a laptop tethers to that does the same, or that runs puppeteer itself. It depends on whether a worker needs a new IP every few minutes, hours or days as to whether a real mobile phone works (as some manual interaction is often required to actively change the IP).
rotating proxies are the way to go with insta, you can't do much about IP blocking besides using the right IPs
although in theory if you had an account(s) you could still scrape data from a datacenter IP(aws), even though the limits were lower than a 4G proxy
you can buy/create insta accounts for less than a $ using throwaway phone numbers
And if you're determined to scrape [a website], sometimes it needs proxies, rotating user agents and some rate limiting.
Using a full blown browser sometimes helps prevent you hitting those rate limits, but they're still there.
Instagram deliberately nuked their own APIs.
IDK exactly why, I think it had to do with the 2016 U.S. election.
Yes, you can overwrite fetch and log everything that comes in or out of the page you're looking at. I do that in Tampermonkey but one can probably inject the same kind of script in Puppeteer.
A while ago, when I was looking for an apartment, I noticed that only the mobile app for a certain service allows for drawing the area of interest - the web version had only the option of looking in the area currently visible on the screen.
Or did it? Turns out it was the same GraphQL query with the area described as a GeoJSON object.
GeoJSON allows for disjointed areas, which was particularly useful in my case, because I had three of those.
There are some, not many, but when possible I would rather just use a simple request library to fetch it than have to spin up a browser.
99% of the time if it's not in the DOM it's an XHR request to a standardised API with nice, clean data.
So we were both irritated and talked back back and forth, ceding no ground, and suddenly an admin banned me without warning. It was the suddenness more than anything that got to me; I felt misrepresented and censored, and I just wanted to be able to re-read it all to gain closure, if you will. If I could find the file today I would probably cringe at what I wrote, or at least how I wrote it, but back then I nodded and thought "yup, I'm correct", haha.
At some point I got pulled in and ran screaming away from selenium to puppeteer -- and quickly discovered the joy that is scripting the browser via natively supported api's and the chrome debugger protocol.
The partners web page happened to be implemented with the apollo graphql client and I came across the puppeteer api for scanning the javascript heap -- I realized that if I could find the apollo client instance in memory (buried as a local variable inside some function closure referenced within the web app) -- I could just use it myself to get the data I needed ... coded it up in an hour or so and it just worked ... super fun and effective way to write a "scraper"!
OnDocumentReady -> scan the heap for the needed object -> use it directly to get the data you need
Every modern browser has native support for WebDriver (i.e. Selenium) APIs.
The advantage to Puppetteer is that the Chrome Dev Tools API is just a better API.
EDIT: https://github.com/adriancooney/puppeteer-heap-snapshot/blob... is the code that captures the snapshot, and it uses createCDPSession() - it looks like Playwright has an equivalent for that Puppeteer API, documented here: https://playwright.dev/docs/api/class-cdpsession
I have a live coding stream I did the other day scraping Facebook for comments https://www.youtube.com/live/03oTYPm12y8?feature=share
If you're interested in seeing puppeteer in action I started doing streams last month where I talk through my method. I’ll be posting a lot more since it's been very fun.
Overall puppeteer is great because you get to easily inject js scripts in a nice API. Selenium is great too but not as developed of a web scraping interface imo. Also puppeteer is a very optimized headless browser which is a given. What really matters is implementing a VPN proxy and storing your cookies during auth routines which I can get into if you have any questions about that.
$ npx puppeteer-heap-snapshot query \
--url https://www.youtube.com/watch\?v\=L_o_O7v1ews \
--properties channelId,viewCount,keywords --no-headlessSearching the heap manually is not working very well. The data I want is in a (very) long list of irrelevant values within a "strings" key. It might have something to do with the data on the page that I want to scrape being rendered by JavaScript.
As I understand it this only works for SPAs or other heavy js frontends and would not work on HTML.
I think that’s fine.
What I’m really excited is this combined with traditional mark up scanning plus (incoming buzz word) AI.
Scraping is slowly becoming unstoppable and that a good thing.
shot-scraper javascript youtube.com 'document.body.innerText' -r
Output: https://gist.github.com/simonw/f497c90ca717006d0ee286ab086fb...Or access the accessibility tree of the page using https://shot-scraper.datasette.io/en/stable/accessibility.ht...
shot-scraper accessibility youtube.com
Output here: https://gist.github.com/simonw/5174380dcd8c979af02e3dd74051a...It's not like you can rely on a dictionary to confirm you've correctly OCRed a post by "@4EyedJediO" - who knows if that's an O or a 0 at the end?
And if you're OCRing the title and view count of a youtube video, for example, you've got to take the page layout into account because there's a recommendations sidebar full of other titles with different view counts.
However this is a nice hack around "modern" page structures and kudos to the author for making a proper tool out of it.
(Fun fact: I believe that Closure Library, and by extension the Closure Compiler, are still used for the Gmail rich text editor! [1])
[0] https://developers.google.com/closure/compiler/docs/api-tuto...
[1] https://google.github.io/closure-library/source/closure/goog...
That would seem to be the actually interesting/challenging part.
https://developer.chrome.com/docs/devtools/console/utilities...
I have a strip-tags CLI tool which I can pipe HTML through on its way to an LLM, described here: https://simonwillison.net/2023/May/18/cli-tools-for-llms/
I also do things like this:
shot-scraper javascript news.ycombinator.com 'document.body.innerText' -r \
| llm -s 'General themes, illustrated by emoji'
Output here: https://gist.github.com/simonw/3fbfa44f83e12f9451b58b5954514...That's using https://shot-scraper.datasette.io/ to get just the document.body.innerText as a raw string, then piping that to gpt-3.5-turbo with a system prompt.
In terms of retaining context, I added a feature to my strip-tags tool where you can ask it to NOT strip specific tags - e.g.:
curl -s https://www.theguardian.com/us | \
strip-tags -m -t h1 -t h2 -t h3
That strips all HTML tags except for h1, h2 and h3 - output here: https://gist.github.com/simonw/fefb92c6aba79f247dd4f8d5ecd88...Full documentation here: https://github.com/simonw/strip-tags/blob/main/README.md
Roughly, my end goal is to do a single or multi-shot with the following information HTML differential (could be selectors, xpaths, data regions, differentials of any of the above, etc...), code stacktrace, related code, and prompt.
For this example, let's consider that the flow involves the bot to login to a website. I have selectors for the `.username` and `.password` inputs and then a selector for the login button as `.login-btn`.
1. The site updates their page and changes up all their IDs, but keeps the same structure. 2. The site updates their page and changes up all their IDs, but changes the structure and the form is named something different and is somewhere else in the DOM. 3. many... many other examples.
Trying to figure out how to minimize the tokens, but keep the needed context to regenerate the selectors that are needed to maintain the workflow.
My hunch is you could do it with a much more complex setup involving OpenAI functions - by trying different things (like "list just input elements with their names and associated labels") in a loop with the LLM where it gets to keep asking follow-up questions of the DOM until it finds the right combination.
I think one of the most challenging part of web scraping is dealing with the website's anti-scraping measures, such as it needs to sign in, encountering 403 forbidden error, and reCAPTCHA.
Does anyone have more experience in handling that?