Git scraping: track changes over time by scraping to a Git repository (2020)
simonwillison.net
simonwillison.net
A fun way to track how people are using this is with the git-scraping topic on GitHub:
https://github.com/topics/git-scraping?o=desc&s=updated
That page orders repos tagged git-scraping by most-recently-updated, which shows which scrapers have run most recently.
As I write this, just in the last minute repos that updated include:
queensland-traffic-conditions: https://github.com/drzax/queensland-traffic-conditions
bbcrss: https://github.com/jasoncartwright/bbcrss
metrobus-timetrack-history: https://github.com/jackharrhy/metrobus-timetrack-history
bchydro-outages: https://github.com/outages/bchydro-outages
As a heads up to anyone trying this stunt, please be mindful that git-diff is ultimately a line oriented action (yeah, yeah, "git stores snapshots")
For example https://github.com/pmc-ss/mastodon-scraping/commit/2a15ce1b2... is all :fu: because git sees basically the "first line" changed
However, had the author normalized the instances.json with something like "jq -S" then one would end up with a more reasonable 1736 textual changes, which github would have almost certainly rendered
diff -u \
<(git ls-tree HEAD^1 -- instances.json | cut -d' ' -f3 | xargs git show --pretty=raw | jq -S) \
<(git ls-tree HEAD -- instances.json | cut -d' ' -f3 | xargs git show --pretty=raw | jq -S)
--- /dev/fd/63 2023-08-10 19:31:03.000000000 -0700
+++ /dev/fd/62 2023-08-10 19:31:03.000000000 -0700
@@ -1,6 +1,6 @@
[
{
- "connections": 5088,
+ "connections": 5089, git config --global user.email "41898282+github-actions[bot]@users.noreply.github.com"
git config --global user.name "github-actions[bot]"
They have the right icon, clickable username and it is as simple as just using this email and name. You or someone else might like to do this, too, so here's me sharing this neat trick I found.https://github.com/TomasHubelbauer/github-actions#write-work...
https://github.com/nmpowell/carbon-intensity-forecast-tracki...
I keep thinking the real power of git-based data-over-time storage is flatter data structures. Rather than one or a dozen or so files, scaped & stored, we could synthesize some kind of directory structure & simple value files - alike a plan9/9p system - that express the data, but where changes are more structurally apparent.
Thoughts?
There's definitely a LOT of scope for innovation around how the values are compared over time. So far my explorations have been around loading the deltas into SQLite in various ways, see https://simonwillison.net/2021/Dec/7/git-history/
Which also one of the inspirations for Git.
Rather than a JSON with a bunch of weather stations, make a directory with a bunch of stations as subdirectories, each with lat, long, temp, humidity properties. Let the fs express the structure.
Then when we watch in git, we can filter by changes to one of these subdirs, for example. Or see every time the humidity changes in one. I don't have a good name for the general practice, but trying to use the filesystem to express the structure is the essence.
Yeah, there's definitely a lot to be said for breaking up your scraped data into separate files rather than having it all in a single file.
I have a few projects where I do that kind of thing. My best example is probably this one, where I scrape the "--help" output for the AWS CLI commands and write that into a separate file for each command:
help-scraper/tree/main/aws: https://github.com/simonw/help-scraper/tree/main/aws
This is fantastically useful for keeping track of which AWS features were added or changed at what point.
While it's easy to gather the data, the friction in analyzing it has always pushed the priority of doing so below other datasets I've gathered.
but conveniently it also serves as a way to track the downtime of github actions, which used to be bad but seems to be fine the last couple months: https://github.com/swyxio/gh-action-data-scraping/assets/676...
I also started but never finished a terms-of-service, privacy-statement tracker. I stopped at the boring part where you'd have find the url for thousands of companies and/or engage others to do it.
By itself a single decompile was hard to parse, but if you do it for each release, commit the decompiled sources, and diff them you can easily see code changes.
So you just run a script to poll for a new client version to drop and automatically download, decompile, commit, and tag.
I'd have a diff of the client changes immediately, allowing insight into the protocol changes to update the private game server code to support it.
I think it's fine to use the term "scraping" to refer to downloading a JSON file.
These days an increasing number of websites work by serving up JSON which is then turned into HTML by a client-side JavaScript app. The JSON often isn't a formally documented API, but you can grab it directly to avoid the extra step of processing the HTML.
I do run Git scrapers that process HTML as well. A couple of examples:
scrape-san-mateo-fire-dispatch https://github.com/simonw/scrape-san-mateo-fire-dispatch scrapes the HTML from http://www.firedispatch.com/iPhoneActiveIncident.asp?Agency=... and records both the original HTML and converted JSON in the repository.
scrape-hacker-news-by-domain https://github.com/simonw/scrape-hacker-news-by-domain uses my https://shot-scraper.datasette.io/ browser automation tool to convert an HTML page on Hacker News into JSON and save that to the repo. I wrote more about how that works here: https://simonwillison.net/2022/Dec/2/datasette-write-api/
That one's a particularly fun demo because it's currently capturing changes to the points and comment count on this thread - a recent example commit: https://github.com/simonw/scrape-hacker-news-by-domain/commi...
Scraping is when you're parsing human-readable content (HTML) and extracting data, as the parent comment correctly points out.
“Git scraping” would intuitively refer to extracting specific information from Git repositories. The naming in the article is therefore confusing. “Snapshotting into Git” would be more accurate. (Git itself uses the term “snapshot” for a reason.)
One of the benefits of catching them in a Git repo is that it helps you spot when their structure changes in ways that may break code that you write on top of them.
> the key element that distinguishes data scraping from regular parsing is that the output being scraped is intended for display to an end-user, rather than as an input to another program
Snapshotting JSON files can be incredibly useful, but I don't think you should call it "scraping".
It has sections covering things called "data scraping" and "web scraping" and "screen scraping" and "report mining", and links to articles about "data mining" and "data mining" and "search engine scraping" as well.
To me, that indicates that the terminology around this stuff is already extremely vague and poorly defined... and the suffix "scraping" is up for grabs for anyone who wants to further define it!
(If you don't like me calling this technique "Git scraping" you're going to /really/ hate the name I picked for my shot-scraper tool https://shot-scraper.datasette.io )
I'm less warm at this point to the general idea behind the hack of dumping the resulting JSON crawl data to GitHub. It's a very roundabout way of approaching basically what something like TerminusDB was made for. It definitely feels like the main motivation was GitHub-specific stuff and not Git, really—namely, free jobs with GitHub Actions—and everything else flowed from that. It turns out that GitHub Actions proved to be too unreliable for executing on schedule, anyway, so we ported the crawler to JS with an eye towards using Cloudflare Workers and their cron triggers (which also come in a free flavor).
I'm not sure what it has to do with git. It seems like any version control system would work. Or, really, the main use of git here is that GitHub is effectively being used as a free database. The snapshots and timestamps are enough to see the changes, regardless of storage format.
Please consider adding a user agent string with a link to the repo or some Google-able name to your curl call, it can help site operators get in touch with you if it starts to misbehave somehow.
Collecting 1 kB every minute might not be a big deal, but collecting 1 MB every minute would cost an AWS-hosted service >$40/year in additional data transfer costs
It's interesting that you can run a scraper at fixed intervals on a free, hosted CI like that. If the scraped content is larger, more than a single JSON file, will GitHub have a problem with it?
Once you get above 5GB I believe GitHub Support may send you a quiet polite email asking you to reconsider!
https://docs.github.com/en/repositories/working-with-files/m... has some more information on limits - they suggest keeping individual files below 50MB (and definitely below 100MB).
It that case the data was about 250gb when fully uncompressed, and IIRC under a gig when stored as a git repo.
It’s a really neat idea, though it can make analysis on the data harder to do, in particular quality control (the aforementioned dataset had a lot of duplicates and inconsistency).
Like everything it’s a process of trading off between compute or storage, in this case optimising storage.
This would have helped so much. Bookmarking this tool. Maybe I will get around to setting this up for this docs site.
Maybe all larger documentation sites should have a public history like this -- if not volunteered by the maintainers themselves, then through git-scraping by community.
pushd ~/data/rfc # this is a GIT repo
rsync -avzuh --delete --progress --exclude=.git ftp.rfc-editor.org::rfcs-text-only ~/data/rfc/text-only/
rsync -avzuh --delete --progress --exclude=.git ftp.rfc-editor.org::refs ~/data/rfc/refs
rsync -avzuh --delete --progress --exclude=.git ftp.rfc-editor.org::rfcs-pdf-only ~/data/rfc/pdf-only/
git add .
git commit -m "update $(date '+%Y-%m-%d')"
git reflog expire --expire=now --all
git gc --prune=now --aggressive
git push origin master
popd
Though I admit using GitHub's servers for this is more clever than me using one of my home servers. Still, I lean more to self-hosting.@simonw Will take a look at `git-history`, looks intriguing!
Not sure how much of a difference it makes to the underlying service but I will also do this with my scraping.
Thank you for point that out
perl -le 'sleep rand 60';
It's the most compact code I could find to do the job. sleep $((RANDOM / 546))
but I guess most cron jobs run with an extremely conservative shell that might not have it.He has tools for parsing them written in Rust: https://github.com/badicsalex/hun_law_rs
and Python: https://github.com/badicsalex/hun_law_py
I'm doing it myself tracking my GitHub star changes: https://github.com/kissgyorgy/my-stars
Parsing the legal acts with the tools you mention looks very interesting! Currently, I simply collect the published XML files whose structure is optimized for laying out the text and not so much for representing a structure of sections and subsections.
https://github.com/Ayesh/Geo-IP-Database/
It downloads the repo, and dumps the data split by the first 8 bytes of the IP address, and saves to individual JSON files. For every new scraper run, it creates a new tag and pushes it as a package, so the dependents can simply update them with their dependency manager.
EDIT: hmmm I realize in addition to that it's also a way to not have to do specific queries over time: the diff takes care of finding everything that changed (i.e. you don't have to say "I want to see how this and that values changed over time": the diff does it all). Nice.
Animation: https://raw.githubusercontent.com/tim-fan/media/main/2021092...
Code: https://gist.github.com/tim-fan/5f601c274a30505b1ae6b989a015...
The trick here is to take some source of information online that's updated frequently and turn that into a historic record of every change made to that source, by setting up a GitHub repository and dropping a YAML file into it setting up a scheduled action.
Achieving the same thing with a time series database would require a whole lot more work I think - you'd need to run that database somewhere, then run code that scrapes and writes to it on a scheduled basis.
If you already have a time series database running and a machine that runs cron I guess it wouldn't be too much work to put that in place.
Git scraping also lets you easily track changes made to textual content, which I don't think would fit neatly in a time series database.