Show HN: git-history, for analyzing scraped data collected using Git and SQLite
simonwillison.net
simonwillison.net
The benefit is every language can read sqlite, theres a rich ecosystem of sqlite tools like datasette and peewee to build with and sqlite can act as a serverless database keeping infrastructure costs low for serving the data.
I also like that sqlite has built in full text search and fuzzy matching capabilities making it quick to cook up a nice webapp for static data.
Also depending on your use case its nice to not have a dependency on pandas or take up a lot of room in your lambda package size for instance.
Edit: PS we use Apache arrow and save as parquet files, which has lots of benefits over SQLite because you can democratize the dataset by using Athena / BigQuery.. and I diplomatically agreed to use Go to consume the data and serve the API.. so that’s been time consuming.. but fun!
So how can I as an engineer do the most good for the widest possible audience? Well, I was in a meeting and I was taking about using SQLite and my peer was like “I dunno it sounds a little complicated” and I took that to heart and responded with, well ya know from where I sit it seems like all the big players are just using flat files in cloud storage as data lakes.. didn’t sleep well that night and then just perusing other options I came across Apache Drill and then it’s spiritual successor, Arrow and it hit me.. well if Athena / BigQuery understands parquet files and Apache arrow writes parquet files without any sort of infra like spark.. what if our data pipeline didn’t save it in SQLite but Apache Parquet and by doing so (with still plenty of complexity ironically), we all of a sudden have a dataset that can technically be read by API developers and the cloud data lake products. All one really needs at that point is access to a cloud bucket..
It’s still been a lot of work, and I haven’t realized all my promises as we’re still trying to ship the product, but I think in the long run, cloud storage + Apache arrow will be the de facto approach to time series data. And in some ways, it also can give event stores in the context of eventsourcing a facelift too.. but that’s not as developed of a thought for now.
"Dolt marries two familiar concepts, Git and MySQL. The first and only SQL database that supports clone, branch, and merge."
it was interesting in that it gave visibility into back corrections for old days. (and provided data for predicting a final count which took 3-5 days to stabilize, before those 3-5 days elapsed)
For example, here is a program which, given a file's SHA1 and path, finds the last commit to have that file version at that path: https://github.com/cybershadow/misc/blob/master/git-find-fil...
Git stores every version of a file as distinct compressed copies (objects). Periodically, it packs these objects into a pack file. To save space, delta encoding is used in this process. As far as I recall, this is done on a line by line level. In this case, simply dumping this JSON data is going to hamper the delta encoding and the repository will grow substantially in size each time there is a change in data.
Then there's the issue of retrieving the data. The basic implementation of reading objects using git checkout and rewriting the objects to disc but there are other ways to read the raw file contents into memory and do your stuff there.
It's not going to be the most performant or best way of version controlling data but in many cases it will do a decent job.