Git for data that commits incremental diffs to Git itself
youtube.com
youtube.com
The burden of checking out and building snapshots from diff history is now borne by localhost, but that may change. Smart navigation of git history from the nearest available snapshots, building snapshots with Spark, and other ways to save on data transfer and compute are all on the table. Merge conflict resolution is in the works. This paradigm enables hibernating or cleaning up history on S3 for datasets no longer necessary to create snapshots, like those that are git removed if snapshots of earlier commits are not needed. Individual data entries could also be removed for GDPR compliance using versioning on S3 objects, orthogonal to git.
The prototype already cures the pain point I built it for: it was impossible to (1) uniquely identify and (2) make available behind an API multiple versions of a collection of datasets and config parameters, (3) without overburdening HDDs due to small, but frequent changes to any of the datasets in the repo and (4) while being able to see the diffs in git for each commit in order to enable collaborative discussions and reverting or further editing if necessary. Some background: I am building natural language AI algorithms that (+) operate on editable training datasets, meaning changes or deletions in the training data are reflected fast, without traces of past training and without retraining the entire language model (I know this sounds impossible), and (++) explain decisions back to individual training data. LLMs have fixed training datasets, whereas editable datasets call for a collaborative system to manage data efficiently.
I am open to everything including thoughts, suggestions, constructive criticism, and use case ideas.
Have you heard about dolt, which is also "Git for Data?"
https://github.com/dolthub/dolt
We also built dolthub, which is like github for dolt databases:
GDPR does require rebasing in some cases as near as we can figure. There are some creative ways to not require taking an outage during this rebase, or some other creative schemes for storing all PII in non-versoined tables. We haven't built any of that yet though, nobody has asked for GDPR support.
On diffs: diffing workload is lazy and needed just once upon requesting a commit for a repo, the bidirectional diff results (incl. index and columns changes) - not the newly changed files - are then committed as csv or pointer objects, git natively supports seeing such objects as new. I had to write an engine to rebuild and serve snapshots from git history, now on localhost with posting to S3, later also in the cloud. Diffing a dataset (csv, Excel, SQL) against the current checked out version, which now resides in a gitignored "datasets" folder in the working directory, now takes ~20 seconds on a 1 GB CSV dataset with 10M rows. Diffing is not always needed, can be bypassed with incremental workflows (new data daily), committing just the diff. I can handle repos of 10s of GBs, with individual datasets of GBs each. Where to put the compute workload of diffing, checking out, and building snapshots is under careful consideration.
Merge will be assisted, with files in S3 or without: 3-way comparing commits boils down to reusing features from the snapshot rebuild engine, starting from the common ancestor and using only the diffs, and handling conflicts. Merging small changes on large datasets involves dealing with the small changes only. Data and diffs can come from S3 if not already on localhost, to merge data, not pointers. Visually presenting diffs in non-adjacent commits requires UI, current git tools would not interpret history correctly.