Git as a Storage
bronevichok.ru
bronevichok.ru
I recently discovered the joy that is ZFS and everything that comes with it. I understand that the technical underpinnings of git are actually extremely different (and mathematical) _but_ just how far is a ZFS snapshot from a git commit really? It seems like the gap between the two might not need a huge bridge. Could a copy-on-write filesystem benefit from more metadata that would come from being implemented in a more git-like way?
The IBM/Rational Clearcase version control system is an example of building a VCS on top of a versioning file system (MVFS), though MVFS uses an underlying database instead of a copy-on-write snapshot mechanism. https://www.ibm.com/support/pages/about-multiversion-file-sy...
But I don't think ZFS has the equivalent of git merge though.
Assume 1 row per block: Original DB "A" has 2 rows, a snapshot "B" is created, "B" deletes one row and adds a new one.
Is it true that the row which "B" took over from "A" and left unmodified resides on the same block for "A" and "B", so that if the block gets corrupted, both databases will have to deal with that corrupt row?
It shouldn't matter if you have a reasonable setup. If you depend on other files on the drive to continue working after blocks have started to go corrupt, that's not a good system.
I really like the ability to use zfs send/receive over ssh for offsite backups.
I'll admit, I haven't kept up with BTRFS features after abandoning it, so some of the features may have improved.
BTRFS has never really achieved "production" status in most people's eyes (at least, that I've seen), and RedHat removed support for it completely (not that they support ZFS either).
ZFS is also very straightforward once you learn how it works, and that takes very little time. The commands and the documentation (man pages) are thorough and detailed. Conversely, trying to figure out how and why BTRFS does what it does has been a huge challenge for me, since nothing seems to be as straightforward as it could be. I'm not sure why that is.
Development on BTRFS is ongoing, but it is starting to feel as though it's never going to actually finish its core features, let alone add quality of life improvements. As an example of what I mean: I run a Gitlab server, which divides its data into tons of directories, some of which are huge and some of which are not, but many of which have different use cases; a Postgres database, large binary files, small text files, temporary files, etc. With ZFS, I set up my storage pool something like this:
gitlab/postgres
gitlab/git-data
gitlab/shared/lfs-data
gitlab/backups
Now everything is divided up and I can specify different caching and block sizes on each one depending on the workload. When I'm going to do an upgrade I can do an atomic recursive snapshot on gitlab/ and I get snapshots on everything.
BTRFS, as far as I can tell, doesn't let you change as many fine-grained things per-storage-space, and it doesn't have atomic recursive snapshots (and touts this as a feature). I'm not sure if it supports a similar feature to zvols, where you can create a block device using the ZFS pool storage (in case you need an ext4 file system but you want to be able to snapshot it, or similar).
<anecdote> I have never once had a single issue with ZFS and data quality, with the exception of cases where underlying storage has failed. Meanwhile, I've had BTRFS lose data every single time I've tried to use it, often within days. Obviously lots of other people haven't had that issue, but suffice to say that personally, I don't trust it. </anecdote>
Meanwhile...
ZFS doesn't support reflinks like XFS and BTRFS do, so you can't do `cp -R --reflink=always somedir/ otherdir/` and just get a reflink copy (i.e. copy-on-write per-file). On XFS, and presumably BTRFS, I can do this and get a "copy" of a 30-50 GB git repository in less than a second, which takes up no extra space until I start modifying files. On ZFS, I have to do `cp -R somedir/ otherdir/` and it copies each file individually, reading the data from each file and then writing the data to the copy of the file.
ZFS also doesn't come as part of the kernel, so you can run into issues where you upgrade the kernel but for whatever reason ZFS doesn't upgrade (maybe the installed version of ZFS doesn't build against your new kernel version) and then you reboot and your ZFS pool is gone.
You also "can't" boot from ZFS, which is to say you can but if you do something like upgrade the ZFS kernel modules and then update your ZFS pool with features that Grub doesn't understand, you now cannot boot the system until you fix it by booting into a rescue image and updating the system with a new version of Grub. Ask me how I found that out.
In the end, my experience has been that ZFS is polished, streamlined, works great out of the box with no tuning necessary, and is extremely flexible as far as doing whatever it is I want to do with it. I see no real reason not to use ZFS, honestly, except for the "hassle" of installing/updating it yourself, and there's an Ubuntu PPA from jonathanf which provides very up-to-date ZFS packages so that you can get access to bug fixes or new features (filesystem features or tooling features) very quickly, with zero effort on your part.
I'm sure it'll get more stable in the future, but right now I wouldn't trust it with any production data the way I would trust ZFS.
With ZFS on Linux being a thing now, I'd choose that over BRTFS every time.
I'd say the git method is actually pretty low in metadata, and the way you'd improve ZFS snapshots doesn't involve making them more like git.
If you did get that huge amount of work done, you could then approximate git with snapshots alone. Right now, you'd probably want snapshots and dedup to work together to approximate git using ZFS.
But copying a file from one 'branch' to another is the only way to emulate a cherrypick or merge.
So after a while of active use with many branches, you're going to have a lot of redundant copies of files all over. You're no longer making proper use of snapshots, and it becomes less efficient than having a working directory and a directory full of commits that hard link to each other.
I'd like to be able to store audio files uncompressed, so that they could be read directly from the CAS, rather than having to be expanded out into a checkout directory.
That said, even if it is compressed, a command like git cat-file could be used to pipe the contents of the file to stdout or any other program that could use them as input without having to create a file on disk.
$ echo "hello world" > HELLO.txt
$ git add HELLO.txt
$ cat .git/objects/3b/18e512dba79e4c8300dd08aeb37f8e728b8dad | \
> zpipe -d | \
> hexdump -e '"|"24/1 "%_p" "|\n"'
|blob 12.hello world.|
$
The header and the content get concatenated together, and the whole thing gets Zlib compressed. The SHA1 is calculated from the header-plus-content before it gets Zlib compressed. $ cat .git/objects/3b/18e512dba79e4c8300dd08aeb37f8e728b8dad | \
> zpipe -d | \
> shasum
3b18e512dba79e4c8300dd08aeb37f8e728b8dad -
$
What I would like to do is record an audio file (e.g. LPCM BWF), take its SHA1 and store it in the CAS as raw content, then reference it somehow from a Git commit. That way it will be part of the history and will travel with `push` and `clone`, won't get gc'd, etc.> That said, even if it is compressed, a command like git cat-file could be used to pipe the contents of the file to stdout or any other program that could use them as input without having to create a file on disk.
That's a neat suggestion! However, I don't see how it would be compatible with random access, which is important for my application.
https://github.com/git/git/blob/master/builtin/cat-file.c
you can absolutely make tools to expand out & load git repos into content stores. it's going to depend on the content store how you do that.
Maybe to achieve what I've laid out, I really would need to write a Git extension a la Git-LFS. But then vanilla Git wouldn't be able to make full use of it, which undermines the purpose of using Git in the first place.
As an alternative, maybe I just commit the darn audio files to the repo.
• In relative terms, audio files grow smaller ever year.
• Large repository size isn't as critical for a music composition tool as it is for perpetually maintained software source code.
• I'm imagining a tool to prune edit history which would consolidate commits and potentially garbage collect audio files that become unreferenced.
I wish there was a way in vanilla Git to just associate a CAS object containing arbitrary bytes with a commit object, though.
It's the same magic you want to do; Really and truly, there's magic there, but it's a pretty thin and well defined layer of magic.
I wonder if I can abuse the pack file format. Mua ha ha. Probably not but learning about Git innards pays dividends even if the experiments don't work out.
One of the examples it gives is storing a music collection. If I understand correctly, I don't think it automatically compresses every file - or at least gives you the ability to not compress it.
It's a generic distrubuted data structure in git, with identities and signature, and conflict merge. At the moment it's used to store bugs, more to come later.
If you assume fully independent git nodes that interact, then there is no authoritative 'branch', 'repo', or anything. Every node will have it's own world view (repo and history) and to enforce a 'global consensus' you will need a BFT distributed consensus algorithm on top of your git scheme so all bank can agree on transactions and apply patches (accept transactions).
If you assume partial consensus -- only interacting banks of a transaction need to be in agreement -- then you still need a consensus for that smaller group to 'merge' their collective actions into a linear transitional narrative.
For most people using 'git', the "distributed transaction log via git" is owned and managed by GitHub, a central authority.
This is related to the idea of using Git as general storage, in that the undo history can be persisted, and then reconstituted by a new process. The trick would be to make all actions compatible with conversion to and from a commit.
I have dreams of implementing an music composition tool / audio editor with a line-oriented edit-decision-list (EDL) text file format that where changes could map coherently onto git history. Ideally, that EDL format would be an open standard as well: I've contemplated the Pure Data file format and AES31 as possible candidates. This is just at the conceptual stage, though.
• PCM audio files which are captured once and then never modified.
• Line oriented text files for which the traditional "diff" functionality suffices.
Just yesterday I lost some work (only maybe 10 minutes worth) when I was updating my org notes. I staged some files, committed them, but made a typo in the commit message. I ended up reverting the commit when I meant to amend. Then I discarded the the staged reverted changes and noticed the status said I was still reverting. So I looked up the command to get me out of that state. I ran 'git revert --abort' and it blew away my unstaged changes. Ah well, those versioned backups I have Emacs do are going to save me this time, or so I thought.
In magit, it's accessable from the log menu
Create a branch, reset it to the badly worded commit hash and then you can try to ammend again.
https://github.com/terminusdb/terminusdb https://github.com/dolthub/dolt
I would love to create some a script to take a measurement of the current tree, then run a tool that runs my script at ~every commit so I can draw a graph of how the metric changes over time.
It's a bit tricky: if you change the script, you need to re-run the analysis at every commit. It starts looking a little bit like a build system, but integrated over time.
I've thought of calling this GitReduce or similar, since it has some similarity to MapReduce: first a "map" step runs at every commit, then the "reduce" step combines all of the individual outputs into a single graph or whatever.
Ideally Git itself could be the only storage engine, so you can trivially serve the results from GitHub.
Then you also have git bisect but that's more for finding when some metric started to show anomalies.
It would be neat if github could store all its data in git, similar to fossil scm. But I suppose microsoft would not want to lose lockin.
I think the article talks about the "What" part of the problem, but the actual code is much more interesting in the "How" sense.
Like the git-ref stuff makes sense as you read the code
https://github.com/ligurio/git-test/blob/master/bin/git-test...
There was a similar set of additions to svn in the past with "svn propedit" in the workflows which I used in a previous workplace.
It was not pretty, because it was like embedding JIRA into svn - but it meant machines could flip state to state with commits during build+test and restart from that point without an independent DB to track the "current state" & people with commit access could nudge a stuck build out without losing "who did what".
Yes, that would be very nice - it is unfortunate that you have to make API calls (over http) to get things like issues ...
I think you can get the wiki with plain old 'git' ? I forget ...
This is correct. The wiki for a repo is accessible as a separate repository named with a suffix of “.wiki”.
So if user foo has a repo bar with an associated wiki, and the repo URL is https://github.com/foo/bar then you can clone the repo and the wiki respectively over SSH by:
git clone git@github.com:foo/bar.git
and git clone git@github.com:foo/bar.wiki.git
I wish they’d do the same for all other repository meta data including issues, repository description, etcThat is, you need API calls to get things like issues.
Is there a single tool that will handle downloading (and the associated API calls) from all of the major providers ? Or is each API tool specifically for either github or gitlab or sr.ht or whatever ?
What I am wondering is are there any tools that use these APIs that have built-in support for multiple provider APIs ? Or does every tool that (helps you manage or download issues, etc.) just built for a particular provider ?
Thanks.
Most of the other such tools I've seen barely have the resources to import/export a single such API. git-issue only has Github import it looks like. https://github.com/dspinellis/git-issue
There's perceval which is designed to be a generic archival tool and supports lots of APIs, but only dumps them to source-specific formats and would still need a lot of work if you tried to use issues from different APIs together: https://github.com/chaoss/grimoirelab-perceval