Some Git internals
yurichev.com
yurichev.com
It was fun. Link for anyone interested(not very high quality but watchable):
OTOH git natively supports sparse commits.
Shameless plug: I recently picked up Golang and used it to implement a mini version of Git CLI using these internals [1].
In particular I'm thinking of the case where there's an append-only file keeping a log, to commit the changes every so often without having to make a blob copy of the file first (which may be relatively large)?
The closest would be to create a pack manually that contains a diff against the previous version but that'd require manual work.
It would be possible, for example, to modify git-fast-import to a) allow to take diffs as input b) allow to store those diffs (these are different things to deal with). The downside is that the more packs there are, the slower object lookup is, which can make everything much slower. Newer versions of git have cross-pack indexes to deal with that, though, but I don't think that's enabled by default.
Another option would be to add a new format for loose objects that allows to store diffs, but that has backwards compatibility implications.
Applying a patch in Pijul is in O(p c log n), where p is the size of the patch and c the size of the largest "deletion-insertion conflict" p is involved in, where a "deletion-insertion conflict" is a situation where Alice deletes a block of text while Bob adds stuff in that same block.
Note that this is a rough bound, since all non-conflicting operations in a patch are in O(log n), except those involved in a "deletion-insertion conflict", which are in O(c log n).
So, Pijul is in fact faster than Git for merging (and rebasing). The only tradeoff at the moment is that going arbitrarily far back in history isn't as fast as it could be (this will be fixed very soon).
Is pijul’s on-disk format stable yet?
Probably. The patch format is very unlikely to change. The repository format may change a little bit still.
I'd say it's probably ok to try and learn it now, but you should maybe wait for a few weeks before using it for something serious. On the other hand, we use it for itself, and I use it personally for most of my projects.
Another (complementing) way of getting an idea how Git works is by reading the source of Isomorphic Git[0]. Being in JS and with fewer (as of now) features, it’s a bit more accessible than reference Git.
[0] E.g. https://github.com/isomorphic-git/isomorphic-git/blob/main/s...
Does anyone have a really good visual explanation of what's being done to the tree and such when commits and merges and rebases are done? I still don't quite grok it intimately enough.
[1]: https://git-scm.com/book/en/v2/Git-Internals-Git-Objects
[2]: https://git-scm.com/book/en/v2/Git-Internals-Git-References
[0] https://sransara.com/notes/2019/build-yourself-a-git/ [1] https://en.wikipedia.org/wiki/Directed_acyclic_graph [2] https://en.wikipedia.org/wiki/Persistent_data_structure