Git’s database internals II: commit history queries
github.blog
github.blog
From my experience and measurements as the author of `gitoxide` both the parsing of commits as well as the binary search to find the object are negligible costs here. The majority of the time is spent in `zlib` which has to inflate objects and deltas prior to use. This makes me wonder what `zstd` would do to these kinds of workloads.
In any case, the object database decoding performance is the bottleneck when traversing commits, and one can expect to get about 120000 commits per second on the linux kernel pack on a modern CPU core like M1 with `gitoxide`s implementation.
Besides that, I find the explanation of the commit-graph file interesting and how it relates to commit-graph queries. There is so much to learn, and so much still to implement :).
That said, here is how to receive a pack from the remote and resolve it: `mkdir out && gix -v no-repo pack receive https://github.com/Byron/gitoxide out`.
> Graph databases need not apply
The previous blog entry had some of the same apologia about database theory and why git seems to be breaking rules. This seems misplaced to me. Git is a foundational technology of the modern internet, used and relied on by... basically everyone to do exactly the job for which it was engineered. It's exactly the place where you'd expect to see a special-purpose data store in use, tuned to exactly the application space where it resides.
That's not to say it's not interesting to talk about graph databases or B trees (from the last post) and how they might be applied to source control. But please don't stop in the middle of your discussion about git to do it. We know git stands alone, tell us how git works.
Commit metadata is decidedly relational (or, rather, can be modeled as such). Commits have zero, one, or two parent commits (thinking of root commits as parentless), they have an author, a date, a commit synopsis and a commit message, a root object, and maybe some other metadata (e.g., who pushed, who signed off, etc.). That's all perfectly appropriate for a relational DB.
The main issue that has come up over time with Fossil-style DVCSes is the size of the metadata you must keep and examine to do things like `git log -- some-file`. The metadata size issue is always going to be an issue, but maybe an ad-hoc DB can wring more compression out than a general purpose DB. Git itself had the second problem, and it had to get solved by using a Bloom filter to reduce the set of commits that need to be examined.
Try git cat-file -p $SHA