Write yourself a Git (2018)
wyag.thb.lt
wyag.thb.lt
If I may do a self plug, I had recently written a note on "Build yourself a DVCS (just like Git)"[0]. The note is an effort on discussing reasoning for design decisions of the Git internals, while conceptually building a Git step by step.
[0] https://s.ransara.xyz/notes/2019/build-yourself-a-distribute...
Almost every rebase user I've spoken with has no idea what the danger is despite it being clearly discussed in the manual page for rebase and despite rebase being listed as dangerous every time it is mentioned in any manual page.
For the sake of new users everywhere, please stop recommending rebase.
In fact, it produces cleaner feature branches for review. Tracking the trunk branch with merges into your feature branch makes for a lot of noisy commits and difficult history to read through when the time comes to diagnose a bug. Rebasing, on the other hand, allows you to neatly put all the changes from your branch (and only those changes) into one or a few neatly-packaged commits.
Comparing a branch to trunk shoudl only shows the actual difference. That you merged trunk multiple times shoudl have zero bearing.
The only way it could ever confuse anyone is if they review every commit and somehow fail to pass over merge commits.
The single most aggravating thing in git are its self-appointed super-users who /almost always/ properly use its power until one day they don't. Then they make life miserable for everyone else while we all "just wait, I'm fixing it".
No individual commit should break the build; one reason is to keep git-bisect working well for future users bug-hunting, without getting stopped because someone didn't keep the commits on their dev branch clean prior to merging (N.B., a maintainer should also reject such PRs). And keeping commits clean usually means needing to rebase occasionally to organize the commits.
And each commit should be reviewed individually, in addition to the whole of the branch / PR.
Not to mention that each commit should be logically laid out, with well-defined changes and well-written commit messages. This usually means needing to rebase a branch when developing non-trivial features or bug fixes, to fold in review feedback.
But as mentioned elsewhere, generally on feature / dev branches, the expectation is that the commits are unstable, subject to change, and should not be built upon (without prior coordination, at least).
Master and stable release branches, on the other hand, should never change or be rebased.
1. I didn't say this affects branch diffing, but rather trawling through history on a single branch.
2. "self-appointed super-users [...] [who break everything]" is a strawman and borderline ad-hominem. If you follow the guidelines I put forth, there won't be any issues collaborating with others.
Also, as a general note, it's actually very difficult to completely destroy information that's been committed at some point. If you're really running into issues with this, don't let fear direct you away from enjoying the greatest features of git. Experiment! Keep trying. Read a good git book (https://git-scm.com/book/en/v2). And learn to use the reflog. Everything you've committed is backed up for a long time even if you've removed all named references to those commits.
Exactly. What's the point of all the "oops, a typo" or "applying code review remarks, part III" commits. Just rewrite. This is the workflow you get eg. with Gerrit.
That said, git porcelain is simply awful. It is inconsistent and full of dangers - there is no way I would dare try new commands without reading up on them, because the names are often misleading. Sometimes I really wish GitHub / GitLab used Mercurial as their foundation. I think the world of programming would be much easier....
There is /nothing/ "dangerous" about rebasing. You just don't rebase branches that are publicly shared without coordinating with the other users, so for some cases (like "master" of an open source project) you don't rebase.
But for your internal workflow, rebase is a KEY TOOL. It's how you write your story of commits. You can't just perfectly nail your commit history the first time you code, unless you are a genius. And what if you're working on a feature, but then you want to commit a certain series chunk of changes to master, so that other features can use that change. Rebase is how you do anything like this. It's core to using and enjoying the beauty of git.
Not to mention reset...
It's non-fast-forward pushes that are dangerous. And they're dangerous whether anyone uses rebase or not.
Yes, don’t rebase branches with multiple authors doing parallel work. But I can’t think of many times I’ve even had to work on a feature branch with multiple authors.
In a gerrit flow, you have to use it quite often, but only on the detached micro-branches that gerrit forces you to use. And the UI provides a nice convenient rebase button.
In a more traditional "trunk" flow, where you're pretending it's svn, you probably want "pull.rebase=true", otherwise you generate a lot of entirely spurious merge commits. This really confused me the first time I used git, long ago.
The "dangerous" case is, as you say, using it to rewrite history that has already been pushed (heresy!). Generally the system will warn you that this requires "--force", at which point you need to stop and think about what you've done.
(It took me a long time to overcome my feeling that history rewrites were inherently wrong - defeating the point of a VCS in some way. I've adapted to it somewhat with the "neatening things up before submitting" view, which took a while to learn. In a traditional VCS there's only one view of the code and everyone shares it.)
People in this thread might also appreciate this essay: https://maryrosecook.com/blog/post/git-in-six-hundred-words
And the more expanded version: https://maryrosecook.com/blog/post/git-from-the-inside-out
It really helped me comprehend Git enough to start understanding the more complex work flows.
[1]: https://pijul.org
In particular, Pijul supports (and depends on) working with repository states that are, in Git terms, not fully resolved. In addition, those states are potentially very difficult to even represent as flat files (see e.g. [2]). Git is simpler in that it mandates that each commit represents a fully valid filesystem state.
That said, I still think Pijul might have a place, if it turns out that it supports superior workflows that aren't possible in Git. But the "VCS elitism" would probably become worse than it is today.
[1]: https://jneem.github.io/merging/ [2]: https://jneem.github.io/cycles/
It also takes a ground-up "build it yourself" approach and has tons of interesting detail.
https://git-man-page-generator.lokaltog.net/
:)
https://github.com/andrewshadura/git-crecord/
Those familiar with Mercurial will surely notice it is, in fact, a port of a Mercurial’s interactive commit functionality, previously a separate extension called crecord.
The only problem now is the time.
There is a video course as well: https://www.git-tower.com/learn/git/videos#episodes
I'm just written a simple git dumper tool (https://github.com/owenchia/githack) a few days ago. Learn by doing is a very good way and I really enjoy it.
Please point them out here
So I think: there must be a better way.
I have often thought about implementing a VCS. The idea behind one doesn't seem particularly complex to me (certainly it's simpler than programming languages). If I did I would quite probably use WYAG as a starting point. My first step would be to define the user's mental model -- i.e. what concepts they need to understand such that they can predict what the system will do. Then I would build a web-based UI that presents the status of the system to the user in terms of that model.
Also to learn that `git reflog` exists, and gives you pointers to all the old states of your repo, even mid-merge or mid-rebase. If you get stuck with a bad rebase, you’re only a `git reflog; git checkout HEAD@{3}` away from being back where you started.
Tutorials like this one and others mentioned throughout the thread do a good job of breaking down the viscera from there.
That's not git, that's the command line - the interface with no inline visualization, no discoverability and no affordances.
git is not its command line interface. I use the git CLI like I use curl: a powerful tool for occasional surgical or automated operations directly on underlying protocol. Most of the time I prefer something that better fits my workflow: for git, it's Git Extensions; for HTTP, it's Chrome.
1) uncommitted stuff in workdir. Potentially can be lost, so commit often.
2) blobs in repo representing snapshots at commit time. Can never be lost.
3) symbolic references to the blobs. Can always recover from reflog.
4) tools to sync the above two things between repositories. (fetch and push)
5) tools to merge, diff and otherwise manipulate the changes between snapshots and files.
I'm confident that git will never lose my data, so long as I commit it. This makes experimentation stress-free.
Technically, you can lose data by explicitly deleting your refs, expiring the reflog, and running gc, but if you go that far you might as well rm -r .git
0) you forgot to explain git's index. Mercurial doesn't have the index, it works how you described git.
2) data (blobs, revsets, whatever) in the repo actually can never be lost, there is no automatic gc
3) no need for some different tool/viewer to view commits that don't have refs
Technically there are several ways to remove data from the repo, but it never ever happens automatically behind your back.
I think I've asked this before, but what exactly are mercurial's branches?
In git, they are a "physical" feature of the repository as it represents a set of lineages, not an actual repository object.
As such, any reference to a commit uniquely identifies a branch, so the concept of a "named line of development" is simply implemented as a reference that gets updated as you make more commits. When you "delete" a branch in git, it goes nowhere. Only its name is removed.
What sort of structure does mercurial use to represent its branches? I know they are not just an emergent thing like in git.
Two articles that might help if you really want to dig into it:
http://stevelosh.com/blog/2009/08/a-guide-to-branching-in-me...
https://bryan-murdock.blogspot.com/2013/06/git-branches-are-...
The first article also states that unnamed branches are useful for small, temporary diversions, and notes that git has to name branches, but I think that's somewhat misrepresenting git since you can throw away names as soon as they are no longer useful. To me it seems kind of silly to have unnamed branches, given that names are free and much easier to remember than commit hashes.
I think Linus' design goal was something that runs as quickly/efficiently as possible on large repositories.
> it is just a complex black box you chant arcane rituals at and hope it doesn't decide to burn your world down
My feeling too!
Git actually doesn’t scale to really huge monorepos.
IMHO the only real advantages over SVN for most users are the better branch/merge functions.
Which is a largish piece of software.
Have you tried looking into any other contemporary DVC systems?
I get Git, more or less. Having tried to make sense of Bazaar or Mercurial on several occasions (to understand the internal data model), I eventually gave up.
http://tom.preston-werner.com/2009/05/19/the-git-parable.htm...
To provide an alternate viewpoint, I have never had trouble with Git. I’m a bottom-up how-does-this-thing-work sort of person so when I first started using Git, I sought to understand how it worked. That part of Git is pretty easy to understand. Knowing that made its CLI a lot easier to grok. Of course, at the time I was having to use ClearCase at work and Subversion on the side so Git, IMO, was a vast improvement to either of those tools.
(dear everyone here and elsewhere recommending git incantations "but of course you have to know what you're doing": if you regularly have to take a backup of your working area before interacting with the vcs, because the interaction may do things you did not intend from which the simplest way back is to reset hard and start over, I humbly suggest that the vcs has failed in its primary purpose)
It's something subversion can't do last I checked, so whenever I need to do a complicated operation with SVN, I import the local state into a git repo first :P
So I guess 97% of users don't really get git.
I think this is really damning. Developers understand complex languages/compilers like C++, Python, Java, etc, all of which have a good deal more intrinsic complexity than a VCS. So if a VCS isn't understood, it is badly designed.
Your idea of a "user's mental model" might land you into trouble though, because all of us come from different backgrounds (subversion, SSafe, git, HG...) and they all maddeningly redefine terms in different ways (eg branch, forks, commits, checkout).
If I do do this, I will explicitly lay out the user's metal model in the documentation at the start. Then it will be the user's fault if they can't be bothered to read it.
> Then it will be the user's fault if they can't be bothered to read it.
can be applied to you about git technically. I think that's just me being a "little" pedantic though.
It sounds like the issue you actually have is that the documentation isn't easily readable in one or two sittings, and you don't have the time (or can't be bothered) to go through it and learn it. Which I totally understand, everyone has different things they need to spend time on, most of the time learning Git isn't one of them.
I have read large parts of the git docs. I don't like them. This is not just the text: the low contract colours and hard-to-read fonts are also factors.
> It sounds like the issue you actually have is that the documentation isn't easily readable in one or two sittings, and you don't have the time (or can't be bothered) to go through it and learn it.
So how long should it take to learn a VCS? And how long does git take to learn?
I guess I would be lost as well just using the command line.
I regularly use Giggle.
> Fork, Git Tower, GitX-dev
None of these work on Linux. GitX-dev also has a website I have difficulty reading.
It would have shaved off another 15-20 lines from the 503 line example ;-)
First, it allows the user to abbreviate flags. They can pass --fl and it will be interpreted as --flag, assuming no other flag shares the same prefix.
This sucks for maintainability: add a new flag and any abbreviation for a previously existing flag that shares the same prefix will now stop working, breaking user workflows.
Since Python 3.5 there's the allow_abbrev parameter that allows disabling this behaviour, but then you also lose the ability to combine multiple single-character flags (so you can't pass e.g. '-Ev' any more, and would have to pass '-E -v' instead[1].
The other issue is that it's tedious to keep all the .add_argument calls readable, while maintaining a reasonable maximum line length.
Click => command line interface creation kit.
Ironic little name...
It’s a magical idea if you haven’t seen it. You just write the help text and it automatically creates the argument parsing code.
import argparse
def add_file(file):
print('Added ' + file)
def main():
parser = argparse.ArgumentParser()
subparsers = parser.add_subparsers(title='Sub Commands')
# add parser
add_parser = subparsers.add_parser('add', help='Add a file')
add_parser.add_argument('file', help='File to add')
add_parser.set_defaults(func=lambda args: add_file(args.file))
args = parser.parse_args()
args.func(args)
[1] https://docs.python.org/3/library/argparse.html#sub-commandsThere are cases where it is not flexible enough but it is good for quick and dirty little scripts because:
* you use the docstring to generate the args. This way you always have a minimum usage that is valid
* Argument parsing is a bit more limited than doimg it yourself, but everything you usually need is there
* Docopt is exists not only for Python, but many other languages implemented it too.
https://looselytyped.com/blog/2014/08/31/gits-guts-part-i/
https://looselytyped.com/blog/2014/10/31/gits-guts-part-ii/
Disclaimer — This is my blog
Update - Fixed formatting / Clarified post