Confessions of a Git Skeptic
blog.artlogic.com
blog.artlogic.com
git has cut the team in 2. The graphic designers no longer know how to commit their work, nor do the frontenders, nor do most of the project managers. In particular, when a conflict comes up, nobody knows how to resolve the conflict (with Subversion we had a simple rule: "If you get a conflict, save your work outside of the project, then "accept theirs full", then re-integrate your work. A pain, but non-technical people could manage it). Instead, all of those people now send their work to the programmers, via email or Dropbox, and they say "Please integrate this into the project". This increases the amount of work that the programmers have to do.
We now use git for everything, because it seems like it is the thing to do, and it helps the programmers merge their work together. But the price is high: the loss of half the team, who can no longer use the version control system.
> We now use git for everything, because it seems like it is the thing to do,
You nailed it. And IMO this current obsession with Git/Github is bordering on unhealthy. I really wish/hope this fad dies out fast!
Also anecdotal, but the graphic designers at our company have no problem using Git.
So much easier conceptually, the cmd line utility was written for actual humans, and it has sequential commit numbers.
If you see references to commits 143, 732, and 2021, you have some instant grasp on the ordering of those. acbe2f, 7624ab, bccc07 not so much.
Git commit ids are a hash of the commit content. This has very nice properties in that:
- Ids are duplicate if and only if the commit is identical (the collision chance is so low it's not worth considering)
- It's content-addressable
And you can refer to it by any unique prefix, so it's not a pain to type out (and you should only rarely be typing out particular commits instead of branches or tags).
That right there illustrates the disconnect that causes some people to dislike git while others love it. Git is truly distributed.
In a distributed system, commits are not totally ordered. It is fundamentally impossible to give them all meaningful sequential numbers, because some are really neither newer nor older than others. They are "in parallel".
Git excels at truly distributed workflows, and distributed workflows dominate in open source. That's why git dominates the open source world. The basic issue in open source is this: code is shared across organizational boundaries. You need to make and deploy a change today, without knowing or necessarily even caring whether the upstream project is going to take the patch. But your code needs to stay consistent and cleanly mergeable at every point through this process, even as multiple patches are flying in multiple directions.
Git keeps getting better, and I'd be hesitant to put anything new on Mercurial at this point. The tooling just doesn't seem to be keeping up any more.
The pain here is not because git is inherently harder, but because you'd gotten used to svn (including your conflict workaround) and I guess nobody wants to do the same learning for git.
Recently I tried Tourtoise git and I came to the conclusion that a beginner will probably rather learn the cmd line version because he will otherwise loose code.
SVN is unbelievably more difficult. The entire conceptual process of merging/rebasing/reintegrating is so much more difficult in SVN than git that I really cannot understand how people cope. Branches in SVN are such a pain in the ass that feature branching becomes a nightmare barely even worth doing. I get migraines trying to keep straight what things in SVN make changes on the server and what things make changes locally only that I can sit on for a while. Being centralized seems to raise the stakes of just about every operation, whereas with git I can take it slow, get everything right, and not involve other people until I decide it is the right time to do so.
A month-long stint of working with another team that uses SVN for some incomprehensible reason left me seriously considering seeing a doctor about some blood pressure medicine. I swear to the gods I would rather use quilt.
For git on the command line, I use day to day 4 commands aliased:
alias status="git status"
alias commit="git add --all .;git commit -a -m"
alias pull="git pull --rebase"
alias push="git push"
So, typically status - find out what's going on
pull - grab any changes from others and replay mine on top
commit - add my files and changes
push - at the end of the day send changes up to the server
Occasionally I'll use stash or tag, which are also pretty simple and straightforward. Yes the command line UI could be improved, but it's not much harder than svn or cvs.This in particular I find curious:
We keep things under control by maintaining detailed sample scripts for everyone to refer to.
Why not set up a file with aliases in it if you are asking people to refer to cheat sheets? The basic commands as above are painfully simple, most people don't even need to branch and can get by just by pulling, committing and pushing, but branches aren't particularly complex in git either...
For non-programmer clients, for any version control system I'd get them to use a GUI like GitX, and they'll never touch the command line - that's very similar to svn GUIs, and frankly I doubt they'd know the difference.
It's high time this sort of version control was integrated into operating systems though - life would be a lot simpler if the Mac OS Finder for example had built in support for a common system like git with a user friendly front-end. I can't count the number of clients and colleagues who have made their own crude version control system by renaming files with version nos or initials to track changes - there has to be a better way.
You can do this in git as well.
The eureka comes when you stop using git by retrofitting Subversion workflow on it. The "Git forces me to learn obscure things" part precisely shows that.
The hardest part is not learning "obscure" git things, it's unlearning the habits and preconceived ideas about how SCMs work.
The terminology used by Git makes things hard for people to understand, remote and local, branches here there and everywhere. The OP is right that the deletion syntax for branches is awful.
BUT
It's not magic. It's explained all very easily.
"The staging area lets you batch up your changes so you can issue multiple commands to add files, or get rid of them if you change your mind."
"Just use `git pull`. If you don't know why you need to `fetch`, don't do it. Git pull is just like an SVN checkout."
"When you commit, you don't have to send it to other people, so your work can be in a broken state, and you won't break it for everyone else. When you're super happy with it, use `push` and it'll be visible to everyone."
These explanations get people who understand SVN to get Git, and they understand exactly what Git is providing over SVN. Once I use these explanations, people never seem to have problems.
What seems to keep happening is they look on the Internet for solutions to issues, which tangles them up even further, with rebasing and such. Even the simpler tutorials like learn.github.com are obsessed with showing pictures of trees and different branching states that beginners simply shouldn't be exposed to. Their mental model is completely screwed up, and they start to blame Git (not unreasonably) rather than themselves. When SVN screw ups happen, they blame themselves because SVN is conceptually simple, so they know the mistake must have happened with them. Then I hear things like "This never happened with SVN" or "SVN was always so much easier". When I prod them about this, they do end up remembering that SVN had it's own issues, but they forget them when their current problem is Git.
EDIT: Oh, and the other problem is that some of the GUIs out there are criminally poor. Particularly eGit for Eclipse. It exposes a metric crapton of functionality that I have absolutely no idea about, and I've been using Git happily for two or three years now. I don't know about refs or anything, and eGit makes it all exposed and seriously painful to do anything. Having tried to help someone with it, I gave up. "I use the command-line. You think I use it because I am clever, but I use it because I am stupid. The command-line is much simpler than this car crash, and I have no idea how you have got as far as you have with it."
I've begun to think of engineer's ability to "get Git" as a sort of abstract-thinking FizzBuzz. DVCS breaks up some composite actions (svn commit/checkout) into their component parts (git fetch/merge/commit/push). If you can't make that mental leap and see the value in it (and accompanying tradeoffs) despite all the learning materials out there, how can you be expected to derive comparable evolutions of design and engineer on your own?
- git rebase -i (the "cleaning" of history)
- git add -p (which requires a separate staging area)
- Using a bare repository (which, just for the record, I've also had to do with SVN, but we didn't have a special Googlable term for it)
- Track the difference between a local and remote branch (this is actually inherent to DVCS workflow, it's like going from DOS to UNIX and complaining that they make you track file permissions... you didn't even have the option on DOS!)
It's actually kind of embarrassing to see an SVN user claim that they don't understand the difference between a downstream concept of a central repository and the remote concept of the same repository, because I can sure as hell remember having to run "svn up" before "svn ci" just in case someone had made a commit to the remote repo in between the time of my last update and my upcoming commit.
master, "origin/master", and origin's own master all have corresponding concepts in SVN, even if master and origin/master had to be the same by definition.
One more thing: I'm pretty sure SVN now has only one .svn directory, just like git.
That was a more recent change in 1.7.x ( around late 2011?) and I know people who are still, in 2013, in 1.6.x
Having a local repo and push/pull to any other repo is not a complication but a huge advantage. It lets you do things like commit while you are on an airplane, or share committed work in an internal team before pushing to trunk, or keep repo history even if the server dies in a fire. It lets you not have to worry about who can be trusted with commit permissions or not. Being able to commit all the time without pushing lets you keep revision history without checking in, so you don't end up with many copies of your working directories, or saving up changes before you are 'allowed' to commit.
Since you can commit all the time, you can destroy the working tree trying things out and then bring it back if they didn't work.
So let's not complain about having a working tree, local repo, or remote repo. The staging area makes it so that you don't have to complete all staging of many different files in one huge command. You can add/rm something, preview what's staged, and only commit when you are satisfied.
In addition, CVS type merge is a pain while it is actually much simpler with git or hg. And being able to keep rebasing your private feature branch on a changing remote branch is a huge improvement.
The "extra steps" (the index, remote branches, etc.) are there for a reason, and that reason is to give you the tools to be careful with your codebase.
I always use `git fetch`, for example, because unless I've been doing something reeeeally weird, `git fetch` will never touch my local work. At all.
The index allows me to make sure I know exactly what I'm committing, and the extra step gives me the chance to look over it once again before I commit. (`git add -p` and `git commit -v` help a ton here, even though it feels like they "slow things down"). I never have "Oops, I didn't mean to commit that" moments anymore.
And the merging model (yes, with those opaque commit-hashes and mutable history and 3 kinds of branches) allows me to fix things when I screw up (or one of my coworkers does). Which still happens all the time, because we're human.
I think you could get a lot more out of git if you were willing to invest the time to learn more about its branching model (which really is elegant once you understand it), and to configure it properly. Otherwise, you may have a look at Mercurial, which has a more svn-like experience.
Git's command line is a UX train-wreck though. Poor consistency, learnability and leakage about how the internals work. I think Git gets a lot of unnecessary flak because even though people realize it's superior, they are always fighting it into doing what they want.
My pet project is giving it a saner command line.
1. In my experience, it's more robust. I had Hg throw Exceptions doing menial tasks.
2. It's more popular, there's more infra-structure available. GitHub is a big factor until we get more competition on this area.
3. Since most commands are just scripts that call lower-level tools, it's potentially more flexible. It's possible to derive new functionality without hacking too much the VCS source. That could include a saner command line.
Is this "pet project" still in the planning stages, or is it published on Github? If the latter, can we have a link to it?
> Git allows the topic-branch development style, where you maintain one short-lived branch for every task you work on. Like many people, I learned about this style from distributed version control. I love it. It’s an excellent way to switch between tasks, and maps well to the way I actually think about the state of the project. But you don’t need Git to do this! Subversion has “svn switch”. It does have an annoying problem that if you specify the wrong URL to switch to, it will trash your local copy rather badly. However, Subversion seems to have fixed that problem in version 1.7.
Quite an annoying problem. I'm not really sure why the author is a git skeptic if it enables and introduced him or her to a wonderful workflow.
The ability to easily work on multiple branches and make offline commits is a huge selling point for any DVCS. I really like git rebase and git add -p now as direct arguments for git, but wouldn't have used them right away.
Edit: added clarification after quote
Perhaps because you can do the same thing in SVN?
At least, that's what he says. I don't really know SVN.
This is largely due to the fact that I have a strong aesthetic preference for languages where you can hold most of the language [1] in your head. Such languages are called compact [2].
[1] Compactness of language and standard library are different. Having a non-compact standard library isn't necessarily a bad thing, especially if it means I don't have to complicate my life with third-party libraries.
[2] http://www.catb.org/esr/writings/taoup/html/ch04s02.html
And, until (apparently) recently, if you make a mistake SVN will "trash your local copy rather badly" which isn't quite the same thing.
My work desktop is Red Hat, so I'm (unfortunately) used to downloading old versions of software...
But underneath, the concepts are pretty simple, and because it's the right simple concepts, the things you can do with it are pretty flexible, and thus powerful.
Git's interface can defn be way better for the default case. But everyone and every team's workflow is different, so the inherent flexibility will always involve a little bit more complexity.
And oh, it seems that he hasn't yet discovered selective commits or stashing, so staging seems like an extraneous extra step.
And rewriting history seems frivolous if you work by yourself. But it makes pull-requests on open source projects way easier to manage if your commits are clean. And if you need to rollback to a working version of your production code, it helps to have clear commits too.
This is a BIG reason why I recommend even to advanced Gitsters to use a GUI tool like Tower or SourceTree so they don't get confused or inconsistent about how their repo is structured. I've used Git for years and I can't keep it all in my head. Surely some can but not me and probably not most.
His explanation was on the lines that if there are 100 developers working on different trees of freebsd and they want to commit their work, it causes a race, as if some one commits before you, you need to do a git pull, incorporate changes and then push, even when they are working on unrelated components.
My response to this was yes, and that is why git allows you to create branches so cheaply, but he still was not convinced.
My thoughts in this entire learning process was along the lines of - "Why the hell are they doing it like this, instead of that? Why are they calling this, like that, instead of just calling it what it already is?"
In my experience trying to learn GIT, I was able to conclude that GIT introduces a LOT of terminologies, and concepts, making the path for beginners like me more and more difficult. Whether spending hours learning these concepts just to get your stuff synced properly depends on your scenario. But for me, SVN rules.
Also, there is another REALLY brilliant way by which SVN helps me. I use Tortoise SVN. I have a couple of external 1tb hard-drives to back-up my code. I like to take regular backups (obviously) because some of the code is in production and I just want to play it safe. Using SVN, on windows, I run an SVN server on my machine, and create respective repo's on my hard-drives and just update them and voila - I have all the code properly backed-up and versioned. You can also automate this (there are scripts to do them), but for me, I like to keep it synced manually. And also, you can keep your stuff backed-up to your Google drive by making a repository inside your drive folder. I can't imagine typing all these as commands in GIT.
This makes some things easier. If the 'server' dies, you can substitute any other copy of the repository for the original. So if github disappears I can still continue to hack away on my code like usual (commit, branch, merge, etc), and if I decide to publish on gitorious or something later, then I just push my local repository up and everything is there (full history, etc). This sort of obviates the need for 'backups', since your copy is already a backup, but making extra backups of any repository is as trivial as copying it (rsync) or cloning it to somewhere else (cd /backup; git clone /myrepo) - so I don't need to run a server or anything to make backups.
As a thought experiment to illustrate the practical difference, consider why pull requests are so trivial to implement on github, and how would they work on the hypothetical svnhub?
So does Subversion, but you have already learned its terminology and concepts and are no longer a beginner.
preemptive disclaimer: I haven't used bzr for years and understand that it's faster now than it was. I liked bzr. blah blah I really don't want to start a flame war/pissing contest. But people should try git out before dismissing its speed as a worthwhile advantage.
_The_ git repository is right there on your computer.
The OP says he is waiting for the git epiphany. The epiphany is that there is no _the_ remote repository. You can use git like it's subversion, and designate one repo as _the_ repo, but because git isn't designed this way it isn't going to help you much in this endeavour.
Subversion got rid of those directories in version 1.7 (http://blogs.wandisco.com/2011/10/11/top-new-features-in-sub...)
Also, for the searching problem: betterthangrep.com. The reason I have perl installed on the Windows box at work.
I appreciate Git and hate going back to SVN when clients demand, but I'm convinced it takes a certain kind of brain to feel comfortable with it.
"You can use “pull” to mean “fetch and merge”, but some authorities will advise you not to. You can also use “commit -a” to avoid the “add”, but you still have to be aware of it."
Whatever tutorial I originally used to learn Git made use of `git pull` and `git commit -a` exclusively, without any elaboration. It wasn't until long into my Git career that I was even aware of `git fetch` or `git add`. To this day I use `git add` only occasionally (but it's so useful when I do need to use it), and I honestly can't think of a case where I've needed `git fetch`.
I'm about to teach my co-workers to use Git for a project I'm spearheading, and I intend to expose them to this simplified view and let them fill in the blanks when needed.
If you do an add, followed by a "git diff --cached", you can view what's going to be committed before you do the actual commit.
I've found this workflow to be useful, since it helps see things like debugging print statements that crept in, or unnecessary files accidentally added ("git diff --cached --stat" is useful for the latter).
This is the biggest issue I've always had with Git. I've struggled to understand why so many developers swear by a system that takes several steps backwards from other systems.
git add is never needed if you always git commit -a
if you don't care what git fetch && git merge does, just use git pull
The biggest problem with learning git is that all git experts love to talk about the advanced workflows, but all the beginners just want to be able to push and pull.
while git has submodules, they no way near have the flexibility of externals. if git would incorporate the idea of externals, I think it would cure quite a few people's complaints.