What does SVN do better than git?
programmers.stackexchange.com
programmers.stackexchange.com
If you're bundling in all of your dependencies, or some images/models/sound files then Git gets very big, very quickly.
In The Witness (http://the-witness.net/news) we have 20GB of data checked into svn. Try that with git.
On the other hand, Perforce is the game industry standard. As much as engineers complain how terrible P4 is in comparison to distributed options, the other teams working on the title (artists, designers, QA/Test) save a huge amount of time by simply not fucking things up for everyone else.
When there's 250 people working on a title, producers will call to lock the entire depot down so that Intern Joe Blow from Test/QA can't check in a broken build and waste the time of everyone who is working on the project.
I've administered depots (along with overseas proxies) with a 1TB head revision. It was amazingly fast despite how huge the data set was. The only exception was syncing a new workspace, which is easy enough to work around with weekly snapshots and rsync.
My last company did just that. The repository wound up being 30gb at the end of the project (which was a 3d facebook game). We tried to host on a company hosted github, but got kicked off because it couldn't deal with a repository that size[1]. Luckily, we only had one person try to check out the repository at the end of the project. It took them 14 hours for the initial check out.
[1] It was never made clear to me whether it was our hardware that had the problem or the github self hosted software. Either way, it was unreasonably slow.
But when we have dozens of multi-MB binary files, which change every week or so, then after a year we're ending up with a massive repository to pull down.
http://mercurial.selenic.com/wiki/LargefilesExtension
[edit: this isn't quite a solution, more of a work around -- but one that can be expected to be installed by default.]
Fortunately that difficulty has a huge payoff that SVN simply can't beat. Git has a superset of SVN's functionality - it can act as a central repository, and can also support any granularity of permissions you wish.
And you can pull in, pull out subtrees using Git using the git subtree module.
What are the advantages that git has over SVN in merging? I've never had any problems merging while using SVN.
git checkout -- <file>
This will revert the file back to the the last checked out version. git checkout <sha-1> -- <file>
Will revert a file to the state it was in, in that particular sha-1.1) set up a centralized repo and tell people to use it,
2) use fetch + rebase instead of merge to have a linear history, so managers are happy
3) when you have non-technical people, you can just tell them a standard set of commands to use (ok the hard part is to convince them to the command line). You may set up all the repo (remotes, aliases etc.) for them and tell them exactly the workflow.
4) regarding the locking, IMO it's good to have a number of integrators pushing to master (after the code review) instead of everyone in the team, that way locking shouldn't be really necessary.
The real things for me are:
1) checking out subfolders in SVN - indeed might be convenient in big code bases.
2) empty directories
Not sure about handling blobs, I didn't have big ones yet.
Anyway, I don't know what would make me crave using a non-distributed VCS now after nearly 2 years with Git.
1. One corrupt repository.
2. The Git tooling for windows is horrible.
3. Stupid people setting up their email address and username wrong then pushing it. We have AD integrated authentication with SVN so you can't screw that up.
Some of you may say that we may not be doing things right but the only reason SVN is preferred over Git is because of externals. It is just easier to reference an external project and keep it updated all the time. This is important for consistency point of view and we want to spend more time developing features verses merging and resolving merge conflicts, etc. When we tag things we simply svn copy and then flatten the externals making it all the same. It works great for us and we cannot see any reason why we should move away from this.
Git is interesting and definitely very powerful. We love the idea of working on a feature and then commit push and merge it when it is ready. It does make a lot of sense to do this over and over again but there is sooooooo much to type all the timeeeeee. It makes it so inconvenient.
Maybe it is just our model of work. Maybe we are weird and old school.
Of course, this feature exactly is not possible in a DVCS (it is almost the whole point of a DVCS). I certainly appreciate other features of the DVCS and would probably take them on balance. But do still miss the simplicity of it.
It's impossible to have global revision numbers for hopefully obvious reasons.
For many types of users, I think that is a thing that svn does better than git, especially since git makes it so easy to destroy/rewrite history.
git commit --amend
will add any currently staged changes to your last commit, and allow you to change the commit message.In the words of the manpage,
git rebase origin/master
will "forward-port local commits to the updated upstream head." In other words, if you have local commits, it throws them away and replaces them with different commits. This allows you to avoid making merge commits.Interactive mode rebase, for example:
git rebase -i HEAD~4
will let you rewrite the last commits. You can use it for lots of things like: changing commit messages, changing commit contents, add new commits in the middle of a sequence, squash multiple commits into one. git filter-branch
can be used to apply a script to all the commits in a branch. Situations where it's useful: "I don't like the username/email I was using, it didn't matter when it was just a local repo for me only, but when I publish it I want to change it", or "Oh crap for 2 years we've had files in our repo we don't own the copyright for, we need to get it out of our history before somebody sues us!"FYI, when people talk about "destroying/rewriting history," they're usually referring to the above commands that take a bunch of commits and replace them with different commits. I wouldn't call push --force, reset --hard and gc commands that destroy history; instead I think of them as commands that help you work with different views of history. E.g. doing git rebase -i followed by git push -f, the rebase is what actually rewrites history while the git push -f merely propogates the rewritten history to a remote repo.
Though this distinction is one of the things brought up in that email, that hg treats it all as one history. That is kind of nice.
Also, you don't need to do the gc; git will run it automatically for you, eventually.
We do in fact need the superiority in branching/merging that you have with git - not having it has been a PITA several times recently. We'd really outgrown CVCS and needed what DVCS had to offer.
That said, we've been having a fun time getting used to git's little ways. Everyone's shot themselves in the foot at least once, me included. If you move from centralised to distributed VCS, you will shoot yourself in the foot unless you have an accurate mental model of what's happening.
A few developers have complained it's all too hard and complicated and they have to use the command line now (they don't) and why can't we go back to the old way. I've been reminding them programmers are supposed to be smart and able to learn things ...
The change is totally worth it, though. The pace of development has accelerated.
tl;dr you do in fact need a basic understanding of the chain of versions and how branching actually works and stuff.
If you have a super important piece of code that you want to gate access to then you're going to deal with this roughly in the same way under both systems. Probably by breaking it into a separate repo and only granting access to specific people since even under SVN everyone gets a copy of the code on their local machine.
A lot of bad press on SVN branching comes from the TortoiseSVN tooling on windows which TBH can be a bit crappy at times.
Not storing directories is definitely a political decision, so by default I oppose it (Apple makes these kinds of unilateral decisions all the time, to me it's the ugly side of programming). I'm hopeful that perhaps the subtree issue has something to do with all of the hashes being nested and maybe there is a workaround (perhaps git needs placeholders when the files aren't present?)
I also didn't realize that there is no revision number in git, but since we can use the hash string I'm ok with it. Conceptually they are both just keys anyway. The chronological order of the hashes is encoded in their structure, so hopefully git makes it easy to retrieve the list.
It's a technical decision. Git doesn't store files or directories. It stores blobs of data identified by arbitrary textural identifiers. These identifiers tend to be paths, but they need not be, for example for a smalltalk-like system. There are advantages to this representation, for example when tracking changes across refactors that merge/split code files.
Does anyone know if git stores the path separator as part of the path? For example does it store a path reference as a string "/var/log" or as a list {"var", "log"} internally? That gives us 2 options:
1) If git DOESN'T store the path separator, then that means there's no way to represent a directory as "var/", to signify that it's a directory (because a file could contain the path separator character). If that's the case, then perhaps I was wrong to call it a political decision. An argument can be made for purity here.
2) If git DOES store the path separator inside a path string, then a mistake was made, because there was a chance there to distinguish between a directory "var/" and a file "var". Sometimes I refer to these types of mistakes as political decisions, for example if it's easier to make the mistake than work through a full implementation. I can't imagine that git works this way, but I'm throwing it out there because I have seen stranger things.
So without some new rationale, I have to come to the conclusion that ignoring empty directories was a political decision. Perhaps I missed something in how the hash is computed, for example if the node's name wasn't part of the hash, then empty directories would have no effect on it, which would mean the same hash could represent different directory structures, which would be incorrect. But that can't be right, because git can store empty files.
This all means that git can't represent an arbitrary directory structure. That's a big deal, because it can't be used to say, incrementally back up a hard drive without help. That's all well and good if its main use is version control, but I'll never look at it quite the same way again.