Torvalds: git vs CVS
marc.info
marc.info
(Assume you may be something other than a coder that carries all files in a laptop.)
1) Large binary or data files -- For this, there are various solutions of varying hackiness.
2) Your project has grown into several sub-projects -- The best solution here is to clone the repository to new names and then refactor and trim the respective project trees with a lot of deleting. You'll thank yourself later for improving your architecture and there will no longer be a partial check-out problem.
The real issue for both of these, is handling unversioned or independently versioned bits with respect to all the other bits. You might want some unversioned binary blobs, but you might also want some versioned giant csv files without having to download all the old versions all the time. Source code has one versioning strategy, data files another, photoshop files yet another. No version control system today really lets you pick which versioning strategy to use for which files. And none of the existing "sub-module" type extensions are really great either.
4) The project has grown into several sub-projects - but you don't control that repository.
"shallow repository A shallow repository has an incomplete history some of whose commits have parents cauterized away (in other words, git is told to pretend that these commits do not have the parents, even though they are recorded in the commit object). This is sometimes useful when you are interested only in the recent history of a project even though the real history recorded in the upstream is much larger. A shallow repository is created by giving the —depth option to git-clone(1), and its history can be later deepened with git-fetch(1)."
There are still issues; you can't commit for instance, but you can update to HEAD to update your patch, with a depth of 1 you only get the most recent changes.
Two of the features that made it difficult to 'sell' Hg for the main SCM of FreeBSD were the lack of Partial checkouts and Partial history
The problem with that, of course, is that merging only makes sense if you are merging branches that occur at the same tree-level in a project. This is why subversion merge tracking is so buggy and half-baked... because any given directory could be a project tree, a subtree, a branch, or a tag, and you could merge any of them. Sure if stick to certain conventions it works pretty well, but technically the whole implementation is a minefield, which is why svn will never be as robust as other systems that don't make this mistake.
Even though git's submodules and subtrees leave a lot to be desired, and have significant room for improvement, they will never be as convenient for partial checkouts, because the requirements for svn-style partial checkouts require crippling the entire system.
[1] http://doc.bazaar.canonical.com/latest/en/user-guide/filtere...
Suppose I sit down for a collaborative session with someone. Over the course of a few hours, we might generate 10-100 commits between the two of us, and several branches, all to implement a single feature. I.e., I commit, say "hey, pull from me, you had an off by one a few minutes ago."
At the end of the session, we only want to generate a single commit to the public repositories.
Mercurial queues are a tolerable way to handle this, but they aren't great.
http://stackoverflow.com/questions/598672/git-how-to-squash-... and
http://stackoverflow.com/questions/435646/how-do-i-combine-t...
[1] http://doc.bazaar.canonical.com/plugins/en/rebase-plugin.htm...
You can try to cronjob the updates, staggered around 4am. Or you can realize that the majority of people are only modifying a subset of these assets, and roll your own link-based asset manager for large binaries.
[1] http://news.ycombinator.com/item?id=1201559, http://news.ycombinator.com/item?id=1219082, http://news.ycombinator.com/item?id=1222905, http://news.ycombinator.com/item?id=1242374
You handle versioning either through filesystem snapshots, or by just using different paths for each version.
Of course not, it's not their problem. What annoys me is that the "You must use DVCS" memo is being sent around uncritically, and anyone who does not fit the use case is left out.
I think that people who deal with images/spreadsheets/data would also benefit from a good VCS tool. If the coders, having scratched their itch, did not declare the problem solved and move on ...
P.S. thanks for the polite answer. So, there is either no problem or no solution ... I think we're done here.
No version control system, centralized or distributed, actually handles blobs better than just treating them as artifacts and using rsync.
The only way to actually solve the problem is to make the data not be blobs anymore, where the diff/merge/serialization code all understands at least the structural container format and can render it usefully. There's never going to be a general purpose tool that does that outside of a live-in smalltalk image (or similar).
The best you're going to get is tools vertically-integrated with the application, and all the ones I've seen to date (MS Office, Adobe Version Cue, etc.) all do a terrible job of even doing diffs/RCS, much less actually implementing a real VCS.
It's interesting that the names are "Distributed Version Control Systems", or "Revision Control", but the ghost of the word "source" (meaning code text, in 1972's SCCS) is still read into it.
Surprisingly, when DVCS are shouted as the bees' knees for everybody, people who have images, documents and data may think they were included. Apparently not, those benighted heathens should use rsync or whatever manually, and not soil the "source control" systems with their binaries. Or go find not-included plug-ins, if they can.
(Those people are apparently too dumb to understand that a binary "artifact" produced with an image editor, spreadsheet, or data analysis tool is fundamentally different from a blessed code file written with a text editor, and thus is not entitled to be under real version control, or have a DCVS tool automatically do something smart about it.)
The Perforce guys must be laughing their asses off.
I don't think you'll find a satisfactory tool that has as its primary use the management of texts and deltas of texts. You'd probably need to use a combination of tools which has its own set of frustrations.
BAM solves this problem with a hybrid approach. BAM adds the concept of one or more BAM servers. Each BAM server contain all BAM data, but clones of the server (work spaces) contain only that data that the user requests. Typically this is the most recent version of each binary but in some cases it may even be less than that.
BAM has been successfully deployed in the game development space with great success. Game developers continue to enjoy the benefits of distributed work flow without the penalty of carrying all versions in every workspace.
By default all files are set read only. This was a big downside when you had no internet connection.
Of course, if you don't mark it writable (and just use :w! in Vim, for instance), then a 'p4 sync' may eat your work if anyone's changed the file.
I don't see it happening anytime soon as long as it justified using "its _source_ control". There are programming activities that require large binary files to be versioned along with source and this is a clearly a limitation of DVCSes IMO.
Apart from that I am quite happy moving to bzr from svn. Life is much easier now.
The hardware part seems limited: "the Alpha group ... hardware description language files into Vesta's source code control", so not schematics or layout binaries as 'source'.
That site seems to be resting since 2006, but there is some activity at http://sourceforge.net/projects/vesta/ (releases in 2009, mailing lists with 2010 messages).
http://lukepalmer.wordpress.com/2008/11/12/sketch-of-udon-ve...
git filter-branch --subdirectory-filter foodir -- --allI did it in the project I'm working on now! It originally had one main repository, and another for an extension that was used as a git submodule in the main one. Active development proceeded in both, with shitloads of commits in the main one just to update the reference to the submodule.
I decided this was retarded, so I took a clone of the extension repository, used filter-branch to rewrite it's entire history so that all the paths were always prefixed with "vendor/extensions/project_name/", and pushed it as an unrelated branch into the main repository.
Then in the main repository I made a new branch, removed the old submodule junk there, did a nice clean merge that melded all the commits from the extension branch going back in time, and made that the new master branch.
- Hard to explain to even fairly technical designers/copywriters.
- With all that easy branching, I sometimes just have a hard time getting a decent version of the codebase together from everyone.
I suspect Linus likes the approach not because it has no politics, but because it encourages the politics that match the way he'd like to manage the Linux kernel development, with this style of hierarchical, cascading commit approval, which is pretty hard to set up in CVS (though Mozilla's sort of grafted it on via their Bugzilla, which keeps track of cascading approvals of patches based on who owns which areas and sub-areas). Not that that's necessarily a bad thing, it's just different (and for a project as large as Linux, probably necessary).
I can count on one hand the number of projects described by that sentence. No surprise, the Linux kernel is one of them. Linus built a tool to satisfy his needs. But most developers work in smaller groups, and these groups have explicit trust. Working at a company, the trust network isn't dynamic. Even in large open source projects, most of the commits are by a handful of individuals. It's not a big deal if the occasional one-time contributor e-mails a patch.
But most of my gripes with git don't have to do with its ideas. Although it's distributed revision control, in my experience everyone designates one repo as authoritative. Like centralized RVCs, certain users are explicitly granted write access to said authoritative repo. So git ends up working like an svn repo with a ton of branches.
My complaints about git have to do with its interface. Coming from svn, git is very frustrating to use. Certain benign commands in svn will erase your data in git-land. For example, "git checkout filename" is the equivalent of "svn revert filename"; it erases any uncommitted changes. Of course, git has a revert command as well, but it doesn't behave like other RVCs. Git checkout can bite you if you have a branch with the same name as a file or directory in your source tree.
My biggest annoyance is if I accidentally commit and push something. Usually it's when I forget which branch I've checked out. Undoing a commit/push means rebasing or resetting, and that's where git drives me insane. I have used subversion, CVS, and even Visual SourceSafe, but only in git have I lost previous commits. Again with the misleading terminology. Why call them commits if you can destroy them with a single command?
A few months after that I tried hg and never looked back.
To look on the bright side first, Git is the only(+) DVCS in which you can clean up these kinds of mistakes without cluttering the revision history with "revert" and "oops sorry" changesets. :)
> I have used subversion, CVS, and even Visual SourceSafe, but only in git have I lost previous commits. Again with the misleading terminology. Why call them commits if you can destroy them with a single command?
The changesets are not really destroyed; you have just redeclared the official revision history not to include them, so they aren't visible. The changesets are still alive within the database until the garbage collector picks them up at a later date (then they will be destroyed), and you can access them via looking up the revision ID in the revision log (or sometimes just by scrolling up the terminal window). Once you have the revision ID, you can tag it (or declare it a new branch) so you don't lose it until you're done with the cleanup. Then you can cherry-pick, rebase, merge or do whatever you want to fix the erroneous commit.
This bit of Git causes a lot of headache among new Git users for the first few weeks (me included), since you need to understand the underlying database and really toy around with it a lot to understand how to do stuff like this properly. Though once I finally got it, Git replaced HG as my favourite VCS.
(+) That I know of at least. :) HG supports a "rollback" command, but that only covers one changeset, and you can't really use it if you have pushed your changeset or if it has been pulled by somebody else.
I'm sure he's no saint, but his most famous rants were responses to sniping. The Minix/Linux debate was started by Tanenbaum, Bram Cohen picked a fight over merge strategies in a list discussion, and the 'C++ sucks' rant was in response to some flamebait.
We settled with Mercurial, because hgweb isn't that memory intensive and our 512MB RAM prgmr VPS can handle it (although, we hope to upgrade the VPS, to allow more checkout's simultaneously). SVN/CVS also consume little ram on the server too though.
People who wish to make a selection should try them all out, and ask around. Because whilst Git users are very passionate about Git, I couldn't find a single one on IRC who had recently tried mercurial or Bazaar. Furthermore very few (if any) actively used Git in Windows
But that's just what I found. I didn't run proper benchmarks and things would be different if we had a better server (Loggerhead for bazaar wanted 2gb when running).
After clicking around for a couple of minutes the only thing I learned about it was that it has something to do with Songbird (which google tells me is a media player).
But Git requires that users either use cygwin or install half a linux environment in Windows.
That was true, but is no longer. There is a Windows port of git called MSysGit available on google code: http://code.google.com/p/msysgit/ .It's nowhere near self contained. Bzr,Hg,cvs,svn (and others), are just a small directory, and need none of those. MsysGit is overrated too. Usable yes, but ideal? Hardly..
Surprisingly, reordering did not affect much the size of the hg store, which is only 60% over doc size. That may be either because hg is being extremely smart about the content, or because OO doesn't move binary chunks around after they are inserted in the file (more likely).
On each commit, a script noted changeset number and output of 'ls -s' on the doc file and the .hg/.../_file.d store. Only started at 11, and trimmed most of the lines.
c.set file file.d
11 320 632
15 488 944
20 736 1336
25 1056 1804
26 2044 2800
29 2336 3260
30 2336 3480
31 2356 3560
35 2688 4104
38 2852 4388
39 2848 4468
40 2856 4540
41 2924 4676Before you downmod me to -∞ for my uncouth approach, consider this:
According to Linus, Git > tarballs > CVS > SVN (he made a statement about tarballs being better than CVS somewhere else). That leaves Git and tarballs.
Now, Visual Studio is my primary development environment, and it does not integrate with Git (as far as I know), and Git just isn't well-supported on Windows (there are some fragile solutions). Secondly, I probably spent a whole day playing with Git where it's supposed to shine (OS X with GitX), and I just find it kind of awkward and unintuitive for no benefit.
Yes, Git is not integrated with VS but I have no trouble switching to explorer and commit my changes using TortoiseGit. TortoiseGit scans my working directory and presents list of all the changes made in a session.
And yes, to be on safer side, I copy and store my working directory + Git to a different location.
http://stackoverflow.com/questions/1500400/is-tortoisegit-re...
It's not just the fragility and lack of integration, but also I'm just not seeing much of a benefit to counterbalance the complexity and awkwardness.
Also, if you could provide some real details on the 'complexity and awkwardness', I could share my experience which could be helpful.
Giving it a try can only prove the presence of fragility, not its absence :-)
Generally, people are very reluctant to criticize "hip" tools like e.g. Clojure, Haskell, Google Go, Git or its accessories. So, when 3 out of 4 people say they had problems with it, to me it weighs very heavily on the negative side.
However, that's no reason to dismiss all version control.
Also Subversion no longer stores it's data in a database so Linus's objection in this article has been resolved.
Finally, Linus's needs are pretty unique in the world. Linus isn't satisfied with Subversion for the same reasons it might work perfectly well for you.
He says SVN is better, but is more fragile (which for source control, I interpret as being worse):
SVN fixes (supposedly) those "implementation
suckiness" issues. ...
I think it's also a much more fragile setup and
there's apparently been people who lost their
entire database to corruption
Even if SVN = CVS, clearly Tarballs > SVN, according to him. His actual quote was Tarballs >> CVS. I can dig it up if you can't.If you find that personally offensive, so be it.
CVS's database is just RCS files, plus a little. Nice and easy to restore if Bad Things start to occur.
Try Mercurial. Git's piss-poor cross platform support has all but removed it from my development routine. Well, that and I find Mercurial's interface far more comfortable to work with.
There's a certain peace of mind to be had knowing you can roll back to what was working yesterday at 17h30.
(P.S. it's easier than the tarball snapshot - right click on folder, commit, type a note. Been down that path ;-)
After I went to all the trouble to explain that my system is the best one for me? :-)
What problem that I have will switching to subversion solve? Suppose I have two versions of a procedure and I can't decide whether the new version is faster and just as correct as the previous one. I keep both with
#if 0
// old one
#else
// new one
#endif
And it's easy and intuitive to see them side by side and switch back and forth between them until I'm sure. The VCS just don't give me this simplicity, convenience and intuitiveness.with git, there's a nice 'bisect' feature that lets you quickly jump back-and-forth between different versions of your code (in a binary-search-like way), so that you can debug performance issues like the one you're using #ifdefs to manually do. just check in a bunch of versions of your code and use 'git bisect' to jump between them and test each out for performance (or correctness)
> with git, there's a nice 'bisect'
How does "bisect" know where the boundaries are? What if you change the original code a bit, like re-indent it or make another trivial change? How can you look at both versions, preferably right in the editor? What happens to time stamps when you switch between the versions? Versions cached in the IDE? Directories? (You may be surprised that Git leaves them around from previous versions)
It's all doable, but not very intuitive. Why bring complexity where there is enough of it already?
Your work flow is limited to the method you've chosen not the other way around. To say it works for you sort of misses the point. Version control can free you to work in ways you can't yet imagine.
> How does "bisect" know where the boundaries are?
Clever algorithms.
> What if you change the original code a bit, like re-indent it or make another trivial change? How can you look at both versions, preferably right in the editor? '
The file gets flagged in your editor as having a conflict. Inside the file, any code parts that cannot be merged are included the file (both versions) and you pick which one you want (or edit the changes together manually). It's actually very easy, very intuitive, and works with your editor. In most cases, you won't have conflicts.
> Versions cached in the IDE? Directories?
I've never had a problem with versions cached in the IDE -- probably because almost everyone uses version control it's not something that usually goes wrong. With IDE integration, it's even better. Directories are handled pretty sanely in Subversion, at least.
You obviously take the snapshot before. No different from the more over-engineered approaches.
>> How does "bisect" know where the boundaries are?
> Clever algorithms.
Really?! You change two methods in a class, and "bisect" knows how to undo only one? No, you have to spoon-feed it, "staging" your changes (I did use GitX for a day). So it's not as simple as taking snapshots after all, is it?
Edit: typo, formatting
Sounds like you'd fit in just fine with ClearCase users.
> (You may be surprised that Git leaves them around from previous versions)
Git leaves directories around only if they contain files that are not checked in version control.
You already have the problem (backups, multiple versions of code) you're just doing it the hard manual way. You could say the same thing about Visual Studio over Notepad -- what problem does it really solve? One is just a superior way to work. You're using stone knives and bearskins.
Version control doesn't prevent you from using conditional compilation. You really wouldn't have to change your work flow at all. But you already make tarball backups -- taking the two clicks to commit your code is going to be much simpler. And if you ever screw something up, you can always go back to a working version.
If you get around to branching and merging the full power of version control reveals itself. I'm currently working on a development branch of my software while the production branch continues to get bug fixes. When I'm ready to deploy, I just merge all those changes together.
And you are saying I'm not doing it right?
I have a 1-line script that does that, automatically adding time stamps to the name of the tarball. How is checking in your version simpler?
> If you get around to branching and merging the full power of version control reveals itself.
This attitude is unproductive, as I witness from other programmers' experience: they branch and then they branch - stuff gets inconsistent, bugs get fixed in one version, but not another, and the same bugs that were fixed before, get magically released to public 2 versions later (REAL STORY).
My philosophy is MAKE BRANCHING HARD.
> This attitude is unproductive, as I witness from other programmers' experience:
You said your are a solo developer so you're telling me if you had the power to branch and merge you wouldn't be able to control yourself? You'd just branch and branch and never merge and make a mess of the whole thing? Even though you could do that right now just using the file system?
The whole point of tracking your changes in version control is to prevent the very thing that you describe. I simply wouldn't be able to function without that ability. Our stable production version is live and we have big changes in development (over 5 months now). Without version control the bugs fixed in production would likely never make it into the new version.
Tarballs let you do that as well.
> You said your are a solo developer
Right. Those people are working on their own code base (total disaster, much of it due to branching and following the VCS "methodology". I feel like slapping their lead whenever he mentions tagging or branching. That stuff ain't free! You only have one brain! /rant)
No, they don't. I can right-click on any file and view all the changes as a diff, I can see exactly what I changed and when, and revert that file back to any previous version. You can't do that with a tarball -- at least not without significantly more effort. It's just better all around.
> Those people are working on their own code base (total disaster, much of it due to branching and following the VCS "methodology".
You're not required to use any particular methodology. I've seen people screw up with every technology in existence -- that's hardly a reason to head back into the woods and live like a cave man.
There are very few things in the field of computing science that are universally agreed on. There are dozens of development methodologies, thousands of different programming languages, IDEs, etc. The closest thing we have in this business to consensus is the use of version control (even if the exact tool to use is still under debate).
You're simply mistaken to assume not using version control is superior in any way to using it. There's nothing wrong with your own methods of development and backup but that isn't version control and it isn't incompatible with it either.
The important question is, do you have unit tests for that script :)
You outline problem with branches that are long-lived, but the majority of branches created in a typical Git workflow are temporary, often lasting only a few days or less before they are absorbed into master.
What if you're working on an experimental feature that makes your program unstable while it's still being developed, and the feature spans multiple files? If you decide later on that the feature is useless, do you modify each of those files to remove all the #ifdefs? I make a temporary branch for this and delete the branch if I decide it's a failure.
Msysgit is stable and fully functional. You can even install it so that the git commands are available on the 'DOS' command line.
If you are working on windows, there is no GUI better than TortoiseSVN. Give TortoiseSVN a try, it won't replace the #ifdef's; but it will certainly be better that tarballs for solo development.
Any new VCS goes through these same rants -- not able to distinguish themselves on merit, they simply attack the dominant player. With Git, one learns quickly -- oh, it's all about the Python 'community' and their politics, I see.
When it ain't broke, don't fix it!
Some reasons I moved to bzr are: - repository backups are free due to the distributed nature - python scripting - fast local operations
So, yes, I don't really have a big issue with svn (or cvs) but given a choice I would go with bzr (or any other DVCS).