Git 2.0
lkml.iu.edu
lkml.iu.edu
Now that I look around, almost everyone in the programming world (at least the part I'm exposed to) is familiar with git as a version control software and github as a social-coding/code-hosting site.
But that aside.. massive props to the Git team on 2.0!
It was a very tearful month.
And git can still be used along with it - For some of the lib/tools that I do, I create a ".git" repository living along with p4, and work like this - then submit in P4.
We bring them into the git repo after running them through our asset pipeline, which means that there are really never any merge conflicts for them and every checkout is a playable version of the game with art.
This reduces the load on git by a lot, though we still tend to have .git folders measured in the tens of gigabytes, which is getting up there.
This isn't necessarily horrid for web apps as the images aren't big. But when you're talking about 500mb PSD files and massive massive 3d maya/max scenes it's horrid.
http://stackoverflow.com/questions/3055506/git-is-very-very-...
SVN and the like will attempt to store the binary deltas instead of duplicating the files. I don't believe git can actually do this due to the way branches / git itself works (it doesn't have a simple #1 - #2 - #3 ... #n commit chain setup)
There are plugins for handling binary assets better but they aren't great. and/or it just makes it less intuitive for the non-programmers which really defeats the whole purpose.
I absolutely _love_ git and I'm 99% sure this issue will be worked out at some point. I've seen teams just use git for the code and then use something else for art assets (alien brains or something).
Edit: The default is to delta files up to half a gigabyte in size.
A system like perforce stores the files read-only on disk, and makes them writeable when you check them out. To do that you visit the perforce GUI and try to check the file out; if somebody else has the file checked out already, the server will tell you who it is, so you can go and talk to them and decide how you're going to order the edits.
(Because of the read-only flag, even if you don't check out at least the files you know you'll need before you start work (which you should do), you at least find out immediately when you try to save. So even if you then end up losing a bit of work - it happens - you at least won't get too far into your work before finding this out.)
[1] Braid here if you haven't: https://www.youtube.com/watch?v=uqtSKkyJgFM
Why? At my current workplace, I converted a svn repository (with 6 years of commits) to git without any hiccups! I did not even lost a single commit!
I also detached a huge binary directory (they committed everything to svn...) and made a different repo for it.
Also you got git-svn, with you can seamlessly commit to an svn repo and still work in git on your local machine, nobody will notice, except you!
Take a little look at this: http://www.highprogrammer.com/alan/windev/sourcesafe.html
At least with NO version control, you know you're not protected. Bad version control, such as VSS, gives a false sense of security that you're protected when you're really not, which is worse.
I put ClearCase on an old version of my CV, because I was nominally the sysadmin in charge of it for six months. I still get calls about it. Those last ClearCase shops, they're getting desperate.
We ended up buying perforce which saved our arses but the days of VSS were dark days indeed.
At my first job (small shop with three devs) they were using VSS and it was so much of a diasaster that I took it upon myself to prepare a sizable document on why do we need to move off of VSS to something like SVN. Manager's response was something along the lines of "I don't trust those fancy merge algorithms, unlock-modify-lock is the way to go". I couldn't convince him otherwise, he wouldn't even agree for a itsy-bitsy pilot project to test the waters.
Having exhausted all my options, I did two things. First, I "Got Latest" of the entire repository. Second, I wrote a smallish script that would overwrite a single random byte in a random number of files inside the VSS database directory and scheduled it to run overnight. Of course it would self-destruct upon completion, so that no traces were left.
The next morning, VSS was on it's knees and nothing could be done. Turned out we had no backups.
And this is how we "migrated" to SVN.
1: Yes, I know that TFS supports git repositories. But I have the same issue: Convincing management that this might be a good idea.
2: Actually I consider TFS's source control part worse than VSS. While VSS was a hack, TFS is sold as 'enterprise' software. And it's unbelievably bad (again, only talking about the source control part and leaving out all the other things it can do. Maybe some are less half-assed than the source control part is?). Given the choice I'd pick CVS over TFS.
I think git is going to prove to be one of those pieces of software that is just transcendent.
In the past, version control was something that existed but you never had to think about. When you got to a good place, you'd hit the big red Save Button and up your code would go. Once in a blue moon you'd step on somebody's changes and have to do a quick merge. No big deal.
Now with Git, version control is part of the workflow. It's never far from your mind. You're creating branches before writing code (versus creating branches three times a year when you needed them). You're doing silly housekeeping tasks like reverting a change that you accidentally checked in to the main branch then branching and copy/pasting your changes back into the IDE just so that the source control system (and possibly some ops overlord somewhere) doesn't get mad at you.
I'm sure there are times where Git has actually made things easier, but most of the time it's just one extra thing taking up my attention.
It's like timesheets for your code. Not a big deal, but ever so slightly insulting that you have to deal with them.
I'm sure if you were a release manager, or the lead of a sizeable team, or the owner of a busy open source project, you would feel very differently about 'silly' modularised changesets and branching.
The masses have spoken and they want git.
(I'm sure there are other statistics out there, but your comment had made me curious and this was the only one I knew of)
You don't have to create a branch before you write some code. Write the code, and if you think it should go on a different branch, type "git checkout -b 'branchname'". But if you don't want to, you don't have to! work on master for all anyone cares. Branches just help you! IF they are too hard, don't use them.
Branches were so problematic in subversion that nobody did them. In git, if you want to, then do it.
If you accidently check in something, just "git rm" it. You have this extra staging area.
If you are finding there is something problematic with your workflow, take the time to look up how to do it, and you'll find that the fix is a command away. You should never be copying/pasting stuff back into your IDE. If you are, you're doing it wrong.
http://gitref.org/ is pretty good if you get stuck.
EDIT: Come on. If you disagree with me, don't downvote, reply so you can contribute to the conversation. Coward.
I've found my workflow in git has changed over the years... just like my coding style does.
But as someone just starting out, I can relate to how demanding git can _seem_ to a beginner. Say I have some refactoring ideas floating around in my head. I could either begin accommodating for these changes in git (namely branching, perhaps also chunking my edits into commits) or start coding and go back and handle commit isolation and branching later.
For me it can seem like git is getting in the way if I select the former method, and that the latter is just messy and cumbersome. But the more and more time I invest into the tool, both my efficiency and the cleanliness of the code base increases. Persevering through the higher learning curve (which exists probably due to a minimal (bad?) interface wrapping a set of concepts that can best be understood visually à la Dwarf Fortress) appears to be a worthwhile investment.
I think its the complexity of git that forces you to think about it. It's really not that user friendly, in comparison to SVN for example. Recently I had to teach a very senior DBA git so that he could check his SQL scripts into our repos, and I kept on thinking it would be much easier to teach someone SVN (actually he already had experience with that).
The big fallacy for me of distributed version control systems is that we all use them in a centralized way anyway (commit loses value, code does not exist at team level until pushed). It just adds another step to your work flow (code->commit->push vs. just code->commit), but the gain isn't obvious.
Another SVN fan unconvinced about git.
At this stage I'd say I'm very comfortable with git at the CLI (I only ever use source control tools in CLI mode), but it took me a long time to get there.
Create some aliases that make it as close as possible or as simple as possible for them. There's no need to force someone to learn the whole tool if they only need a subset.
Anyway, I think I get what you're saying wrt. less experienced users, but I found that with the proper explanation it wasn't actually that hard to get the point across. When explaining it to my coworkers, I've always tended towards focusing on the human factors. Namely things such as: "Don't you wish you could have taken that commit back?" (rebasing), "Everyone has a backup of the whole history"[1], "Hey look, if you mess up a merge, you can just say 'git reflog' and figure out where you were before everything got messed up", etc. If people can see the advantage in terms of their professional goals they are usually very willing to learn even relatively arcane tools such as git :).
[1] I know that's a white lie, but...
A screwdriver is much more simple to use than a drill, but if you have to choose one, which would you build a house with?
I think this statement highlights the fact that there are really two camps of developers when it comes to version control software. The difference between the two is what they consider to be the primary artefact that they are working on.
One camp considers the primary artefact to be the source tree that they have in front of them. The other camp thinks that the history of that tree over time, as represented in the version control system, is the most important and valuable thing.
Where things get rough is when members of the 'source tree' camp have version control imposed on them, usually by some corporate policy change. For them, the version control system is just like a backup. Commit messages are perceived as valueless as the developer never expects to see them again, so quite often they will be blank. Any change to the workflow that gets between the developer and making a change to their source file is a hindrance.
I would not recommend that an organisation with developers who are primarily 'source tree'ers try to adopt git. As the parent says, it requires too much knowledge of the tool, change to workflow, and, above all, commitment to having the git repository represent a detailed history of the development process. For them, a simpler tool the stays out of the way is probably more suitable.
I like a good code history, but not at the expense of good current code!
The only change to my workflow coming from a little bit of SVN and a whole ton of CVS was pull and push. Other than that, you can work exactly as you did in CVS and SVN if you want to. Sure the commands are a bit different, but the workflow can be the same.
We switched from a Subversion-like proprietary VCS to git about six months ago.
We needed to because the proprietary thing was a stupid waste of money (it was basically equivalent to Subversion in capabilities), and because its ability to do the sort of branching we found ourselves in need of was severely limited. So we switched to git.
We've had more than a little culture shock since then - and basically it's because we work in source-tree mode, and history is something you delve into only when you have to. In fact, when we moved to the new system (git on Github) we didn't even port history - we started the Github repos with a historyless copy of the source, and left the old stuff in the old system should anyone need to look at it. Github basically serves the role of the central repo in the old system.
We would never go back - git centred on Github is a vast improvement on the old system - but it's a bit too fully-featured for what we want from it, and has the problem you describe where we have to think too much about the tool's view of things.
Honestly, I'm not sure you have realized the true power of git branching.
Here is our typical workflow at work:
1. Branch master to topic/feature branch
2. Work on feature branch, push feature branch to remote
3. Push of feature branch to remote triggers a Tddium test run.
4. Deploy feature branch to a staging server for internal review.
5. Make fixes, adjustments to feature branch and push to remote, which triggers more Tddium test builds.
6. Once things are looking good, we pull --rebase master, which pulls down everyone elses changes to master.
7. Rebase the feature branch off of master, resolve any conflicts.
8. Merge feature branch into master and push.
Above, branching allowed me to:
* Save my own personal branches to a remote repo where they are backed up,
* Offload my test suite to a 3rd party,
* Allow me to deploy a feature branch to a staging server all without messing with the master branch at all.
I can also squash or rename commits on my branch easily and cherry-pick other commits from other branches without worrying about master. I can also VERY EASILY switch to other branches to work on other features and bugfixes as they arise. I normally work on 2-4 diff branches per day. Try doing any of that in SVN and be in for fun ride...
Offloading tests to 3rd party will not help you ensure that tests run fast as they should.
Working too long on your branch will only make mergers more painful later.
We use svn and git, but the workflow is still local fast tests before even branches commits. No personal branches, only features, and update those often because they're not exclusively your silo. And commit/merge to head often. And head is the only place where you get the 2h integration test run. Because otherwise, you'd wait 2h after every branch commit (which we know nobody can do)
Not sure what you mean. I state above how it helps my workflow.
> Offloading tests to 3rd party will not help you ensure that tests run fast as they should.
That's not the main point of CI. The main point is both: as a sanity check for tests pushed to the remote and also to not waste your time and CPUs running tests locally.
> Working too long on your branch will only make mergers more painful later.
I easily rebase off of master on a daily basis with git, no issues there.
what does it help you if you and other dev have their, often rebased or not, branches for a long time? the first one will be fine, second will have merge hell.
honestly, you are only comfortable with that setup because you probably work alone on a small team.
edit: and dont get me wrong, if that is the case, you are using the right tool for the job. i would love to be able to not rely on regular atomic commits in my huge dysfunctional group.
> rebase is not the same as avoiding merge problems.
Yes it is. The point is this: Everybody rebases against "what is to become our next release" on a continual (daily if not more frequent) basis. If you have acceptable levels of unit testing and integration testing on your feature branches there's usually no problem with early (that is, way before the sprint ends) merging of feature branches to master. If your team is really paranoid and all your team's feature branches tend to pile up until the last day of the sprint[1], you could forcibly (via process) stagger your feature merges so that they're always at least 1 day apart to catch any bad interactions between features.
[1] This is indicative of other problems to do with process/management, so take the next bit of advice with a grain of salt.
I have exactly this situation at hand at the moment. We are contributing as much as possible to an open source project. However as buisness goes, we need certain extensions to this project which clearly fall under our NDA, so I can't contribute them back. That's precisely where version-management becomes important, and where real benefits of git kick in.
In this case, I can do a bit of trickery by chaining repositories: I have a public repository for our contributions, which can open pull request to up-stream. Behind that, I have an internal repository, which has two read-only remote repositories (upstream, and our public contribution repository to fast-track certain changes). From there, you can do the changes in the right repository and need to obtain the changes through the right remote repositories.
The other thing that Git has achieved is that it is now the de facto standard for source control in a way that no other source control tool has ever been. An increasing number of ecosystems don't support anything else and most best practice guidelines that you find on the web assume that you're using it, or else tell you to switch if you aren't.
I suspect that it won't be long before not using Git will start to count against you in the jobs marketplace, whether you're a candidate or a recruiter.
Cool, I was just looking for something like this with "git merge". Turns out "git merge" already supports it, and I need to get better at using the reflog (@{...}).
> "git add <path>" is the same as "git add -A <path>" now.
> The "-q" option to "git diff-files", which does NOT mean "quiet", has been removed
More intuitive.
> The bitmap-index feature from JGit has been ported, which should significantly improve performance when serving objects from a repository that uses it.
Improves clone performance if you're pulling lots of history, but there still doesn't seem to be a way to sparse-checkout without fetching the entire .git repo.
I'm not entirely sure if this is what you want, but sparse-checkouting only parts of the repo is possible: http://jasonkarns.com/blog/subdirectory-checkouts-with-git-s...
For cloning a particular branch, there's also (since 1.7.10)
git clone -b your_branch_name --single-branch git://some.remote/repo.gitOne feature I really hope git to add is an easy way to clean up deleted files in the repository. Some times I accidentally check in some large zip files or built files and that really blows up the repository size. Those files stay in there even if I've deleted them. It's a pain to clean them up.
This was probably complicated by the fact I thought you were saying "the feature is there, go RTFM." This is not a regular feature of Git and I think architecturally it's the wrong idea. You want to delete objects, you better go learn rebase, and be prepared to piss off anyone who already has a clone since you're going to have to break fast-forwarding.
Here are the steps I know of to get rid of big blobs. Long and complicate. Is there a better way?
- Find out size of the big blobs and their HASH.
git verify-pack -v .git\objects\pack\pack-HASH.idx | sort -k3n
- Find out the name of a blob hash to verify the blob is the one to delete.
git ls-tree -r HEAD | grep HASH
git ls-tree -r COMMIT_HASH | grep HASH
git log --all --raw --no-abbrev | grep HASH
- Expire all working reflog now, and then prune deleted blob
git reflog expire --expire=all
git gc --prune=now
- Remove a blob.
git push -force
git filter-branch --index-filter 'git rm --cached *.zip --ignore-unmatch ' HEAD
git filter-branch --index-filter 'git rm --cached *.zip --ignore-unmatch ' --prune-empty -- --all
git filter-branch --index-filter 'git update-index --remove webapp.zip' <introduction-revision-sha1>..HEAD
git filter-branch --index-filter 'git update-index --remove webapp.zip' HASH..HEAD
- Clean up reflogs.
rm -Rf .git/refs/original
git reflog expire --expire=now --all
git gc --aggressive
git prune git filter-branch --tree-filter 'rm -f path/to/bigfile.zip'
is one command.It may be obvious that this solution is also going to cause the same issue as with rebasing, but just in case, from the man page for filter-branch:
WARNING! The rewritten history will have different object names for all the objects and will not converge with the original branch. You will not be able to easily push and distribute the rewritten branch on top of the original branch. Please do not use this command if you do not know the full implications, and avoid using it anyway, if a simple single commit would suffice to fix your problem. (See the "RECOVERING FROM UPSTREAM REBASE" section in git-rebase(1) for further information about rewriting published history.)
I even use pu on my work machine to be able to test and use the latest features and never had a serious problem.
Maybe I lack of the domain knowledge, but writing VCS must be a very difficult task if you care about preserving history.
You could see pushing as making a backup, but even that wouldn't backup the exact state of your repository (private branches, configuration, reflog, rerere, etc..).
If an object gets corrupt, it's almost impossible to fix it, unless you have a copy of that object somewhere.
That's why it's still important to make backups of your repository.
If they don't use source control at all, that's obviously a problem, but don't discount a company just because they don't use your favorite software.
Subversion (and Perforce and other centralised tools) will let you restrict who can check out what from a repository. That's useful if you want someone to only work on a small part of the code but cannot risk giving them access to all of it, for security or secrecy reasons.
And finally, as much as I love git and DVCS, for some teams Subversion works just fine. They may well have Continuous Deployment set up with strong CI tests on commit or they use Subversion's "reintegrate branch" functionality which is somewhat similar to rebase. Honestly, merging in Subversion isn't really that bad.
E.g. I've worked on systems that were ostensibly locked down like this. Except that for any developer who cared, getting physical access to any number of machines where the "secret" code would be sitting checked out, with suitable credentials to do stuff to it, was trivial.
I'm sure some people get it right. But I'd be very critical of that as a reason. Especially given that you could easily solve this in Git too, by putting the sensitive stuff in a separate repo.
Your point regarding binary assets is better, and if someone had raised that as an objection to git, it'd be worth listening. But that's a very different argument than the hypothetical argument posed above.
"SVN works for us" is also a very different argument.
Had those arguments been given, I'd be sympathetic. But the hypothetical "we need control and auditability" given above would be a red flag to me.
How about the removal of a sensitive file accidentally added to a git repo? It can't be done without rewriting the entire commit history after the file was committed and then forcing a push which will screw everyone up with an existing clone.
You could argue that the file should never have been committed, but then I'd argue you have never worked with humans.
Given that nothing stops you from running a centralized Git repo for things that are candidates for audits or release. If someone uses that as their reason for not switching to Git, it indicates to me that they don't understand their requirements and/or is using it as an excuse because they don't have any real reasons.
I'd even argue that if you want more control and auditing, Git is an improvement over alternatives like Subversion, because it makes it more manageable (and so more likely to actually happen) to set up workflows with far more frequent commits and subsequent reviews by a gatekeeper. You don't need to let your lower level developers even have any kind of access to the main repository, yet you can retain full history of changes that happens before their code is approved and pushed to the main repo by a suitably anointed person.
You sound pretty close-minded and relatively inexperienced if you would turn down a great career because they used something other than git. It's borderline idiotic to make such a large decision based on something inconsequential. Maybe that's okay in the Web app industry where most of the jobs are just variants of the same thing so you have plenty to filter through, but the idea of turning down a job at somewhere you've been hoping to work like NASA or Google just because they don't use a tool you like is one of the dumbest things I've heard in a long time.
I really would not get your hopes up on how well things are run. In your career YOU WILL meet people who refuse to try new stuff because "this is the way we've always done it" and unreasonable reasons.
I've met people who refuse to move on from visual source safe for example.
The majority of your first years of your career will be learning how to navigate office politics, how to introduce change, how to get things done without getting fired, how to manipulate your boss to make the right decisions without imposing yourself on him. Those are more important than raw technical skill. Most of the value of 'experience' comes from those skills.
Thankfully, being cocky is somewhat tolerated even expected from overeager graduates. However, that eagerness turn into zealotry and, later on, obsolescence/cluelessness for a significant part of them. You know those guys, since they cause you the most problems at work. They built their reputation by embracing new stuff and now prevent everybody to move on.