A beginner's guide to Git version control
developers.redhat.com
developers.redhat.com
But after looking at other source control options, I find it to be an absolute joy to use -- even for very complex tasks. The VC problem itself is where the complexity lies. Any tool that deals with collaborative working on document will present the same issues that Git does. Maybe worse.
Be thankful that you can use any tool you want to create plain-text diffs; that git performs operations quickly; that resets, undo's, etc are possible; that the precise history (both of the actual state and the steps taken to get there) is entirely legible; that each command performs a single, well-defined, and well documented atomic operation; that the tool is extensible, command-line script-able, usable locally, free; that it keeps the size of a repository small; that it is scalable across any number of contributors working simultaneously.
Some, all, or none of these may be true with other tools.
Anyone ever had to do a diff of a Microsoft Visio document by hand? Anyone had to manually type in the name of a document and its revision by hand into a web form? Anyone ever spent an afternoon working on a document, only to realize that someone else already made the changes you made but forgot to update the filaname which caused their changes not to be visible to you? Programmers are spoiled with the best tools in version control. People in other domains are doing this, without even knowing that they are doing this. They are making commits, merging, rebasing, etc.. without even having a word to describe it.
Git beautifully matches the problem of collaboration at a fundamental level.
The abstractions, the data structures... if not flawless then at least very, very good. (And that is enough to overcome any number of other issues.)
The complaints are the inconsistency of the interface. E.g. `git branch`
Instead of:
"Oh, email this to doc management. Depends on who you get. If you get Nancy, she will double check that your rev C filename matches with your rev D filename in the changes section. Matteo does it differently though, sometimes you do need to include redlines for minor AND major changes, even if the major version gets rev'd. Oh, that doc doesn't have any minor changes at all. I don't really know why, but you better not include minor changes because that will be flagged for sure, only on this document though, for every other one you should have C->D in the footer. Oh that one, the template never got updated so you actually have to type it in manually. Yeah, no idea why it does that.
It would be great if you got the change review back by Wednesday because regulatory has a preliminary product summative scheduled Monday. There's a confluence page for which doc manager should be CC'd for each document to get the review back in time."
"Just memorize these shell commands and type them to sync up. If you get errors, save your work elsewhere, delete the project, and download a a fresh copy."
But being the guy who sort of understood svnmerge.py was much worse.
I've helped people in my own open source project when they hit git roadblocks. They've never needed to save out their work and wipe their repo. There was always a solution within git and once they learned it, they really understood how version control worked much better.
git fetch origin main
git reset —-hard origin/main
Why does the first “origin main” not need a slash and the second does? I’m sure there’s a “perfectly reasonable” explanation for it but it screams UX inconsistency to me.
In before people tell me that it makes perfect sense because that’s how Git works. Well duh, because Git was designed that way, not that the problem that it’s encapsulating demands it
Surely “git branch” is “the abstractions”?
`git branch` is a CLI command.
I think it was the performance.
It was GitHub. GitHub changed it from "use one of several analogous products"[0] to network effects.
[0] Joel, founder of StackOverflow, had a GitHub like competitor that used either git or mercurial commands on the same repository. Just as one example of how similar they are.
IIRC git grep search took seconds compared to more than a minute with hg grep.
Because it's called Github, not Mercurialhub ;) (I bet if Github would have chosen Mercurial instead, the popularity would be reversed, at the time I switched to Github I was just looking for an alternative to SourceForge for hosting my open source stuff, but didn't care much about the actual version control system).
I see a lot of people attributing Git's "win" to GitHub, and that's likely the biggest nail in the coffin, but the GH devs chose it because it was already reasonably popular and it was popular because of Linus Torvalds.
Regardless of what the benchmarked real-world performance of Mercurial was or how well optimized it was, there is a class of developer that thinks all Python is slow and may never change their mind about it.
Python seemed a good decision: git had a two week headstart and Mercurial was faster to hit many usability milestones (and arguably git may never hit some of them, jk). Mercurial had good Windows support from early on (and doesn't need to ship like half of the GNU userland to do it).
I just don't think it should be surprising that some Linux kernel developers dismissed Mercurial off-hand just for being written in Python. I also don't think it is a coincidence that early GitHub ignored Mercurial just about as dismissively. (Ruby and Python aren't entirely "competitors" but there isn't always a lot of shared love between them.)
Whether it was the language or not, Mercurial was objectively slower than Git.
Nowadays that doesn't matter much, but 15 years ago, it was a much different story.
> Whether it was the language or not, Mercurial was objectively slower than Git.
That was not my experience, but I was on Windows at the time. Mercurial "objectively" had good Windows performance. (Not just because it ran at all, versus how much work it took to tame Frankenstein's Monster of shell scripts that was early git to run at all on Windows. But also because it's file system transaction model always fit Windows better.)
(ETA: Also, to be fair, my "team" at the time was darcs and my opinions on VCS performance at the time were from a very different perspective.)
It's especially vexing that mercurial throws a fit if you try to do anything like pull with changes in your repository while git does the sensible thing and just works.
One thing that perhaps let me understand git better later was how `hg pull` works actually like `git fetch` and makes more sense than `git pull`.
The team has eventually switched to git anyway.
If I could go back to Git I would in a heartbeat.
git-svn has a lot of good features that are specific to it, so it’s worth learning it like a new tool rather than a git feature. Get comfortable with it from some tutorials, then read the man page from top to bottom. It’s worth it.
The cheapo version is to git init and checkout of your source code, pretend it’s a git project, and then produce a patch for your branch at the end. You might even be able to use svn and git commands in the same directory? git for managing your private “work in progress” branches, locally, and svn to track upstream and submit work.
For me it was annoying and janky any conflicts were blocking me for ages. With git it’s a breeze and I would never ever want to go back.
Still, it is interesting to compare workflows and common gotchas.
All of those need you to know what you’re doing and what exactly you want to undo
* git branch hidden current
* git do-the-command-here
* if we call git undo, checkout the hidden branch, otherwise delete it
In practice, I think a lot of people call `git lg` often enough that they can just git reset --hard to a recent commit hash, or they manually create a just-in-case branch if they're going to do something risky/dangerous.
Microsoft Word has undo. Git can as well.
git reset —hard HEAD~1
It is already part of the CLI API.Or it could use the ref log (possibly augmented) as a literal undo log, and follow that.
> That would break a lot of the guarantees gained from the append only graph that git exploits for nearly everything it does.
The graph is really not append-only.
The commit graph is effectively append-only, but the named pointers into that graph can be reassigned at will, which means some commits can become (more or less) unreachable and eligible for garbage collection.
Sound to me like the commit graph is not append-only.
Yes, it is. You can't change a commit: you have to create a new one, and create copies of any child commits, which are also different commits. This generates a different graph. What changes are things like tags, branches and HEAD, but they're just pointers, not part of the graph itself.
> Yes, it is. You can't change a commit
You can remove commits from the graph. That’s not append only in my book.
In the same way you can’t rollback a db rollback, you can’t undo a git gc which cleaned up a bunch of unreachable objects.
edit: there's even prior art. Sapling has undo: https://sapling-scm.com/docs/overview/undo/ and is GPLv2 licensed. Just look at what Sapling does.
Technically nothing prevents the database automatically creating a subtransaction / savepoint for every operation you perform.
> This command would have to decide what to do for every previous command, no?
Yes? Not sure I see what the issue with that is.
And technically much of the information is already encoded in the reflog. The two big additional informations you’d need are:
1. Tracking changes to refs as they’re not necessarily tracked by the reflog, but there should only be a small number of plumbing operations manipulating refs so it’s not really a concern.
2. Tracking changes to the index, would likely be a lot more difficult.
The issue is that it sounds even more confusing in practice than status quo.
If you don't understand git, you can't rely on `git undo` cause you don't know what's undoable and what isn't. And if you do understand git, you don't need `git undo` in the first place, you can already undo things by yourself.
To undo something you gotta know what thing-that-you-did to undo in the first place.
This is often achieved using the command design pattern, but git does not record commands, nor has such a concept. It only has state.
Let's imagine it does record commands, what should a purported magical "undo" right after each one of these commands?
- git branch topic
- git branch -f topic cafebead
- git checkout topic
- git checkout cafebead
- git checkout -- file
- git commit # detached
- git commit # on branch
- git commit --amend
- git reset HEAD^
- git reset --hard HEAD^
- git add foo
- git rm foo
- git rm --cached bar
- git merge a b c # merge succeeds
- git merge a b c # merge fails
- git merge a b c # merge fails, resolve conflicts, add, continue
- git merge --abort
- git cherry-pick
- git cherry-pick --abort
- git rebase main
- git rebase --onto main cafebead^ topic
- git rebase -i # pick, reword, drop commits, one fails
- git rebase -i # pick, reword, drop commits, one fails, fix, add, continue
- git rebase --abort
- git stash
- git fetch
- git pull
You could come up with answers to that, but then, they would make sense only with one use case of what you intended to do with each command. To cater for that you'd start to need "git undo --this-way" or "--that-way". Any command that cannot be undone would break undo. Any external tool would need to operate using only commands, or it would break undo. Any external too or script extending git by chaining commands would need to implement new commands with undo for it not to break undo (otherwise "git undo" would only undo the last one). Also, consider that there may be files untracked in the tree and tracked in other branches, become tracked midway through a merge or rebase and whatnot. It becomes stupendously complex stupendously fast.If "git undo" were made for beginner users to make it more approachable, then it would fail as soon as the user faced a non-undoable situation and not understand why ("this undo is stupid!"). If "git undo" were made for advanced users to make it more convenient, then it would fail as soon as the user used advanced tools or scripts breaking undo.
What git chose to do is to be the simplest possible tool: it's a DAG made out of commits that can be labeled, and its commands operate on the DAG and labels (with a staging area so that commit creation is transactional). Commits are not removed until they're GC'd and "undoing" is moving back labels to be pointing to commits as they were before, so not having "git undo" is not even dangerous.
To make "git undo" robust, git would need to hide and/or change its core design + API + UI, which would make it not-git. I'd argue that part of the success of git is these "internals" (which are its actual interface) made it spectacularly easy to script, extend, compose, write alternative implementations of...
The only way to implement "git undo" is to embrace the database-like design (stage manipulation with git add and git rm are preparing a transaction that gets committed with git commit) and explicitly start (git start / git begin, records current work tree + refs + index state) and end (git end, a.k.a discard recorded state) or rollback (git undo / git rollback, restore recorded state).
BEGIN TRANSACTION git begin
INSERT INTO ... echo >> foo && git add foo && git commit
UPDATE ... WHERE git cherry-pick bar
DELETE ... (SELECT ...) git rm baz && git commit
COMMIT // or ROLLBACK git end # or git rollback
Now you gave me some ideas... hold my beer...I don't think so. Lots of software products that support "undo" support undoing across a greater many actions than git provides, and yet they are able to perform the undo just fine; they don't need a special "undo --paint" and "undo --move-object" and "undo --reset-to-default".
There's literally nothing special about what git does that can't be covered by a general "undo" command which looks at the last git command executed in the local repository and maps that specific intention to an action.
When this is not possible (in the middle of a rebase, for example), simply tell the user (as git currently does) that $ACTION is not possible while performing a rebase.
Having an undo command is likely the most leveraged thing by far the git developers can do to make new users unafraid of it. Screw up somehow? Just undo.
For example: a number of operations are not reversible in the current git model because they're not tracked. We'd need git to start tracking uncommitted changes so that it can restore them on `undo`. That's a fundamental change to how git works and I doubt that the idea would be entertained
No? If an operation is not reversible then it’s not undoable. That’s not abnormal. The opposite really.
Git is objectively pretty janky, though, so it seems to me this statement is an indictment of Subversion and CVS rather than some statement about the virtues of git.
I think it's much like the case with Wireguard. It's not that wireguard is some super amazing thing to deserve such adulation; it just sucks (a lot) less than what came before it.
I am talking about patched-together workflows developed at regulatory agencies, law offices and the like; collaborative editing of... non-code "information".
People who are doing "version control" but don't know that they are doing version control. They are editing documents in parallel. They are changing documents and recording the changes with a paper trail and diffs. They are working on separate parts of the same document in parallel, and managing the conflicts that appear in those documents when they are rev'd. They are sending their changes to someone else to have them reviewed before they are rev'd. Version control.
Their workflows and their tools are positively prehistoric compared to what is being discussed here.
Half the industry runs on excel spreadsheets. I feel your pain.
...
> Anyone ever had to do a diff of a Microsoft Visio document by hand? Anyone had to manually type in the name of a document and its revision by hand into a web form? Anyone ever spent an afternoon working on a document, only to realize that someone else already made the changes you made but forgot to update the filaname which caused their changes not to be visible to you? Programmers are spoiled with the best tools in version control. People in other domains are doing this, without even knowing that they are doing this. They are making commits, merging, rebasing, etc.. without even having a word to describe it.
I don't see what that last paragraph has to do with git being any good. You're looking at the difficulty of collaboration without revision control, and saying "this shows how good this particular revision control system is".
If you think git is good because of that last paragraph, then it's no better than CVS.
> Any tool that deals with collaborative working on document will present the same issues that Git does. Maybe worse.
I think mercurial is a lot better and more intuitive.
So does the recent "jj" appear to be: https://github.com/martinvonz/jj
I think maybe both fossil and bitkeeper are more intuitive too.
Did you try any of those?
Genuinely curious: what other VCS are you referencing here?
Personally, I have not used many, but of the total of three that I have used, Git is the hardest.
I've used SVN and I will start by saying that Git is objectively better. But that "better" is a balance between ease-of-use and features/power. SVN is much much nicer to use than Git in certain (admittedly narrow) contexts. To be more specific, the SVN (cli) user interface is a lot nicer to use than Git's. Merge conflict resolution is improved in Git, and decentralisation is pretty useful - the ability to setup a local repo and persist changes to disk immediately with no remote is essential in modern version control. These things make Git better but it is not easier to use.
Then of course there's Mercurial: somehow delivers on all the things Git has over SVN while still being almost as easy to use as SVN. No competition here.
The issue is git's interface is terrible and very powerful. Which means when something goes awry and they land outside those 5 or 6 commands they often have no idea how to fix it. Which invariably leads to a copy paste of their changes and a delete and re-clone.
Honestly I really like mercurial. I found its interface better but in this day and age all the tooling is built around git so ...
Yes, this. I have no problem adding, rm-ing, branching, making and accepting PRs, etc. These are what the overwhelming majority of git for dummies tutorials cover. (I've even paid for a very well-known and well-reviewed course and completed it.) And, these photos upthread are fantastic for that initial git-101 progression.
Then, I eventually landed my first tech-sector job and still haven't picked up much more since. My solution to every ounce of apprehension is deleting and re-cloning out of paralysis and fear.
I find great difficulty in being able to glean a repository's "environment" in a sense analogous to gleaning the environment of an initially compromised machine (the foothold) when you've popped a fresh shell on a boot2root box. (I thank the digital gods for IppSec and his hands-on videos that teach you how to clear up the fog of war around you!)
Having worked with older version control systems, I found that to be one of the best things about git. If I had a problem with a version control system like AccuRev, that sort of simple fix was not possible. Your every action resulted in a change of the server state, so if you got your workspace into a confusing state, the confusion would be synced to every machine.
These days I understand git well enough that I can fix just about any mistake, but the 'delete and start over' workflow was extremely valuable to me as a beginner. Knowing that there was a foolproof fallback of just deleting and recloning let me experiment without fear.
Git is basically that a programming language. The time you need to "master" git is comparable to "master" a programming language.
I'm being writing code for more than half of my lifetime so far, and every time I need to use an unusual git command I feel like a script kiddo who copies and pastes from SO.
“git is like a terrible programming language” doesn’t change the criticism one bit.
> Jujutsu is a Git-compatible DVCS. It combines features from Git (data model, speed), Mercurial (anonymous branching, simple CLI free from "the index", revsets, powerful history-rewriting), and Pijul/Darcs (first-class conflicts), with features not found in most of them (working-copy-as-a-commit, undo functionality, automatic rebase, safe replication via rsync, Dropbox, or distributed file system).
git diff > ../i-messed-up.patch
cd .. && rm -rf $REPO
git clone $REMOTE
cd $REPO && git apply ../i-messed-up.patchIt’s always been much easier to understand (and explain) what’s going on when you separate fetch from rebase/merge*, and I feel like all these “just re clone and start again” memes are all because people’s branch tracking broke and they wanted magical “git push # no further args” to work properly again.
Once you know how to “git push remote myref:theirref” you become much less dependent on magic. Knowing about / having to know about how it works internally is the fun / tedium of git.
*Just kidding about doing git merge, btw. Linear history for life!
The official docs are dogshit so I just don't have time to figure out the minutiae.
Pull means “do what fetch does, but with the spicy bonus risk that your precious work will be modified in an unexpected way depending on when and how you checked out the branch you’re on, what version of git you are using, and what config options you have set.”
In any case, I always teach people to commit more. As long as you commit, you can't lose anything.
However, I'm still not sure git will actually fuck things up with a pull. I think that's the users later when they go nuclear and delete the repo.
That's going to happen anyway when you need to rebase/merge. May as well type one command (pull), deal with the conflicts and continue, rather than type command, type another command, deal with the conflicts and continue.
Unless you have some other strategy that does not include rebasing or merging upstream changes, there's no advantage to not pulling. And, TBH, if you workflow is "this branch is worked on while specifically ignoring other changes by other people", then you have bigger problems than version control.
Of course there is: after fetching you can see if there are differences between upstream and local, you can inspect those difference, and you can decide how to reconcile the two.
“Pull” is a big hammer which unconditionally performs an integration operation, and the default integration is one you almost never want too.
There‘a no advantage to pulling IME, in the best car scenario it’s just an alias for “git fetch && git merge”, if that’s what you want you can just do that and create your own alias.
[FWIW, it seems you know more about this than I do, so don't think that I am purposefully trying to be annoying. I'm not, I hope :-)]
I agree that you can see the differences, but I'm asking how helpful is this.
For me, anyway (not an advanced git user) seeing the differences between master and my feature branch before doing the rebase makes no difference - I'm still going to do the rebase no matter what I see in the feature branch.
It is going to be rare for me to be able to see, of the 10 merges to master, if an of them are going to break the code in a way that I cannot continue (in which case I won't rebase).
The longer I put off rebasing, the harder it is going to be to do it, so I am highly motivated to rebase on whatever master has, even if it just got broken, because it will be more painful to rebase later.
Another situation I've run into is that sometimes I've made a small change that's stacked on top of a lot of other branches. And if those other branches get squashed and merged, regular rebasing can be really annoying - it is often easier to cherry-pick (I use rebase --onto, but same idea) your changes onto main/master instead.
That is true. But a few times, I've finished rebasing and regretted how I handled conflicts. And once the rebase is complete, you can't abort any more.
And I haven't put in the time to learn how to use reflog. Maybe this is my sign to do so.
git reflog <refname>
git reset --hard <refname>@{1}I mean, isn't that what a backup branch is? How are you going to get the hash of the original commit if you don't create a branch?
I suppose you could use git log or something to print the hash to the terminal, but that's just way more fragile than using git to keep track.
Checkout modifies local files to be like whatever you’re checking out.
And pull’s only purpose is to screw up your local files. :)
Therein lies the chief problem with git: it's a leaky abstraction. Nobody actually should give a flying shit about commit hashes and ^HEAD or whatever it's called.
All the newbie tutorials you find waste so much time on that, but what people really care about when they start using git is "what the fuck happened to my files and how do I get them back the way they were".
Git docs and tutorials are breathtakingly bad at showing this.
By default.
There's an option to make it fetch + rebase.
But anyway it's a complicated composite command.
I agree with the others. Don't use pull: use the underlying simpler commands directly.
The most powerful one most people don’t use is reset. Soft and hard resets are my bread and butter. I don’t even bother with interactive rebases for squashing. I do a soft reset against origin/<branch> and create a new commit.
To do it sanely you need to add the contributors remote, and fetch, and checkout, as usual. Would be happy to be educated here if I'm missing something.
- https://nitter.net/search?f=tweets&q=from:@NikkiSiapno+git&f...
- https://nitter.net/search?f=tweets&q=from:@ChrisStaud+git&f-...
For folks who enjoy visual guides, here are a few more I'm aware of:
- http://marklodato.github.io/visual-git-guide/index-en.html
After using git for well over a decade, I'm completely convinced that if you find yourself frequently rebasing/cherry-picking/reflogging you're using git wrong.
(... && Git checkout main && git pull --rebase && git checkout -B clean_branch && git apply ~/patch.out)
(I like the light rhyming of the first part)
Here's how I think everyone should use git:
1. Create a new branch for your changes 2. Make commits and merge from main with wild abandon 3. One final merge from main 4. Squash everything into a single commit, push a PR
If you keep your branch focused on only a single change, the end result is a tight, focused, single commit PR that merges cleanly into main and didn't involve any complex or error-prone shenanigans.
It's there when I need it, but I also work in such a way to never need it.
The opposite is true: resolving a conflict during a rebase is much easier, as you get to resolve the conflict in the context of a single commit and its parent. In some cases it may end up being more work than resolving a whole merge, but it's much easier to reason about.
I'm curious why you don't like it?
I always squash‡ before pushing a PR, so the end result is identical to a carefully rebased PR.
† occasionally branches will need to be split into separate commits, but that's not my default working style
‡ I know `squash` is a rebase under the hood, but it won't ever result in conflicts, so I'm happy to use it with every PR
When I hear about people not using rebase in their daily workflow, I imagine myself 10 years ago when I barely knew git and couldn't really use it as a helpful tool like I do today. It's almost like looking back to before I started using VCS in the first place - somehow I did manage to not use one for years (even collaborated via FTP!), but now it seems impossible. Usually most of the useful magic with git happens before anything gets pushed out, and `git rebase` has a huge part in it.
Reordering is pretty powerful. If you made a mistake, commit the fix, then move it to the commit where you introduced the mistake, and squash. Removing broken commits makes `bisect` nicer to use when you're desperate enough to use it.
Obviously don't do this on commits you've already published.
But for the rest? If you're working in a repo with more than 5 people, rebase, cherry-pick, and squash are necessary to keep your sanity. Merge nodes are awful once you get beyond more than maybe 3 developers.
What's supposed to be wrong with that?
If you just need to checkout the last branch you can also `git checkout -`
use it like
> git recent
2023-08-07: jimkubicek/add-journal-table-creator
2023-08-07: main
2023-07-28: backup/git-squash-to-main
2023-07-28: backup/cleanuprebase makes roll backs extremely easy if you need to roll back specific commits because of bugs and makes releases easier via cherry picking (so you don't slow down trunk merges just to do a release) and allow for fine grained continuous deployment that is harder to achieve than without it.
It is my experience however, that either everyone needs to rebase or you end up with issues eventually when only some developers are and other ones aren't.
I don't care as much for squashing myself as a general case, as you lose fine grained per commit rollback strategies though
The only time I merge is when I'm working on a shared remote branch. I haven't found a workflow (although I'm all ears if you have any suggestions).
1. write some code on a local branch
2. upstream has new revisions? rebase my branch on top
3. if not finished with my task yet, go to 1
4. if ready for review, open PR
5. if accepted, squash and merge
6. if changes are requested, write more code
7. upstream got more commits causing a conflict? don't rebase! it will screw up the PR history on GitHub and can cause issues for reviewers who might've checked out your branch locally and maybe done some experiments. merge upstream into your local branch. then you can push fast-forwardable commits.
8. push new commits to PR and go to 5
I used to think of rebasing as just rewriting commit history. But now I also think of it as altering the history of collaboration that is captured in a PR. So I switched from rebasing onto new upstream base branch commits and force pushing to PRs that already had reviews, to merging in new upstream base branch changes. I only do this after someone else has done anything on my PR; if I open it but nobody has reviewed yet, I'll do the rebase/forcepush to keep it current until someone does.
I prefer squashing to merge because I prefer the default branch to have one commit per unit of collaborative work. The way different people split up commits on a branch is arbitrary and varies widely; you'll never get more than 2 engineers to agree on a convention here. Keep all the messy stuff in the PR, and you can always revert one of those individual commits if you want finer-grained rollback. If you want a PR to have generated more than one commit, then it should be more than one PR.
I believe the only reason to do so is GitHub's lackluster PR UI. Force-pushing with an updated version of a branch after a review works reasonably well with GitLab's MRs.
> I prefer squashing to merge because I prefer the default branch to have one commit per unit of collaborative work.
There's no reason to squash when you can create merge commits from fast-forwardable state instead (again, one of the easily achievable options in GitLab's UI; GitHub doesn't make it easy AFAIK). This way you don't lose commit granularity while you can still obtain the "one commit per unit of work" view with simple `git log --first-parent` (or do the opposite and skip the merge commits with `git log --no-merges`).
The problem I run into is that other people have different workflows.
If they `git checkout remote/branch`, then everything's fine. But if they want to make a local copy of the branch, it'll get all messed up if I force-push. And I only want to adopt practices that are as robust as possible in the face of the possible ways other people could work.
The only thing that is sacred is the main/master branch. Everything is else just a speculative idea that, until reviewed and applied to master, is ephemeral.
(I’ve tried collaborating on a branch before but at the end of the process it’s hard to review because you either feel like the other party is rubber stamping their own code alongside yours, or you need to find a third party reviewer which spoils the 1:1 nature of almost every other code review I do.)
In my experience, that's part of the negative side of pair programming as well, although I really like it in general.
Sure there is. Less noise commits.
It's literally just a matter of a single command line argument to switch between views of whole MRs and individual commits, and both those views can be incredibly useful (especially during bisection).
It's also useful to skip noise if you happen to merge the upstream branch back into your topic branch for some reason.
Also, there's always `git log main..`, or even `git log main..topicbranch`. Combined with `--oneline` and perhaps `--graph`, `git log` is a really powerful tool to visualize repository state (and something that's incredibly lobotomized on popular Web frontends, unfortunately - I often end up cloning a repo to browse its history just because the Web UI is useless).
> This way you don't lose commit granularity while you can still obtain the "one commit per unit of work" view with simple `git log --first-parent` (or do the opposite and skip the merge commits with `git log --no-merges`).
I prefer not to squash before committing. I like having smaller commits in my history.
And if you don't squash, then using merge to fix conflicts with main/master before merging looks a lot more confusing.
> Who wants to spend time resolving meaningless conflicts?! Every time I try it, I instantly regret it.
Typically it doesn't take a lot of time unless you do something weird. Although to be fair to you, it does take more time than using merge.
A previous employer had a multi-tenant application that was deployed as a client-specific application which loaded the "core" as a dependency. They didn't really know how to do versioning and most version changes were just arbitrary "I feel like we should call it 1.8 now".
At one point I ended up maintaining a client-specific branch of the core dependency on version 1.10 (branch was 1.10-$CLIENT) while the "main" branch was 2.3 or something. For context, it went 1.10 to 2.0 because general cognitive dissonance.
This meant any change that needed to be made in the application core for this particular client also needed to be cherry-picked in some direction, usually by making the change on the client branch and cherry-picking it back as necessary. In some cases another client -- naturally, they would be on a separate branch like 2.3-$CLIENT -- also wanted that change so it needed to be cherry-picked again to that branch.
The result was a minimum of two PRs, one a cherry-pick of the others' commits (one commit unless I felt like spending my time in self-loathing), that I would make for every change. Not knocking cherry-pick at all; it's wonderfully useful when used correctly. That's just the result of non-technical decision-makers making decisions about technical tools.
On the plus side, I learned a ton about git in that job.
There’s absolutely nothing wrong with rebasing/squashing/amending/resetting heads on personal feature branches. In fact, it’s a pretty good practice if you make messy history and can make PRs less of an eyesore. I think the confusion comes up about when destructive history operations are appropriate because the git cli client does not have a concept of protected (shared) branches vs feature branches.
As long as you keep history destructive operations away from shared branches, you’re good.
At work we're maintaining a downstream Linux tree with a few hundred patches on top of mainline. The tree gets frequently rebased on top of new upstream releases, and some changes are being progressively upstreamed. It's much easier to reason about the remaining downstream changes and deal with conflicts when rebasing than when merging upstream releases back into the downstream tree. Of course you can't expect to be able to carelessly `git pull` in such workflow, but if you're working with people who actually know how to use git it's not really a big deal.
Naturally, this particular project uses a special workflow that fits its needs. It doesn't usually make sense to rewrite shared branches in projects where you're the upstream.
> If you’re doing collaborative trunk based development then you’re only cherry-picking.
All my work is collaborative trunk based development, and I never cherry-pick.
> There’s absolutely nothing wrong with rebasing/squashing/amending/resetting heads on personal feature branches.
I agree that there's nothing _wrong_ with it, just that it's unnecessary. If your branches are focused on a single feature and you're always squashing your PRs to main, the cleanliness of the branch while you're working on it is unimportant.
I'm personally not a fan of always squashing, for large features you lose a lot of history. I like a merge commit in some cases, you can still undo everything easily and most git commands support --first-parent so you can "pretend" everything was squashed in certain cases. But when you're got blaming, you have a lot more context to go off.
Instead I squash away the garbage and push out a reasonable looking chain of commits with nice descriptions.
If you break this rule you could be in for dealing with some atrocious merge conflicts though, so I try not to do it unless the branch I'm cherry picking from is a definite actual dead end (e.g. the change was an urgent hotfix against an old release branch and your workflow doesn't involve merging those back into main/master).
I will occasionally chery-pick something from master, do my work etc. Before making my PR, I'll rebase against master and potentially squash/reorganize my commits. When the PR eventually gets merged to master there aren't any problems.
I don't think I ever merge without rebasing though, so maybe rebase has been saving me from any potential problems.
A lot of people want to use git as a checkpoint/backup system, and commits and associated changes reflect that. The rebasing/cherry-picking/reflogging is one way to update the set of commits on the branch in order to make a set of meaningful commits for the feature branch they're working on.
https://opensource.com/tags/git
Unfortunately, the team who ran site got caught up in Red Hat's layoffs earlier this year and the site has been sitting in limbo ever since, so I don't know what will happen to it long term.
https://wiki.archiveteam.org/?title=ArchiveBot https://archive.fart.website/archivebot/viewer/domain/openso... https://opensource.com/article/23/6/new-developments-opensou...
I would say a bigger mistake is starting with the command line. A good GUI is absolutely instrumental to understanding Git, and it lets you avoid Git's terrible CLI for as long as possible.
Depends on the user.
If they are already an active user of the Terminal, they should be able to learn the git cli without ever touching a GUI.
The git cli has some warts for sure, and some weird inconsistencies. But with a bit of practice and some good documents about the correct mental model to have, you get used to it and you learn to use it very effectively.
1. the git commands map so cleanly to the states 2. there are so many terrible GUI interfaces that try and coddle the developer, really hiding the intent
I think the real problem is the flexibility allows for a lot of totally unintended but "legal" actions, from which it is really hard to recover because it's not your standard workflow.
I disagree. "checkout" does literally 2 unrelated things, one of them destructive with no safety checks. "reset" kind of moves HEAD around, but does a bunch of other stuff in the process. "Rebase" has a lot of magical (albeit useful) behavior involved in determining what exactly gets rebased. Etc.
> there are so many terrible GUI interfaces that try and coddle the developer, really hiding the intent
I agree. The CLI is confusing and occasionally obfuscates the data model, so adding yet another confusing obfuscating layer isn't going to really help. That said, a really good GUI that doesn't try to hide what's going on, would be a useful learning aid.
> the git commands map so cleanly to the states
They absolutely don't. Someone already mentioned the mess of "checkout" and "reset" but that's only half the issue. The naming is a huge issue with learning.
The worst is the "index" which is apparently named after the data structure used to store it. Such a bad name that it's often called the "staging area" instead. But even that is bad. Why isn't it just called the "draft commit"? That immediately tells you what it is.
Another example is reset's soft/mixed/hard. Terrible meaningless names. They might as well have been called 1, 2 and 3.
And there's more! It's not just the mapping and the naming. The actual CLI is stupidly inconsistent too. Flags have wildly different meanings depending on the command. You do the same action in totally different ways depending on insignificant differences (e.g. git reset --hard Vs git branch -f).
I have to look up the command to delete a remote branch every single time I use it because it's just so unintuitive.
I think people love the Git data model (which is great) sooo much that they think they love all of Git.
> It will be good if someone can suggest the git commands for that.
> I tried creating diffs from the tagged commits, but those are a bit messy too. I'll not try to fix more, unless someone can tell me the git commands that will make something better.
Consider that even the creator of Vim, a text editor so esoteric that the running joke is that no one even knows how to exit it at first, was repeatedly asking for help with Git commands. Let that sink in.
For me, ultimately, Git is best understood when there are visuals. The command line is cool, but if there's a tool that begs for a UI it's Git.
Yeah, I know there are Git UI tools, so many the article should have suggested some of those as well?
And, yes: it is good to learn the underlying fundamentals of the technology we are using. But, on the other hand, it denotes a rather poor abstraction from the UX point of view, imho.
Now, when I deliver git training, I start by explaining the DAG and how there is no magic, only git. By the way, the notes and exercises of the course are in my GitHub account[1], feel free to check it out if you think it can be useful.
[1] https://github.com/ciberado/git-workshop (https://ciberado.github.io/git-workshop/)
The problem appeared when people like I started basing our workflow on it ^_^. There are a few good books[1] and tutorials[2] out there, but I totally agree that the official documentation is only useful as a technical reference.
[1] Mastering Git, https://www.packtpub.com/product/mastering-git/9781783553754 [2] Atlassian Git tutorials, https://www.atlassian.com/git/tutorials
git was originally explicitly a VCS core for anyone to build on. Remember the "plumbing vs porcelain" days? I do.
And after 18 years what have the "adults" done in the meantime? Come at them sideways for having the gall to release an unrefined tool that got popular?
The core of git is cool. Few simple commands, easy to use. The remaining stuff is something only a mother could love.
Mistakes were made.
They haven't caught on because it turns out git's CLI isnt actually bad. It feels complex because it feels like what it is doing is simpler than is presented. But in fact it is solving quite a complicated distributed database program, transactionally, with editable history. The CLI hides a lot of that but cant hide it all. But you do need the flexibility.
Perforce is idiot proof. I can teach a non-programmer who has never even heard of revision control how to use Perforce in literally 10 minutes. They will never shoot themselves in the foot. They will never lose work. They will never, ever need to nuke and reclone their repo.
Perforce has other issues of course. But Git has both a bad CLI and a bad model. Maybe its particular model is strictly required for the Linux kernel. However for 99% of developers that are centralized on GitHub the model ranges from “mediocre fit” to “downright broken”.
> Perforce has other issues of course.
Claim: You can't solve the issues with perforce without making it no longer idiot proof.
I've used perforce. For distributed development, it's a total non-starter.
But “distributed” must mean something other than that, especially how git is presented (i.e. technically there is no center).
So, do folks doing distributed development routinely push changes in a peer to peer fashion? Alice, Bob, and Charlene are collaborating with Alice and Bob working on one feature while Alice and Charlene work on another, pushing incremental changes to each other but only sending the completed feature/branch to their non-collaborating peers when they’re complete.
Does that happen often or is it just the “commit early, commit often to the local copy” that distributed devs are really using? “I can edit on a plane” scenarios.
Yes. The prototypical example is the Linux kernel, which is what git was originally created for. There, there are a large number of different trees. Linus's tree is 'standard' Linux, but there are the various stable trees, there are trees for various subsystems, there's the continuous integration tree 'linux-next', and others. Those trees' changes are all intended to eventually reach Linus's tree. But there are other trees which aren't, they host patchsets that sit on top of Linux "proper" but aren't intended to ever be upstreamed.
>Does that happen often or is it just the “commit early, commit often to the local copy” that distributed devs are really using? “I can edit on a plane” scenarios.
In practice, not many open source projects are big enough and distributed enough that they need to do what Linux does. This aspect of it is very useful too: that you can code on a plane, that you can code in the bath, that you can code in a shack in the woods, etc.
> This aspect of it is very useful too: that you can code on a plane, that you can code in the bath, that you can code in a shack in the woods, etc.
Even turbo-centralized Perforce supports offline mode. Distributed systems enable offline mode, but offline mode does not require a hyper distributed system!
git simply has a handful of commands that require network access, none of them a part of "day to day" development work, nor necessary to fully utilize git if your canonical repository is on your own machine.
my biggest problem with perforce (it's been years since i used it) was that it had a "checkout" model where it was necessary to do something before starting any code changes that you might want to commit later. i found that quite problematic, and reminiscent in some way of older systems like cvs. git manages to retain the sense of "the codebase is just a bunch of files" all the time, and that works better for me.
Define distributed development. Do you mean like Linux with thousands of random contributors? Or do you mean a game team distributed across the globe? Or a AAAA team with big, scattered offices?
Perforce is not a good fit for Linux! It’s effectively the only game in town for almost all game devs.
> You can't solve the issues with perforce without making it no longer idiot proof.
I think I’d take that bet. Becoming idiot proof isn’t hard. The trick is for all commits to be automatically backed up in the central hub. And for commits to be locked and stable once made.
Git’s ability to re-write history is, imho, a huge mistake and I don’t think actually necessary to support Linux. Flattening on PR merge doesn’t require a rewrite.
That's the last thing I want. Random WIP commits being sent off to some central hub? Fuck that, man. Fuck that.
>And for commits to be locked and stable once made. Git’s ability to re-write history is, imho, a huge mistake and I don’t think actually necessary to support Linux. Flattening on PR merge doesn’t require a rewrite.
Being able to re-write history is absolutely necessary. I seriously doubt you've ever looked at a patch series posted for any free software project if you say that rewriting history isn't necessary.
To put it quite simply: my data is under my control. I can do whatever I want with it. I commit frequently because it is useful to be able to go back in history through changes as I make them. For the purpose of publication, it is not useful to see the various stages I went through when thinking about how to solve a problem. That's not what git history is for. It's for presenting a logical series of changes in a way that is easy to understand and bisect. Flattening on 'PR merge' is abysmal. I don't want one massive commit. I want a series of logical commits.
This is where you’re objectively wrong. I have this feature available to me today. It’s a killer feature. It’s amazing. Having it has zero downsides. Not having it is a pain in the ass and makes life worse.
Imagine this. You’re working at a company with thousands of engineers. Everyone is making stacks and stacks of local commits. At various points in time people push their commit(s) to code review. If approved it gets merged into master.
Now imagine if anyone could check out any commit from any employee just by typing “git checkout #####”. That’s it. That’s the feature. If you browse the repo it is perfectly clean. There’s no dirt or noise. This includes letting you checkout your own commit on one of your five different machines/platforms/cloud servers without having to push or pull or any of that shit. Commit on one machine and checkout on another. It’s pure automagic.
> Flattening on 'PR merge' is abysmal. I don't want one massive commit. I want a series of logical commits.
Sure fine. Shape the series of commits however you want. As few or as many as you want. With nice clean messages. The world is your oyster. But those are new commits. The initial commits should be, imho, unaltered (and unmerged). They can be GC’d months/years down the road if needed.
No, I don't think I will. No problem requires thousands of programmers.
You never need to do these things in git either. People only 'nuke and re-clone their repo' because they google something and get awful StackOverflow answers written by idiots that tell them to do that. It's not how you're meant to do things in git.
You're very unlikely to actually lose history in git unless you go out of your way to do so. I mean, you might not actually commit your changes, but I'd hardly call that 'losing work'. What's the alternative, autosaving into your history? No thanks. But once something's committed, it is hard to delete it. The reflog exists.
The philosophical difference I've found using both is that Perforce is file centric and git is commit centric. In perforce you have the file tree and then the history of each file, with git it's flipped. This is why I find it so hard in git to see how a particular file has evolved, with Perforce it's second nature. Perforce is so good at telling you why a particular line of code is there, I miss that so much in git.
I've even considered trying to code something similar to ease the pain!
https://mail.python.org/pipermail/python-dev/2014-May/134528...
I've never used it so I'm wondering what I'm "missing out" on!
I could explain Git's data model and what the operations do to my wife. I don't think I could actually teach her how to use the git command line though.
Not to worry, you'll probably still be alive for that.
* The staging area/"index" is a draft commit. You can add or remove things to the draft. When you run `git commit` it turns the draft into a real commit and the draft becomes empty again.
* Stashes are just "WIP" commits. They don't have a branch name pointing to them but you can list them all. When you apply or pop (apply & delete) a stash it takes the diff from the stash commit (versus its parent) and applies it to your working tree.
Submodules conceptually are relatively simple but the actual implementation is stupidly buggy and confusing. Like, I still have no idea what "submodule init" does. Why are there two places where submodules are stored - `.gitmodules` and a second secret place that you can't see? Why isn't `--recursive` the default? Why doesn't switching branch also switch the submodules?
Most of the answers to that are "well it could work properly and be simple, but actually it's half-arsed and full of bugs". You can seriously break your repo with `git switch --recursive` or if you try to use work trees with submodules.
And given the choices in things that get first class support in git, "which branch was this commit originally made in" or "no you don't need to care about that intermediate commit" or "what commits were derived from this commit" or "what branch is this commit in" seems like they would've been more useful problems for the freeping creaturitis to solve.
Have you used LFS? That's what happens when you try and use external scripts to hack in features. No thanks!
You absolutely need to know the right tools and techniques, of which there are many. For instance, nails and screws are not interchangeable, and they're made with a variety of metals and coatings for specific applications.
Just about anyone can hammer a nail into a piece of wood. It is quite difficult to put together a high quality finished piece of woodwork. Same is true for software.
I don't think Git is stupid or fundamentally wrong, just that the interface we're all using (even the GUI interfaces) are just thin veneers over the underlying API, which is confusing to many people, and that Git could be more productive for most developers when we have an abstraction that is easier to understand.
I've seen people much smarter than me fumbling around to solve problems with Git.
That's your problem right there.
You're a craftsman. You have three sets of inputs: raw materials, designs and tools.
You've chosen to focus on raw materials and designs, and have decided that mastery of the tools (or some subset of the tools) is not important.
A cabinet maker who said "my expertise is in complex corner joinery and standalone rectangular forms" would be laughed at if they also added "i find japanese saws and routers problematic".
Yes, tools are tricky. Tools are hard. But tools are an equal part of the task triangle, and you owe it yourself to build up your own scaffolding of understanding for the tools as you do for the other components.
I'm not arguing against source control, I'm arguing against a tool that so complicated and is obviously so hard to master might work against us as much as it work for us.
Nobody is pretending it is simple. It is simple.
No, it isn’t.
Wouldn't the grass be greener if a variable were stored as a unique key, making refactoring trivial? Wouldn't the birds sing louder if formatting were just a view on the underlying data? And wouldn't the sun shine really brightly if diffs were to operate directly on the abstract syntax tree?
I fear that this has to do with the great problem of interoperability, and of people not always wanting to work together. What would be a constructive way to coordinate ourselves out of this silliness?
The tool is great thou, and is my default diffing tool.
Basically, you need to sit down and build the greatest programming language ever conceived, complete with a world class ecosystem, and then convince people that this is truly a revolution software development, and you probably won't make a dime off it because proprietary programming languages are evil.
Good luck?
It's about "i want to be able to pick the tools i use", which in turn means "the underlying data we operate on needs to use a format/storage mechanism that any reasonable candidate for a tool can use"
And the answer to that quandry is: plain text files.
For example, 'git config --global user.name "username"' and 'git config --global user.email "useremail"' are required for Git users on any system before a commit is made, but since it follows the 'Git Configuration: Windows' title, it reads as if it's a Windows-specific configuration.
Additionally, $HOME/.gitconfig is also used by Linux (and UNIX and macOS) systems to store this configuration.
https://ardour.org/files/gitintro.pdf
If I say so myself, I think it's better than TFA from redhat.
> and then tell git to continue attempting to apply your changes.
Sorry, I just don’t know what that means. Redo step 4? Does rebase fail on the first conflict? It seems that when you get a rebase conflict your in the middle of a process but dumped to the command line. It’s not clear how you move forward (I guess repeat step 4) but more importantly how do you go back to the start before you did step 4? git rebase master
and get a conflict report. You fire up your preferred editor/IDE/whatever on the files that have conflicts, and you fix the conflicts, either by picking one of the versions (they will be marked sections delimited by <<<< and >>>>) or by manually merging the versions into a single new version. You save those files back to disk.Then you run
git add ... list of files with conflicts that are now fixed ...
and then you run git rebase continue
You may have to repeat this process multiple times before the rebase can complete.This advice will mislead beginners for whom the article is written.
Earlier than this post, my thought was that "What value can I add by writing about git? Theres already tons of articles for that."
Seeing this post from Redhat, motivated me to write whatever I feel like writing, even if it has been written before.
Thank You
Everything else can be googled/learned when needed. I think I just wish the default flags were more sane and command's name would be less "internal".
https://www.youtube.com/watch?v=hZS96dwKvt0
By far the best non-beginner git tutorial I've ever found.
https://github.com/martinvonz/jj
I've been using it myself lately and it's pretty cool. Takes some getting used to, but actually pretty easy to recover from screw ups, unlike git.
Most git clients handle this automatically since for most "normal" flows there is no need to separate the two.
An explanation on the need for separate staging would be nice to include the in the philosophy of git.
$ history 1000 | cut -c 8- | sort | uniq -c | sort -nr | head -10
238 gs
73 gd
68 gcz
67 ga .
56 npm run test
53 gp
29 code .
25 open .
24 git pull
20 ls