Like, I've had to explain a lot of times why you `git pull origin master` but when you want to interact with that remote branch otherwise it's `origin/master` instead. The lack of clarity is in what commands operate on what levels, with many of them operating on several at once.
There have been some efforts to reform the command set to be more clear, like `git switch`, but the old commands will persist forever along with a lot of other footguns (like `git push --force` really ought to be replaced with `git push --force-with-lease` and moved to `git push --force-I-really-mean-it` so it hardly matters.
`git fetch` (and by extension `git pull` when given a remote) and `git push` copy data to and from a remote. When you specify `git pull origin master` you're saying "pull down a copy of the remote ref master from origin", which it then saves locally as the ref `origin/master`.
Everything under `origin/` (or really `refs/heads/origin/`) is just a cached pointer to the last known state of that ref on the remote.
All other commands operate only on these local references. So when you want to refer to what you know to be the state of things on `origin`, you can use `origin/master`. Otherwise that command has no particular knowledge of how to talk to origin.
Incidentally this is a shortcut I use all the time to update my local master from a remote:
`git fetch origin master:master`
Which is super unclear in its meaning but it means fetch origin's master HEAD and put it in my local master ref. I actually use this more often than git pull nowadays.
Pushing merges is great. Pushing random (unreviewed) local commits directly to master is bad, but it's no worse when those commits are merges than when they're not. Conversely, rebasing master (which is quite easy to do if you're inexperienced but have been advised to use git pull --rebase) and pushing that creates a self-perpetuating mess that is very hard to fix (because even if you fix what you did, any other user who did a rebase-pull of master in the meantime is going to reintroduce the problem). Using rebase also trains you to force-push which makes messing up published branches much easier.
It is worth the time to fully understand refspecs. Once people do, they tend to understand all essential ramifications of branch and repository naming.
- origin master <=== the actual remote version of the master branch
- origin/master <=== a local branch that you cached from the "origin master" remote, may or may not be in sync with the real "origin master"
`git pull remote_repository_name branch_name` is the generic way to look at it instead of some magic incantation.
I like to call origin "upstream" to differentiate them.
and then git pull is another way to think of git fetch and git merge as one command roughly.
Comparing against that locally cached ref is also what git uses to tell you how far behind/ahead of the upstream you are in `git status` or whatever. Fetch and push are the only git commands that actually talk to a remote (at the "user level" of the command set anyways, those are also composed of lower level commands).
If it's difficult to keep your mental model of some system up to date, I doubt that doing bigger steps at once makes things easier.
So
1. run `git fetch`
2. if the textual output does not tell you what has happened, run `gitk -all`
3. Decide what to do. Rebase, merge, whatever.
Of course if you know exactly what you are doing, pull can be fine. If you changed the repo yourself on another computer that is the case. Otherwise, how can you know your second step, before having even seen the data you are operating on? Well, it can work, but if it doesn't, don't complain.
I agree. For a DVCS like git, separating the network transaction from updating the working copy on disk is the best way to go about it. Going in the other direction, this is the default since git add, git commit and git push are executed separately.
Essentially the whole concept of "upstream" is weird and non-orthogonal. Another one that bothers me is that as far as I can see there's no way to globally turn off setting an upstream on newly created branches (I can pass a flag to the specific "git branch" command, but that's tedious and error-prone).
This is the crux for me. Command naming is completely unrelated to and unindicative of state.
It feels like surely there's an opportunity for the basic CRUD operations to be collapsed down into a standard "{action} {source} {target}" style.
There will be nuances, specifically around branching, but the basics should be basic. As opposed to a Swiss Army knife, where you have to pull out the scissors and squeeze them three times before you can unfold and use the blade.
As part of a security-related project some years ago, my team and I hacked jgit to use SHA256, which required changing the length of pretty much every on-disk data structure. Sadly, there was (probably still is) no HASH_LEN constant, just a lot of magic offsets strewn throughout the code. I had to compare lengths against the git spec at every step.
And yet I still scramble for stackoverflow every time something goes slightly amiss.
It was actually a pretty cool system. I don't think it was ever sold though.
Wow, that description feels spot on.
jGit is actually a separate project from core Git, but once it gets adopted into core Git we can expect that jGit will follow suite, given that it's critical to Gerrit and other projects.
[1] https://lore.kernel.org/git/20191223011306.GF163225@camp.cru...
All it does is add this simple check before actually pushing:
if (remote_ref("blah") != local_ref("remote/blah"))
fail();
Most of the time it doesn't matter, and for most people's uses of --force it would have no effect (because most people are just pushing to a branch they're the only one pushing to). But every now and then it helps a lot to avoid losing data.I try to be pragmatic about this sort of thing yet push —force is one of those cultural no-no’s for me.
git init --submodule --recursive
Or is it
git submodule --init --recursive?
God I hate this UX so much I usually have a ./fetch-subrepos.sh that runs a bunch of "git clone" commands.
And if I push without first pulling, must it always punish me with a merge commit? Can't I say "oh shit I don't want to do this, go back and git pull"?
This is a source of probably 50% of my "ah, fuck, time to undo..." moments with git, these days. I hate that shit. Muscle-memory gets ahead of me and I commit on a shared remote branch, which would be fine given our workflow except that I didn't pull first. What a pain in the ass.
I would guess there's an easy way to make git do this automatically for you via config so you never forget, but I just never, ever `git pull`
Or:
> git config --global alias.up '!git fetch && git rebase --autostash FETCH_HEAD'
From:
You probably also want:
git config --global rebase.autostash true
[pull]
ff = only
If it does fail I can decide whether to merge or rebase.I agree with having as many commits as possible be compilable, but that's not the sole criterion, because there's a tension between that and having granular history: if you squash the whole history of the repo into a single commit then that means 100% of commits are compilable, but it's still a bad move. Conversely, a non-compiling commit in between two compiling commits is not a big problem (you just make sure your git bisect script skips non-compiling commits) - what really matters is keeping the diff between two successive compiling commits as small as possible. IME the best way to achieve that is never rewriting history.
Git submodule runs commands on submodules.
What is hard about this UX?
And it's not punishing you, its doing what you asked, to pull into a non matching head, how does it know you're not using git in the intended and distributed way?
Btw, just quit the editor without saving, it aborts.
What they're saying is: if you quit that editor without saving the commit message, git will abort.
AFAIK whatever's opened does need to block the CLI, so you can't use a command that opens a GUI editor then returns immediately or git will interpret that as your having closed the file without saving, but otherwise any editor should work, CLI or GUI, and can be assigned in your git config.
Most people want a single source of truth workflow that corresponds to the old total ordering imposed by svn or p4.
On the contrary, you want those commits for bisection, which is the main reason to have a VCS history at all.
> Most people want a single source of truth workflow that corresponds to the old total ordering imposed by svn or p4.
People think they want that, but I've never seen a convincing case for why. Bisect works better if you use merge. Blame works better if you use merge. And if you really want to see the history without merges (why?), it's one flag to do that.
I think I know git well, but you got me confused. I've never heard of pushes causing merges. Surely you are talking about pulls, right?
The only theory that makes sense is that this person doesn't know how to `pull --rebase`, but the order of `push` vs `pull` wouldn't change the presence of merge commits, so I'm still confused.
If I pull from origin before making my changes, I don’t have to merge, obviously.
But correct me if I’m wrong: I think that if I don’t pull first, but my changes don’t conflict with any part of what was done by the previous commit(s) I missed, I’ll still have to merge if I touched a file they touched.
This is a common scenario for me. Correct some typos in comments for example, and I get forced to figure out how to merge using vim, which I don’t know how to use at all (being a nano user). I’m sure I could and should switch to at least using nano by default, but I don’t know how merging really works, either.
What I really want to do is undo my commit, pull, and redo my commit. Then I don’t have to figure out git merge.
I do understand the problem being discussed; what I don't understand is what it has to do with pushing first. You have the same problem no matter which order you use `git push` vs `git pull`.
> I think that if I don’t pull first, but my changes don’t conflict with any part of what was done by the previous commit(s) I missed, I’ll still have to merge if I touched a file they touched.
Yes, that's true.
> What I really want to do is undo my commit, pull, and redo my commit. Then I don’t have to figure out git merge.
You can do that with `git pull --rebase`, which, as others have mentioned, you can set as the default behavior of `git pull` like this:
1. After making a PR, there are conflicts when merging into main. In a merge-based workflow, I would merge main into the feature branch, resolve any conflicts, then push. In a rebase-based workflow, I rebase the branch onto main, resolve any conflicts, but now I need to push --force. As some of the other comments have mentioned, this can be improved with --force-with-lease, but still isn't the greatest.
2. After making a PR, there are some typos that need to be fixed. Fix these in an interactive rebase, to edit the same commit that introduced the typos. Also requires either --force or --force-with-lease.
3. When the PR is accepted, the result is rebased on top of main. My local branch still exists, and must be deleted. I would prefer to use `git branch -d` to delete the feature branch, but this rightfully says that the feature branch hasn't been merged in. I instead need to use `git branch -D` to forcefully delete it, introducing a point of human error. (There are some cases where git can delete the branch safely, which I think occurs either when the feature branch has only a single commit, or when the feature branch can be applied on top of main without a rebase, but I haven't exactly determined it.)
#1 and #3 are cases where a safer option cannot be used due to a rebase-workflow. #2 would exist in either case, since even in a merge workflow, rebasing of branches before they are pulled makes sense to do.
FWIW: it occurs when the feature branch was based on the tip of master (because no-one else has committed to master since you branched/since you rebased onto master) - in this case rebasing your feature branch onto master is a no-op and the commits that go into master have the same hashes as they had on your feature branch.
For the rest, there's shell autocomplete and muscle memory.
I'm the guy people go to fix Git screw ups at my jobs but I just click a few buttons or drag a few commits...
Maybe you've already read it, but this is what let me grok the underlying data.
Comments like this, which points to a resource intended to help people "grok the underlying data", has the effect of seizing the focus of conversation and implicitly retargeting it to be concerned with with people who don't understand the underlying data model. When you been through this enough times, it just comes off as incredibly annoying and a source of tiresomeness.
Even worse, I'm not sure I correctly remembered the weird combination of actions and flags to use to get back to the state where I can continue with what I wanted to do in the first place.
That article is a good example of the problem. It tells me `git rebase` is an easy thing to do but I better not use that distributed VCS to publish my work that way, where 'publish' probably also applies to different machines of mine.
It's just that the command-line interface is very opaque regarding what it does to that data.
For instance, say I want to apply the last three commits I made in one branch to another branch. It's a very simple operation conceptually.
Good luck remembering that the command that does it is rebase, and what the arguments for it are.