Proper use of Git tags
blog.aloni.org
blog.aloni.org
I've never understood this practice of not immediately bumping the version in source after a release. We update the version in source to the next logical version (usually a patch bump with an alpha0 pre-release tag) immediately after tagging and publishing a release. This way you just need to look at the version in source to understand where you're at in a release process, and only a single commit has the full release version in source, which is also tagged as such. This doesn't seem to be a common pattern, so shat are the downsides of this approach? Am I missing something?
To illustrate to readers, let's say we are developing over a tag `v1.2-pre` in `main`, and now the code has stabilized. The idea is to release a tagged `v1.2`, then in the `main` branch immediately release a tagged `v1.3-pre` that follows it.
We obtain the following:
1) `main` branch commits look like `v1.3-pre-<number>-g<hash>` in `git-describe`.
2) A maintenance branch may have been forked from the commit that was tagged with `v1.2`, so for the commits in that `release/1.2` branch you automatically get `v1.2-<number>-g<hash>` as output of `git describe`.
One possible downside (of version bumping immediately after release) is that if you use the major/minor/patch scheme, you can never be sure that you’re actually bumping the correct number. The upcoming release may be a major, a minor, or a patch release, and it might take you a while until you know enough about completed work items so you can make up your mind which one it’s going to be.
Seems like any way you do it you have to keep track. IMO version numbers are slightly easier to keep track of than breaking changes. Changelogs are not always structured.
Minor annoyance, but still.
An alternative model I suppose would immediately have major bump, minor bump, and patch bump branches; then you just commit to the appropriate one, and I suppose keep major rebased on minor rebased on patch. (And master = major I suppose.)
I wonder if it really matters that release version numbers only increment by one. If not, just bump anyway when appropriate change is made - no need to check.
In practice I think the problems would be a) having to be very disciplined about this on every commit, rather than having a reminder to consider it as part of a release process, and b) ordering version numbers from different branches when merging to mainline
Or you could bump every time as you describe, but on every major/minor bump make sure the parent commit is released first (which would be >=1 commits since the last one and have at least a patch bump). And I suppose you'd never need to bump patch if that was just a lone thing that happened post-release in prep for the next.
Assuming that you do a strict hierarchical versioning scheme and not some multidimensional thing where numbers increment independent from each other.
Benefits of keeping the last version
- Easier to detect when a change occurred since the last release
- (Rust specific) It makes it harder to patch a registry dependency with a git dependency because the versions will never align [1]. This is why the default changed in cargo-release.
- The "next version" is just speculation. Will the next release bump the pre-release version, patch, minor, or major?
A downside to keeping either last version or next version is if someone builds from master, the bug report could be confusing unless you include the git hash (and whether the repo was dirty).
I've seen some advocate for not keeping a version in source at all [2]. This article advocates against it but doesn't give the reason. I guess one is it requires you to have all tags locally which is a silent failure mode when you don't and disallows shallow clones.
[0] https://github.com/crate-ci/cargo-release
it was a never ending source of merge conflicts
branching also causes problems (and mistakes)
vs. the build process deciding (possibly with some human input) what the tree's version is, then tagging it
For GitLab there there exist a workaround by using .gitattributes to create a VERSION file with "$Format:%(describe:tags)$" which will get expanded to the git describe string on archive creation, so the version number even survives in a .zip. GitHub however ignores .gitattributes and so far I haven't seen another way to get a version number into the .zip other than just including it in the source.
Gives a HEAD.zip that contains a "<repo>-<sha1>/" folder and the latest source.
Then, running `make_release tag`:
1. Sets the version to `4.6.1`, commits it, tags it.
2. Sets the version to `4.6.2-dev`, commits it.
I then merge the 4.6.1 tag to main and push main, develop, and the 4.6.1 tag to the repo.
If development calls for a minor or major release, that's what the `bump` option is for. It prompts for which part of the version string should be incremented and creates a commit doing just that.
Setting the version usually involves touching a couple of files. e.g. a source code constant and a podspec or some other metadata file. The script really helps preventing any mistakes.
Where I work we don’t increment until we know what branch/version we are working on, and determine the semantic version, and then we still use tagging because we tag the version + build id which is computed at build time
I personally think that's messy because it requires two changes to the version instead of one and there's no way of knowing that the next version will be 1.4.3. If a developer makes a breaking change, will they update the version in the code to 2.0.0.alpha0 or do they fix the version when they release?
I suppose if you have nightly releases then it could make sense but then I think using the date or the commit hash as the version number would make more sense.
That’s some impressive prognostication! I recently released a semver major for a minor corner case bug fix. I couldn’t have predicted this would be version 4 after over a year maintaining version 3, but it was a significant breaking change to fix a minor bug.
Unreleased software doesn't require a version number at all.
What you can do is incorporate some build identifier into local unreleased builds. It could be from the git sha, plus some indication of whether it was a dirty state or exactly that git sha.
The command does it:
git describe --tags --dirty
It will produce a string consisting of the most recent tag, the number of commits since that tag, a portion of the hash and the word "dirty", all separated from each other by a dash. If there are no uncommited changes, then -dirty is omitted, and if there are no commits since the tag (we are at the tagged commit), then the count of commits and hash are omitted.You can incorporate this sort of identifying string at build time; nothing is checked in. If you build the tagged release in a clean repo, the string will just be the tag. Only if you make new commits and/or have uncommited changes do you get a different string without having to commit anything.
My personal experience with go modules and versioning has been really positive. In that ecosystem - you only define your major version in the mod file and rely on the VCS for everything else.
I know it is working well for Linux kernel hackers. Linus does all the tagging, and I don't recall issues with merge conflicts on the top Makefile with the part that holds the version.
In many environments, multiple branches of software are deployed and maintained at the same time.
Even your example with Linux; I don't think that is correct -- Linus doesn't tag the stable branches, and others. There is a centralised agreement of how the version numbers are maintained though.
Also ties in with the author's recommendation to begin tags with "v". My experience is that excluding it is better. Then a simple "git describe" readily gives the version number in scripts with no sed or reprocessing.
I've seen many conventions with Git. It's interesting to hear some rationale, but a stretch to describe these suggestions as the "proper" way.
As also pointed out, many build tools that want or need version numbers often also have command line flags or can take environment variables instead of using source files.
In practice that's a tiny bit more complicated to do it really well, with the version as a dependency in the build process. When developing, you don't want to trigger a full rebuild if the version number changes, but you do have rebuilds to do when it does.
Related to that "tags as source of truth" is using tags as a deployment trigger. A release manager applying a tag can be a signal or gate for a version to go to later environments. (For instance: CI builds from main branches stop at Dev environments, CI builds triggered from new tags automatically move on to UAT and Staging environments.)
Also, another tip I've found useful for people with more "monorepos": tag name restrictions match branch name restrictions and you can use a "folder structure" of tags. You can name tags things like subcomponent/v1.0.2. Some git UIs even present such tags as a folder structure. Doing that can confuse git describe, of course, so finding an arrangement that works for your project may be a balancing act. I've used lightweight tags for subcomponents so that git describe doesn't "see them" by default and then you can use the git describe --tags that also takes lightweight tags into account if you need a "subcomponent version" for subcomponent tag triggered deployments (and then you just need to remove to remove the folder prefix).
Using the --match option, `git describe --match='subcomponent/*'` fixes this problem. It filters the tags that are considered to only those matching the pattern, so that a later tag for another subcomponent will not be used.
Agreed entirely on restricting tags (they tend to have meaning and expected semantics outside the repo, which makes them risky), but also! Teach people more about git, or unix CLI patterns in general. `git log test` is ambiguous about it being a tag (or branch) or file/folder, but `git log -- test` is not. `--` as an ambiguity-preventer is a very common pattern, it works in many, many CLI tools.
---
One that hasn't been included here: don't rely on tags for any kind of business logic, if you have literally any way to avoid it.
Tags are mutable, and do not have a history of when you changed them. They're plenty handy for "human, enter a thing to use" purposes, but you should immediately turn that into a commit sha and then only use that commit sha.
$ rm .git/refs/tags/v1.0
or in git syntax
$ git tag -d v1.0
For this reason one should never use tags when pulling code from untrusted repos, instead use the SHA1 hash which is much harder to forge.
That is not what anyone means when they talk about immutable objects or refs in Git. I can delete a commit from `.git` but they are still considered immutable.
A branch is mutable since you can push it forward without `--force` as long as the current commit becomes an ancestor of the next tip. Can you change a remote tag without `--force`? I don't know off the top of my head. But I doubt it.
To /tmp/foo1/
! [rejected] ddd -> ddd (already exists)
error: failed to push some refs to '/tmp/foo1/'
hint: Updates were rejected because the tag already exists in the remote.
I think any ref is 'mutable' by these standards, and any hash is immutable (that's the entire point of a hash).In any case, the syntax to delete a remote tag is just e.g.
$ git push origin :mytag
After that you are free to push a new tag to replace the old, with no warnings or --force flags.
When we use a tag to specify the version of a dependency, we trust the maintainer not to do this. If we don't have that level of trust we can use SHA1's instead. We should not pretend that there is anything in git that tries to prevent a tag from being rewritten by someone who wants to.
Tags do not do that. Moving or removing a tag retains no history about the move, nor is the old value left hanging around somewhere (after pruning), nor will anyone who had not yet pulled the tag notice the removal, as with branches. Most configs will complain about the tag disappearing or moving, as with branches, but that depends on your config and the command you ran.
(annotated tags do have their own sha and creation date and whatnot, which is great, but next to nothing references them. and removing them from a commit leaves no evidence that it ever was on that commit, as the commit is unmodified)
Needing `--force` with a default setup is literally all I meant by "immutable", apparently a cursed word in this context (the tip of a branch is supposed to be able to move in that common-history sense of movement). Gawd.
normal tags are just labels that can be changed or deleted at will. removing a normal/lightweight tag from history doesn't change any commits and doesn't require a force push.
that is why the article recommends annotated tags. you have to force push to rewrite those. then the git describe tags based on those will be immutable unless someone goes out of their way to rewrite history.
If you move a tag, annotated or no, the .git/refs/tags/x file will contain the new sha it points to... but the history of that `x` is not stored anywhere. The fact that it used to point to [old sha] is gone for good, and the old sha will eventually get garbage collected (via pruning), just like if you remove a branch.
Now I'm not sure what the point is of annotated tags other than the default git describe behavior.
Everyone else can enjoy Cunningham's Law working like a charm.
Off at the less-structured side of things though, have you seen git notes? You can add notes to any commit: https://git-scm.com/docs/git-notes . Tags are kinda just notes with a display name.
At times I've wished that code review feedback were just stored in notes, so it'd survive changing hosting systems... but they're not quite reasonable for that :| What I think I really want is a notes tree under any commit, which kinda exists since you can add notes to notes, but there isn't really enough structure or tooling to support that kind of use out-of-the-box.
Similarly, many repository hosts can help you setup tag protection as a part of branch protection tools, but while that helps with your own repos it doesn't generally help with remote repos.
And agreed, tag protection rules do exist and are fairly common. Though by far the majority I run across do not protect tags or branches by default. And even if they do, external systems may or may not honor changes - that's why dependency management lock-files exist, to detect changes like this where the "name" (i.e. tagged version) stayed the same but the content changed.
Or in a different flavor, you have Go modules, where you cannot ever remove or mutate a tagged version in the main proxy... but you can change it in github, and now your web-UI-visible code differs from what people download. Which may be worse, because while go.sum will store the go module checksums and can complain if you pull the wrong contents, that checksum doesn't match the sha it pulled. If you have a module-compatible tagged version, the git sha isn't stored anywhere, you just have `require thing v1.2.3` and the go module content hashes. Trying to "recover" the sha from this can be rather painful, as you essentially have to check the module checksum for all shas in a repo... assuming it even still exists.
Right, as with many things in computing often you want "trust, but verify". Trust a good tag by a good author not to change, but also go ahead and store a hash in a lockfile and verify it, just in case.
I believe the right thing functionally is a child branch, updated automatically with either a merge or a force-push on new successful builds. But it feels not quite right conceptually, and it's harder to explain to new developers (who can already get overwhelmed with handling multiple branches). Is there a non-branch solution for having a moving unique label?
You can force push it, but why not keep a trail of what your past state was?
And you furthermore have no immutable record of what happened anymore.
You can always create one tag that you force-push to point to the "latest tag" to make it convenient for one-liners, but not force-push any of the "real" ones.
I once made a blog post / video around signing and verifying your commits and tags at: https://nickjanetakis.com/blog/signing-and-verifying-git-com...
In the first section, are you saying that:
a) tags can and should be named to include the sha at the end; and
b) git commands silently discard everything before "-g" if what follows is a sha?
I feel like I've missed something as I find b) very surprising. (Consider giving someone the string $LATEST_VERSION_NUMBER-g$MALICIOUS_VERSION_SHA; I can't work out an exact exploit but it seems wrong to only process the SHA.)
Of course I'm not suggesting that tag names should include the hash. The existence of the tag `v4.11-rc7` allows other commits to have nicer derived names.
EDIT:
Also to your last inquiry, it may be indeed a surprise that Git resolves `<anystring>-g<githash>` to `<githash>`, but some may argue it's a feature, not a bug :)
I'm curious if the git resolution algorithm that accepts describe strings also verifies/checks the commit count matches. Glancing at the documentation [1] it says that it specifically matches describe output strings (the docs call it <describeOutput>) and not just "<anything>-g<hash>", so it may actually check the tag existence and verify the commit count.
Has there been a better proposal to semantic tags which makes organizing and sorting them easier?
[1]: https://calver.org/
I searched for a bit and found that github does this ( e.g. https://github.com/TryGhost/Ghost/tags ). Also there is a git config that sets this globally for the git cli: `git config --global tag.sort version:refname` ( found here: https://gist.github.com/loisaidasam/b1e6879f3deb495c22cc ). That's another git config going onto every machine I use...
Git has built in version sort, though it's not the default. GitHub shows tags in version sort order in the drop-down tag selection.
Here are some commands that use that sort order:
ls -v # Linux only.
ls | sort -V # Most modern OSes.
git tag | sort -V # Git tags in version order.
git tag --sort=v:refname # Same as above, it's a git builtin.
git config tag.sort v:refname # Set default sort order for `git tag`.
git config versionsort.suffix -pre # Make 1.2-pre1 sort before 1.2.
git config versionsort.suffix -rc # Now 1.2-pre1 < 1.2-rc1 < 1.2 < 1.2-patch1https://www.gnu.org/software/libc/manual/html_node/String_00...
(... well, it's copied and pasted into the tree but close enough)
[1] https://en.wikipedia.org/wiki/Perl_5_version_history
Just to be clear, I don't actually think this is a great solution, but is _is_ easy to sort.
So instead of a "v4.11-rc7-87" tag you would have a "v4.11-rc7-87" branch and a merge commit that holds the meta info about that branch.
Compared to a tag, you’d have to deliberately go out of your way to move a tag (`git tag -f`). Generally, you should be able to trust your team members not to move a tag.
Branches can change over time, whereas tags are largely immutable. A change to a tag always requires a force push and if you've got branch protection tools in your repository host's arsenal you can often entirely prevent tag changes.
I'd also like to point out that often if you are using merge commits as your "version management tool" there's a lot of people that use "branch per environment" strategies and need to merge between environments. There's a lot more risk in that approach than a tag-based approach: if a tag doesn't change after it is applied, then you don't need to rebuild binaries to deploy it again. If you don't rebuild binaries between environments you have to make sure the same binaries work in all environments, which is good practice. I've seen too many times "branch-per-environment" strategies wind up with per-environment code that is difficult to untangle and makes testing and debugging difficult and merge conflicts more likely and more complicated and increases the risk of per-environment bugs.
git init ./repo
cd ./repo
git remote add origin $uri
git fetch origin $tag
git checkout FETCH_HEADIf you have a checkout, you can find the tag name by doing a git tag -l *14* and looking through the output; I'm intentionally providing a very general glob because one may not know beforehand their exact tag name conventions.
I would argue that any tool (such as a GUI app) which doesn't provide the depth of usability that a shell offers is, in fact, a deficiency in the GUI or tool. If your GUI doesn't offer completions similar to a shell then your GUI tool is inferior to a shell.