Semver violations are common, better tooling is the answer
predr.ag
predr.ag
What we did:
1. Scan Rust's most popular 1000 crates with `cargo-semver-checks`
2. Triage & verify 3000+ semver violations
3. Build better tooling instead of blaming human error
Around 1 in 31 releases had at least one semver violation.
More than 1 in 6 crates violated semver in at least one release.
These numbers aren't just "sum up everything `cargo-semver-checks` reported." We did a ton of validation through a combination of automated and manual means, and a big chunk of the blog post is dedicated to talking about that.
Here's just one of those validation steps. For each breaking change, we constructed a "witness," a program that gets broken by it. We then verified that it:
- fails to compile on the release with the semver-violating change
- compiles fine on the previous version
Along the way, we discovered multiple `rustc` and `cargo-semver-checks` bugs, and found out a lot of interesting edge cases about semver. Also, now you know another reason why it was so important to us to add those huge performance optimizations from a few months ago: https://predr.ag/blog/speeding-up-rust-semver-checking-by-ov...
- cargo-semver-checks: https://github.com/obi1kenobi/cargo-semver-checks
- Trustfall query engine, which powers cargo-semver-checks: https://github.com/obi1kenobi/trustfall
- Trustfall playground, where you can query Rust library APIs in your browser -- for example, "which structs in `itertools` are importable by more than one path": https://play.predr.ag/rustdoc#?f=2&q=*3-Structs-importable-f...
- 10min conference talk on Trustfall: https://www.hytradboi.com/2022/how-to-query-almost-everythin...
I'm also giving a talk at P99 CONF in a few months about how Trustfall's new optimizations API made cargo-semver-checks over 2300x faster: https://twitter.com/PredragGruevski/status/16893002495908003...
More recently, we got tasty-mima [3], which checks compatibility at the type system level, rather than the binary level.
[1] https://github.com/lightbend/mima
1. If a new version of a dependency contains a desirable/necessary improvement (however that is defined), we want to automatically pick up the new version.
2. If upgrading to the new version will make things worse (however that is defined), we do not want to pick up the new version.
But there is no automated way of determining which of those is the case sort of actually trying the new version and fully testing it, even in the presence of "proper" semantic versioning (or any other kind of metadata you might provide).
- you may already have a workaround for the library bug in your code, so a non- interface-breaking fix in the library may break your usage until you remove the workaround. Semver can't help you with this because the library maintainer doesn't know about your workaround
- an update may have both properties (fixed one thing, broke something else). Semver can't make the judgement call about whether the tradeoff is a net improvement or not
- the update may change a behavior that is not formally guaranteed, but which your use has an (unknown to you) dependency on. So the semver can "correctly" label it as a non-breaking change, but it still breaks your use case.
- etc.
This doesn't mean semver is worthless. Knowing that something is definitely a breaking change because an API was removed or modified in an incompatible way will still save you time.
But you still need to actually test to know if updating a dependency is actually safe.
> Knowing that something is _definitely_ a breaking change because an API was removed or modified in an incompatible way will still save you time.
We want to help maintainers spot accidental breaking changes early — before they are released — instead of both maintainers and downstream projects having to scramble for fixes after an accidental breaking change is released.
Iff it's a part of the API you use directly or indirectly. Sometimes you're only using part of it, and changes to the rest don't matter.
The authors advocate that they can have the semver declaration test by a computer which unquestionably will increase the reliability of the declaration in the long run.
It is like testing but for backward compatibility, which of course can be part of the testing infrastructure some day.
I like their approach very much. What a great bachelor project... And the witness approach is great. Bonus points for the systematic approach.
Congrats to all involved.
The CUE Unity project is an interesting take on how to "solve" the bigger issue of breaking changes or regressions.
https://github.com/cue-unity/unity
Basically, the idea is that users can register their project, tests, or benchmarks in Unity. The project maintainers can then test new code against these projects before releasing, or even before pushing a commit, because a regression is found for example. This works by injecting the version (or local code) into the dependency management system. For CUE, this is a `go mod edit`, or setting up the environment with the cli at version / local. How this is set up and managed isn't as important as getting the process in place and making it an easy workflow for both sides. Hence why I call it a coordination problem.
We have a similar project Harmony, inspired by Unity, powered by Dagger.
https://github.com/hofstadter-io/hof/tree/_dev/ci/harmony
(it's really just a dagger pipeline that has a few flags and loops over data in a file)
it's probably impossible to automate an entire ecosystem, and there is value to enabling a tighter integration within a project ecosystem (a subset of the language ecosystem).
For example, Rails and React are the two most important dependencies in our system. They are extremely good about semver with major versions including highly detailed patch notes.
v43 to v44: fixed "teh" in a comment
v44 to v45: restructured the whole thing, steps for users to migrate are expected to be in the man-months for most projects
Semver is a human convention that can also frequently help machines. It'd be nice if it were a machine system that also helped humans, but very very few compatibility systems achieve that.
There are no other options currently to achieve this, and it works pretty well in practice. Use semver, don't make everyone else pay for your willful ignorance.
Have fun.
Instead of choosing v56.1, because that would be semver.
Rust compatibility (thus versioning) is fairly fragile due to its relatively powerful type system, and yet:
>Around 1 in 31 releases had at least one semver violation
Abandoning semver means making it worse 30 out of 31 times for a fragile language, or likely more (because you do not always use the thing that technically broke, or in a way that notices the break). Other languages generally have even stronger arguments in favor of semver.
It's not a "both sides" thing, semver wins by an absolute knockout in every conceivable way, and especially in practical day-to-day ways. Anyone opposed to it does not know what they're talking about, or they're simply trolling, and it shows. There could definitely be better systems, but "just don't use semver lol" is not one of them.
If the hints are sometimes incorrect, should we trust them? Isn't it the same release-verification legwork either way? If so, what's the point of the hint?
> Rust compatibility (thus versioning) is fairly fragile due to its relatively powerful type system, and yet:
Yes, I'm speaking about SemVer, which was neither invented by nor solely implemented in Rust. I understand that this blog post is about SemVer in Rust.
> Abandoning semver means making it worse 30 out of 31 times for a fragile language, or likely more (because you do not always use the thing that technically broke). Other languages generally have even stronger arguments in favor of semver.
I read this as: tooling is required to verify a release does or does not break a contract defined by a library, regardless of versioning strategy. It's not like we're talking about 1 in a million here.
> Anyone opposed to it does not know what they're talking about, or they're simply trolling, and it shows.
Ahh yes. "I'm right and if you disagree you're wrong". The ultimate debate tactic. I'll skip the rest of this thread. Enjoy the day!
No, as it turns out it very much is not even close to the same amount of legwork.
> I read this as
I don't think your attempt at a re-statement is accurate to the source. Also, "regardless of versioning strategy" is simply not an option in Rust, since the entire ecosystem is opinionated about how versions should work and overall everything works very well.
This post is about taking something that is already very rare (though painful when it does happen!) and making it even more rare. It shows a concrete way to do so, and also shows that so far we've been "getting lucky" by experiencing breakage less often than we might have been, given the same code changes.
In other words: `cargo update` today has a tiny chance of breaking your Rust project. But if everyone used `cargo-semver-checks` every time, as demonstrated by what the blog post's analysis caught, we'd reduce the number of accidentally-breaking releases by *3% of all releases*. That is an astoundingly high reduction, so coupled with good testing practices and other static checks (type system, borrow checker, etc.) it means that `cargo update` will be even more astonishingly unlikely to break your project.
It works the vast majority of the time. Many, many studies across many languages repeatedly show it. There are flaws, but there are extremely clear benefits - most updates just work, automatically. It works so well most people don't even realize how much they're being protected from.
So what are you gaining by abandoning it? You're losing a lot, if you're not gaining a lot in return then there's no reason to choose it.
---
>Ahh yes. "I'm right and if you disagree you're wrong". The ultimate debate tactic.
I agree in principle, but it is also simply true in some cases. It's like arguing against vaccines or range checks on arrays - there's a massive load of trolling or willful ignorance in semver debates. There can definitely be better systems, but none are used widely enough to be noteworthy to most, and most presented alternatives are laughably bad to anyone who has done anything but throw nonsense at the internet's wall and see what sticks. Some things are just not that complex, and armchair bullshitters can be spotted miles away.
If semver suggests a release isn't breaking, in Rust it overwhelmingly isn't breaking. Around 3% of the time, assuming you use every obscure bit of public API in the library, you might get broken by a release when you shouldn't be. In practice, that percentage is way lower than that: nobody uses the entire public API of anything, or comes anywhere close. That's why the post says these semver violations were unknown, even though most of them were multiple years old. If something breaks, and you weren't using it anyway, you were not broken.
The post is emphatically in favor of semver as used by Rust, because it has clear benefits and works great the vast majority of the time. The part that we want to make better is to lower the odds even further that an accidental breaking change breaks a bunch of people's code and ruins their day.
No single technique can do that: not testing, not fuzzing, not cargo-semver-checks. But a sound combination of techniques can drive the odds to damn near zero for all practical purposes.
I'm sure someone could build a semver-checker for Python or JS or Java or whatever, and I'd even be willing to help them and walk them through how cargo-semver-checks works. But in Rust, a big driver of adoption is that Rust and cargo are opinionated about things, so people adopt cargo-semver-checks not necessarily out of love for semver but out of a desire to not wake up to 100 people being upset in their GitHub issues about broken builds. The adoption situation might be harder for other languages with less opinionated versioning stories.
In the meantime, semver is good because it gives us a small number of additional semantics to impose on version numbers if we wish. Semver is not a solution to versioning, it's just a foundation to build better automatic versioning tools atop of, such as the tool linked here.
The best solution I've found is: each compiled library doesn't just have a single API, but rather has multiple.
libfoo-1.2.3 might implement the following APIs:
* libfoo-1.0
* libfoo-1.1
* libfoo-1.2
* libfoo-1.2.1 (because people don't actually use the "no API additions in patch releases" rule)
* libfoo-1.2.2
* libfoo-1.2.3
* libfoo-experimental-1_2_3 (no semver in the version here)
and if you enable experimental exports for the library, you gain a package-manager-level dependency on libfoo-experimental-1_2_3. Then later you can declare "libfoo-1.3 includes libfoo-experimental-1_2_3", so old programs can use newer libraries without change.
That said, mutable metadata is still an important part of the ecosystem. Blacklisting is essential (since the symbol-list isn't actually the whole API and it is impossible to automatically detect violations), but easily abused for "that's insecure" and "I don't support that" purposes so there should be some friction.
Note also that ABI (and sometimes API) is fundamentally platform-specific. And there's enough complexity in the world that the tooling shouldn't think in terms of semver, but rather arbitrary compatibility DAGs (but with sparseness when possible). Consider the GCC 4.7.0/4.7.1 std::list breakage that was reverted in 4.7.2, breaking software that was compiled in the mean time (the only reason this wasn't catastrophic is because nobody actually uses std::list) - that intermediate software needs to know it can't use new binaries.
`cargo-semver-checks` by default doesn't check for any semver obligations for code behind such features. (I'm its maintainer )
(also I'm not sure under what circumstances recompilation is required)
For us, we audit every update of every package. We read the changelog or release notes for every version from the current version to the new version (some packages make this very cumbersome, looking at you Vercel). Then, we use our judgement about the order in which to update packages; usually we apply all patch-level updates, run our tests; then minor updates, re-run our tests; and then we apply major updates one-by-one and re-run tests after each of them.
We've found plenty of semver violations in this process, to the point where we don't really trust the semver of NPM packages at all.
Is this a common approach, or do you just blindly update everything at once and fix things if they break?
Sometimes it isn't fine, of course. Those cases fortunately are few enough and not damaging enough that this approach hasn't been that problematic thus far.
And read release notes for everything you upgrade too.
Most established JS/Ruby codebases I've seen have dependency technical debt and they _know_ they're going to run into breaking changes trying to upgrade. These companies need to do the project management work to break their debt down into as many small incremental changes as they can. Any upgrade that has a breaking change gets done standalone (you need to read the changelog to figure out if there's a breaking change, you don't trust semver). Upgrades that are necessary but safe can be batched together. When you have to make code changes for breaking changes these should also get done incrementally ahead of time. Usually these changes are backwards compatible (e.g., if a deprecated method is going away in v2.0, you can switch to the new method in v1.X so you're not upgrading the package and changing your code in the same PR). Even if they're not trivially backwards compatible you can often make them so, and you should. I've done individual upgrades consisting of dozens of small pull requests. When you do it this way you don't get stuck with a long-running branch that gets abandoned, you roll back fewer deploys, and changes are easy to review.
Once you're on the latest major version of all your packages then you want to get into an ongoing cadence of dependency maintenance. Good small-medium teams tend to do a maintenance rotation where you spend ~1 day/sprint of one developer's time on this upgrade work. Here project management is also important - ideally someone is looking at CVEs, stale packages, and other indicators of package risk, reading changelogs to identify effort, and breaking out maintenance tickets. Some of these tickets will be individual package upgrades, some will be batches of safe upgrades, and some (when a new major version of a framework comes out) will get treated like the previous paragraph. Larger teams might consolidate this work in a platform engineering team, but that has its own organizational challenges.
Is there similar guidelines for other languages like Java or Python?
def do_it(*args, **kwargs) -> None:
...
You might keep the signature constant, but let the behaviour evaluate across versions, but semver-wise, you could consider that API unchanged.Rust generally has more exceptions than most languages about which breaking changes are technically considered non-major, so writing such guidelines for Python or Java shouldn't be too difficult.
The same tech and ideas that power `cargo-semver-checks` could also be repurposed for those languages as well. If a company is interested in sponsoring such work, I'd be happy to help build something like that!
At first I thought it was insanity. Now, buried in a maze of transitive third party and first party dependencies scattered across multiple git repositories and crates, I really really miss it.
Of course like many other things there it depends on having lots of well-paid staff, without deadline guns to their head, committed to code quality and infrastructure.
Yes, no matter how hard you try semver will sometimes get it wrong, but as long as it usually doesn't it's still useful to let my tools do their best at automatically selecting a version of the library that will be compatible with every part of my application that cares, and that has had as many bugs fixed as possible. That's what semver allows cargo to do, and 99.9% of the time it just works. 0.1% of the time you have to manually intervene, and that's ok. It's better than having to manually intervene 100% of the time, or the old-school C solution of "just pick whatever is lying around on the OS and hope it works/is compatible".
...
This article starts with the wrong assumption that semver is itself meaningfully correct and useful. Semver is neither of those things except in a very specific circumstance that almost nobody has.
Semver violations are common because:
1. Semver as defined by semver.org is founded on the fundamentally false/flawed premise that you can predict whether a change will break compatibility. You can't. You should still try, but you should also recognize that sometimes you will guess wrong. And sometimes when you guess right you will still be wrong. The notion that just because you didn't add/remove an input or output parameter means you didn't change your API is extremely naive.
But this isn't the important reason. The important reason is...
2. Even if that weren't the case, it still would only make sense for approximately 0% of all versioned software projects, because almost no projects in the wild continually preserve separate feature branches with backported bug fixes, and continually preserved feature branches with backported bug fixes is the only scenario where semver's separation of features that aren't expected to break compatibility and fixes that aren't expected to break compatibility makes any kind of sense, because there is no such thing as a clean separation between feature and fix.
I hope that the truth of #1 is obvious to everyone, but if it isn't please refer to https://xkcd.com/1172/ Less pithily, internal changes alter outward behavior literally all the time in unintentional ways, and no software project on the planet defines and constrains the expectations for its interface with sufficient clarity to avoid that. Not one. It's probably not even possible.
As for #2, consider the following scenario...
You have released version 1.0.0 of something. Then you add a feature and fix a bug unrelated to that feature. Are you at version 1.1.0 or 1.1.1? Well, it depends on the order you added your changes, doesn't it? If you fixed the bug first you'll go from 1.0.0 to 1.0.1 to 1.1.0, and if you add the feature first you'll go from 1.0.0 to 1.1.0 to 1.1.1. And if that difference doesn't matter, then the last digit doesn't matter.
The problem here is that Semver treats feature and fix as fundamentally distinct from each other. But non-breaking is non-breaking. The user on the other side of the fence wants the best version of whatever you have that doesn't break compatibility with what they're doing. If they trust you to make that distinction, then they only care about your major versions. If they don't trust you to make that distinction, then they only care about strictly matching the whole string. If you trust yourself (lol), you can cater to both groups with two-part major.minor. If you don't trust yourself, you can just use a single value that increases over time. But semver has three fields, and one of those fields is basically always completely useless except to satisfy a contractual obligation.
There is only one scenario where three version fields matter, and that's when a government defense contract forces you to fork development at the start of each project and then maintain separate code branches for separate projects, where they only get the features they ask for and you only fix the bugs they ask you to fix (this is exactly how defense contracts work), and you laboriously backport bug fixes to all of them, and the major and minor versions indicate the point of the fork, and the patch version is all the changes applied to that fork.
Great, you agree then that this tool is useful because it will help you try to predict whether a change will break compatibility.
I can barely follow the rest of your comment. Semver is useful in, at least, a decentralized set of actors trying to minimally cooperate with one another by using a short number to communicate coarse but important information with respect to compatibility. Guess what that describes? A FOSS software ecosystem. So...
> Semver is neither of those things except in a very specific circumstance that almost nobody has.
Nope. I would say FOSS is pretty popular.
No. I guess you didn't actually read what I said, and decided instead to construct a strawman by carving very few words out of it completely ignoring the rest of the words around them.
> I can barely follow the rest of your comment.
Ah ha.
> Semver is useful in, at least, a decentralized set of actors trying to minimally cooperate with one another by using a short number to communicate coarse but important information with respect to compatibility.
It's not any more useful than a system with only either date/incremental (this is newer than that) or major.minor (I promise that this change breaks, and that other change might also break but I hope not) versioning except for the one case I mentioned, which FOSS decentralization has absolutely nothing to do with.
So I'm a bit confused about the position you seem to be arguing against, when the tool and the entire ecosystem are doing exactly what you (AFAICT) consider a good idea.
Semver.org says:
Semantic Versioning 2.0.0
Summary
Given a version number MAJOR.MINOR.PATCH, increment the:
MAJOR version when you make incompatible API changes
MINOR version when you add functionality in a backward compatible manner
PATCH version when you make backward compatible bug fixes
Three fields. If you're saying that cargo-semver-checks doesn't need three fields, then you're saying that it doesn't benefit from semver as defined by semver.org.From what I can tell, you're trying to argue that semver is useless while also saying that not following semver means that you don't get any benefits, but that just doesn't seem like a proper argument, so I'm surely misinterpreting.
I think SemVer is a social construct, but that most of the time we can predict what changes will be breaking. We may get it wrong of course, but that doesn't mean we should throw our hands up and not try.
For tooling it's very hard to detect between a feature and a bug fix so it's almost impossible to automate that series of checks
Indeed, why? And yet semver itself does, which is senseless.
As a user updating my software it's nice to know the gist of an update at a glance.
As a piece of software looking for breaking changes, it is unnecessary to consider.
Semver doesn't help you do that. If you're on version 1.0.1 of my software, and I update to version 1.1.0 and then 1.1.1, have I fixed a bug that exists in 1.0.1? There's no way for you to know based on the version number.
Knowing whether you need to update because of a bug is something you can only find out from change logs, not from version numbers.
If I didn't change the behavior with any set of input parameters valid for the prior version, I haven't broken compatibility.
> because almost no projects in the wild continually preserve separate feature branches with backported bug fixes
Maybe not “continually”, but nore than 0% of real software projects using SemVer do release bug fix releases of m.f.b after feature releases of m.(f+1).0
> The user on the other side of the fence wants the best version of whatever you have that doesn't break compatibility with what they're doing.
They may not, because feature releases may be assumed to have greater risk of expanding the risk surface, whether or not they break backward compatibility.
You don't anticipate that the behavior will change in a way that you care about, but that's not the same as not changing behavior. Different sides of an API nearly always have different definitions of "valid" and "behavior" here, and API documentation is never as concrete as would be needed to prevent that. So are you exhaustively fuzzing the input space? Because if you are, then great. But I bet you aren't. And even if you are, nobody else is.
> more than 0% of real software projects using SemVer...
I carefully said approximately. It gives us some leeway when talking about this without falling into the trap of believing that a small degree of adherence is meaningful.
> feature releases may be assumed to have greater risk of expanding the risk surface
I know it's common for people to feel that way, but there isn't a meaningful interface safety dividing boundary that any of us can point to between development of a feature vs a fix. Code is code is code is code. The developers can either reliably identify when changes will break users or not. If they can, then breaking.nonbreaking is plenty. If they can't, then only a single newer_than_before value is fine.
Like, if you want to argue for a new Dragonwriter's TrustVer where we talk about degrees of belief about whether a change will break something, we can do that. I'm all for it. But more than two values will still not be more useful to the user than at most two (except in the one contractual obligation case I presented), and it won't be Semver.
Given two versoins, 1.1.10, and 1.2.18, which is newer? Who knows!