Package Management: The problem with using version ranges
lucidchart.com
lucidchart.com
This article is terribly outdated. Version ranges are great and the only way to manage correctly your dependencies without having tons of duplicates. Also they are a security guarantee, because with them you can get security fixes “automatically”.
Yet there have been problems, that’s why both Yarn and npm@5 implement a lock file which gives you all the guarantees of a range-less package management with the power of a range-based dependency tree.
If module developers following this (IMHO not lucid) article start bundling specific versions in their npm packages everyone will suffer. It’s as stupid as it get.
Please, do your research before calling out on practices.
Too many open source projects fail to ensure non-breaking changes or long term support of multiple release versions.
So either you rely on greatest(>) that has no guarantees to not break and not just patches (~), you try patches and a new release is effectively end of support or you rely on greatest, fail to do a regular audit and realise the dependency is no longer in active development and will receive any updates (which upon typical discovery would require a change of dependency used).
Security isn't just about vulnerability, and I don't like to ask it (because it isn't universal), but is not a risk of system failure often worse than a risk of system vulnerability? It depends somewhat on how critical the system is and defaulting to explicit versions instead of latest is probably a better practice in my opinion.
I've done development using NodeJs and on the JVM and found I had far fewer problems with Java dependency management.
Relying on those projects is not a good thing.
Automatically when you request an upgrade to your dependencies. Older npm's defaults were terrible here, but as of npm5, it works more like Bundler/Cargo. You get deterministic installation, and "bundle update" lets you get a tool-assisted upgrade.
Semver ranges (dependency 1.0.x) along with version pinning (lock files that represent the currently installed versions) are the best-case scenario. One can avoid the exercise of "bumping versions" every once in a while (just upgrade all packages within supported semver ranges), get the same semver compatibility guarantees, makes sure all other team members are using same versions, and most importantly the CI process is reproduceable and can be sped up with reliable, non-broken dependency caching.
The elephant in the room is Java. Despite its maturity, both mvn and gradle have poor support for ranges and pretty much no concept of lock files. Everyone is left to "bumping versions" by hand all the time, and snapshot versions are pretty much hell and lead to irreproduceable builds.
Vendoring the whole tree is only an optional add-on to semver ranges. You can't vendor a whole tree and resolve all nested dependencies if everyone is using directly pinned versions! I think the article might have a fundamental misunderstanding between versioning as a library author and as a user.
TL;DR - Semver ranges are great when coupled with version pinning and lock files. If not, there are some trade-offs, and as long as people are aware of the trade-offs, its still better than no ranges.
Maven-like snapshot are a requirement for an environment with poor support for transient semver-ranged dependencies, which maven is…
Haven't worked with them myself, but colleagues are singing high praise.
If you read the whole article, I conclude with recommending lockfiles and npm shrinkwrap, the traditional JS lockfile solution. Yarn is the new kid and awesome and deserves mention too.
Lockfiles are an appropriate choice for JS because of unique dependency loading. For Java and many other ecosystems, fixed versions are appropriate because they automatically deduplicate.
There are pros and cons. The absolute most important thing is to achieve determinism. Programming is hard enough as it is.
You try maintaining an ecosystem where every dependency constraint is a hard specific version. You'll spend 80% of your time fiddling with bumping versions. Fuck that shit.
Yes, the developer might like a library, middleware, application server or whatever to be a specific version, but that ignores that in a production environment you can be behind on security fixes.
It's completely reasonable to lock your code a specific major release of some dependency, assuming it's still being patched, but minor release should be something your patch management solution just applies.
From an operational point of view, I think the original approach of Go, where HEAD of a dependency was just pulled from the repository of the dependency, is the right way to go. If something break, you fix it, because it needs to be done a some point anyway.
I see this all the time with Java middleware. Some developer specifies that we need JBoss version X.Y.Z and not patch unless they say so. At the same time we're required to patch the operation system with the latest security patches. Well what's the point of that if the only service exposed to the Internet is the only thing not being patched?
But version pinning is probably the best way at the moment.
This way the maintainers can use ranges and the developers can pin what they really need in the end-product.
There are more developers than maintainers, so this scales much better.
Wrong comparison. The majority of software is built, deployed, maintained and used by entirely different organizations over years.
Companies doing end-to-end CD and running their software only internally are the minority.
This kind of bad practices from few developers turns into countless hours of work to maintain systems years later.
It would be a completely unacceptable situation for application A1 to depend on DB abstraction library v0.8651, A2 on v0.8653, A3 on v0.8670, A4 on v0.87…
How do you ship this to your distro users? Almost always there is no easy way to install several libraries of the same name but different versions in parallel, and it's never supported by the distro auto-packaging tools.
Patching to use the latest version of the dependencies and verifying the patched version works ok is a huge burden on package maintainers. The application author knows the code much better!
The current solution of specifying minimum versions of dependencies exists not because we like the suffering caused by occasional breakage. No, it exists because it is a practical, working solution for the real world that is the best balanced trade-off.
In PHP we usually have a version range in libraries' dependencies and lock file that specifies exact dependency versions for the entire application. So you can be sure that you get tested combinantion of libraries.
Some do. Some others target the versions shipped in a stable Linux distribution.
And even if the upstream don't test some combinations of versions, distributions will test what they ship.
How can they check future versions of libraries that will match the range?
How is this possible is beyond my comprehension. Someone, anyone break something and I have no standard way to bypass "dll hell" (putting space problems aside)? what the actual f.
/usr/lib/perl5/vendor_perl
/usr/lib/python*/site-packages
/usr/lib64/ruby/gems/*/gems
/usr/lib64/node_modules
You have only one version of each library. This is the exact opposite of DLL hell.Perhaps go for a range + last tested version, e.g. "1.9.7 worked" so when 2.0 comes out and breaks, you know what to use instead.
If you change the output of "ls", you know everything that uses ls will break and you can bump a major version. But if you change e.g. the date format to include a timezone (visible when -l is specified), does such a change really warrant a major version bump? It's just a minor change but it might still break things. Or maybe if you refactor some stuff in curl, the curl-dev package might suddenly be incompatible with something that depended on it.
Not sure how node would handle the same type of error but I'm guessing it would involve subtle compatibility issues.
However.
The reason people use ranges is because they don't want to handle the administrative burden of tracking upstream changes and updating their software.
The OWASP Top 10 (2017 draft) shows that using known-vulnerable components is a common security weakness in software. Nobody sets out to do this, but the flipside of having access to massive troves of dependencies is that you have dozens, hundreds, perhaps thousands of dependencies you might not know about.
I've worked on Cloud Foundry Buildpacks. A large amount of effort has gone into exactly this problem: tracking upstream versions, pulling new ones immediately, building them and testing them. But it's at varying levels of sophistication.
For some dependencies, you get structured data. NodeJS publish an index.json file[0] which is fairly trivial to keep up with. But for other sources, which I will leave unnamed, it's necessary to do flaky, unreliable HTML parsing to detect new releases. Data about what's been found so far is kept in a ci-robots[1] repo to ensure atomicity and simplicity (I'd prefer a Concourse resource that emits versions one by one, but that's just me).
What's needed is a uniform way to describe available versions of software packages. Uniform across all the major tributaries of versions: package managers, distributions, the works. If it becomes possible, as Buildpacks have after a great deal of effort, to automate fetching and testing new dependencies, then the need for ranges vanishes. You can rely on your CI/CD to run each dependency change through its paces and update the firmly fixed list for you.
[0] https://nodejs.org/dist/index.json
[1] https://github.com/cloudfoundry/public-buildpacks-ci-robots
For example, there's a long tail of one-off packages on Hackage that are really valuable to keep around, even if they're not actively maintained and bumping versions would cause breaking changes. I'm thinking of all the various logic frameworks, theorem provers, etc. which are out there; maybe to accompany some paper from the 90s. They're certainly not "core functionality", but every now and then it may turn out that one of them is exactly what is needed in some app/library. Every few years the major repos may drop the last working version of GHC, so someone may have to step up and do the required fixes, but that's few and far between.
These days I use a combination of Hackage (all versions of all packages), Cabal (to solve dependencies), tinc (to make Cabal deterministic) and Nix (to cache everything).
Stackage puts considerable effort on the whole community to ensure that packages remain compatible, and forces package publishers (developers) to update their packages in a timely manner (otherwise the package will be dropped from the snapshot). That works in the Haskell community, but I'd be hard pressed to believe any javascript developer who has a package on npm.org would put the same amount of effort into keeping it up to date.
That means there is no such thing as 'too big', because the number of all existing packages in a ecosystem doesn't matter. It's the commitment that counts, and that's what I don't see in the JS community.
now when you talk about actually developing the software, this is a completely different set of requirements and that one is best served with very coarse grained (but safe) version ranges. different tool for a different use case.
I completely agree with the article. Version ranges are evil. Avoid them at all cost.
Because other people have different priorities and have come to a different cost/benefit ratio value than you?
Another reason is that many developers are obsessing on getting the latest version of their dependencies for fear of security issues or just missing out on the latest and greatest - and they often completely ignore retesting the application since they now have someone to blame if it fails (that other developer should not have pushed the breaking change with a minor version bump!)
I agree with you that it should be the standard to have fixed versions and update your dependencies at a time of your choosing so that everything can get tested properly - but it seems to be an uphill battle.
Getting the latest version is how you get new vulnerabilities.
Various software distributors, including some Linux distros let software bake in for this reason and can be even faster than the upstreams in developing and applying patches to known vulnerabilities.
Also, unfixed but known vulnerabilities are less dangerous: security and system engineers can work around them, also IDS/IPS can detect and often block attacks.