Never Update Anything
blog.kronis.dev
blog.kronis.dev
This is what applications used to be like, before the web and internet hit and regular or even push updating became easy.
It was simply so difficult and expensive to provide updates once the software was in the customer's hands that it was done intentionally and infrequently. For the most part, you bought software, installed it, and used it. That was it. It never changed, unless you bought it again in a newer version
The internet has normalized poor QA. The bosses don't give a shit about QA anymore because it's so cheap to just push out a patch.
I mean just look at old video game magazines that talked about the software development process: the developers would test the hell out of a game, then test the hell out of it again, because once it was burned onto a $100 cart (in 2024 dollars) it wasn't ever going to change.
Now games can remain buggy and unstable for months or even years after "release."
Luxury! With Nintendo you'd often get only one ticket. Any further bugs would cost you further submissions, and months of slippage.
aka The Correct Answer™
For the products I managed in the 90s, I put QA/Test in charge of releases. Very unusual. The results were awesome.
Back in the 80s and 90s if you had bad QA you'd ship very buggy software and customers hated you because they had to live with it for months and months until the company managed to do another release. And then it was costly to ship off those floppies to every customer. So there was a very real price to pay, in money and reputation.
Then it became possible to do updates online. Initially it was a nice way to deliver an emergency fix if necessary, but you mostly continued to do nice QA'd releases every now and then.
But as with everything, when something becomes too easy it gets abused. So companies realized why do much QA (or any QA in extreme cases). Just push updates all the time and hope for the best, if customers (who are now QA) scream, push another update tomorrow. Break quick, fix quick.
It's mostly unsatisfactory if one values stable quality software.
Microsoft Corp. would like to have a word with you. /s
And it was much better than the current situation, if you ask me.
The first company I worked for in the 90s had a C codebase which seemed like it was half #ifdef or runtime version checks because they had to support customers who rarely updated except when they bought new servers, and that meant that if some version of SunOS, DOS/Windows/etc. had a bug you had to detect and either use a work around or disable a feature for years after the patch had shipped.
I do agree that stability, especially on the UI side, has serious value but my nostalgia is tempered by remembering how many people spent time recovering lost work or working around gaps in software which had been fixed years ago. I think the automatic update world is better on the whole but we need some corrective pressure like liability for vendors to keep people from pulling a Crowdstrike in their testing and reliability engineering.
I think that we went to another extreme. Because it's so easy to update, we just ship bad software saying "we'll fix it later". And we don't.
These days, it's just "ship it and we'll fix it later" instead, which is a big part of why (in my opinion) software quality has been declining for years.
I don't think the article itself holds up that well, it's just that updates are often a massive pain, one that you have to deal with somehow regardless. Realistically, LTS versions of OS distros and technologies that don't change often will lessen the pain, but not eliminate it entirely.
And even then, you'll still need to deal with breaking changes when you will be forced to upgrade across major releases (e.g. JDK 8 to something newer after EOL) or migrate once a technology dies altogether (e.g. AngularJS).
It's not like people will backport fixes for anything indefinitely either.
Never Update _Anything_ :)
In the old days we need CMS mostly because generating links and update to certain pages were expensive. Hard Disk were slow and memory were expensive. Now we have SSDs that eliminate 99.9999% of the problem.
Of course security updates are very hard, but if an old version has some good community you have the option of forking or upstreaming the updates yourself
For some languages and applications it can be trivial to backport the changes then trying to keep up with the new features. If it's tested and stable it will likely be more stable than a new version, I do this for some smaller programs and I'm not even a real programmer but more of a hobbyist
https://wiki.alpinelinux.org/wiki/Nginx
Also, may want to consider a flat html site if you don't have time to maintain a framework/ecosystem. =3
I did end up opting for Ubuntu LTS (and maybe the odd Debian based image here or there) for most of my containers because it essentially has no surprises and is what I run locally, so I can reuse a few snippets to install certain tools and it also has a pretty long EOL, at the expense of larger images.
Oddly enough, I also ended up settling on Apache over Nginx and even something like Caddy (both of which are also really nice) because it's similarly a proven technology that's good enough, especially with something like mod_md https://httpd.apache.org/docs/2.4/mod/mod_md.html and because Nginx in particular had some unpleasant behavior when DNS records weren't available because some containers in the cluster weren't up https://stackoverflow.com/questions/50248522/nginx-will-not-...
I might go for a static site generator sometime!
Ubuntu LTS kernels are actually pretty stable, but containers are still recommended. ;)
A bit off topic, but I rather enjoyed the idea behind mod_auth_openidc, which ships an OpenID Connect Relying Party implementation, so some of the auth can be offloaded to Apache in combination with something like Keycloak and things in the protected services can be kept a bit simpler (e.g. just reading the headers provided by the module): https://github.com/OpenIDC/mod_auth_openidc Now, whether that's a good idea, that's debatable, but there are also plenty of other implementations of Relying Party out there as well: https://openid.net/developers/certified-openid-connect-imple...
I am also on the fence about using mod_security with Apache, because I know for a fact that Cloudflare would be a better option for that, but at the same time self-hosting is nice and I don't have anything too precious on those servers that a sub-optimal WAF would cause me that many headaches. I guess it's cool that I can, even down to decent rulesets: https://owasp.org/www-project-modsecurity-core-rule-set/ though the OWASP Coraza project also seems nice: https://coraza.io/
Gets rid of 99.999% of problem traffic on APIs.
It is the most boring thing I ever integrated, and RabbitMQ has required about 3 hours of my time in 5 years. I like that kind of boring... ;)
Rate-limiting token-bucket firewall settings are a personal choice every team must decide upon (what traffic is a priority when choking bandwidth), and often requires tuning to get it right (must you allow mtu fragging for corporate users or have a more robust service etc.) These settings will also influence which events trip your fail2ban rule sets.
Have a great day, =)
Yes, it’s a static website. It’s amazing how little performance you actually need to survive a HN avalanche
I regularly snowboard with someone still at the company. They’re still on AngularJS.
AngularJS never dies.
AngularJS is forever.
The idea is to realise that there are two different classes of consumers who want different things, and rather than try to find a compromise that would not fully satisfy either group (and turns out to be more expensive to boot), we offer multiple release trains for different people.
One release train, called the tip, contains new features and performance enhancements in addition to bug fixes and security patches. Applications that are still evolving can benefit from new features and enhancements and have the resources to adopt them (by definition, or else they wouldn't be able to use the new features).
Then there are multiple "tail" release trains aimed at applications that are not interested in new features because they don't evolve much anymore (they're "legacy"). These applications value stability over everything else, which is why only security patches and fixes to the most severe bugs are backported to them. This also makes maintaining them cheap, because security patches and major bugs are not common. We fork off a new tail release train from the tip every once in a while (currently, every 2 years).
Some tail users may want to benefit from performance improvements and are willing to take the stability risk involved in having them backported, but they can obviously live without them because they have so far. If their absence were painful enough to justify increasing their resources, they could invest in migrating to a newer tail once. Nevertheless, we do offer a "tail with performance enhancements" release train in special circumstances (if there's sufficient demand) -- for pay.
The challenge is getting people to understand this. Many want a particular enhancement they personally need backported, because they think that a "patch" with a significant enhancement is safer than a new feature release. They've yet to internalise that what matters isn't how a version is called (we don't use semantic versioning because we think it is unhelpful and necessarily misleading), but that there's an inherent tension between enhancements and stability. You can get more of one or the other, but not both.
Even some JDK vendors can't resist offering those who want the comforting illusion of stability (while actually taking on more real risk) "tail patches" that include enhancements.
But they can't, because this is not a possibility that is given to them. All updates are put together, and we as an industry suck at even knowing if our change is backward compatible or not (which is actually some kind of incompetence).
And of course it's hard, because users are not competent enough to distinguish good software from bad software, so they follow what the marketing tells them. Meaning that even if you made good software with fewer shiny features but actual stability, users would go for the worse software of the competitor, because it has the latest damn AI buzzword.
Sometimes I feel like software is pretty much doomed: it won't get any better. But one thing I try to teach people is this: do not add software to things that work, EVER. You don't want a connected fridge, a connected light bulb or a connected vacuum-cleaner-camera-robot. You don't need it; it's completely superfluous.
Also for things that actually matter, many times you don't want them either. Electronic voting is an example I have in my mind: it's much easier to hack a computer from anywhere in the world than to hack millions of pieces of paper.
* Updates that add a theoretically independent feature, but which other software will dynamically detect and change their behavior for, so that it's not actually independent.
There was a time when Windows had a description for updates. Now the only distinction is between KB3587690 and KB67457770.
Author proceeds to add to two updates to the article, epic troll.
The current BSOD epidemic demonstrated the folly of mass concurrent versioning.
*nix admins are used to playing upgrade Chicken with their uptime scores. lol =)
Leaving something alone that works good is a good strategy. Most of the cars on the road are controlled by ECUs that have never had, and never will have any type of updates, and that is a good thing. Vehicles that can get remote updates like Teslas are going to be much less reliable than one not connected to anything that has a single extensively tested final version.
An OS that is fundamentally secure by design, and then locked down to not do anything non-essential, doesn't really need updates unless, e.g. it is a public facing web server, and the open public facing service/port has a known remote vulnerability, which is pretty rare.
I don’t think it’s necessarily this, but the fact that being able to update anytime is a great source of pressure to release untested software at any cost.
We used Microsoft office 2000 for 12 years. Never had to retrain people, deal with the weird ribbon toolbar, etc.
It's only the deranged use of OSs with ambient authority that gums up what would otherwise be stable systems.
Example: for my (personal) projects, I only use whatever is available in the debian repositories. If it's not in there, it's not on my dependency list.
Then enable unattended upgrades, and forget about all that mess.
The 2021/2022/2023/2024 version-numbering schemes are for applications, not libraries, because applications are essentially not ever semver-stable.
That's perfectly reasonable for them. They don't need semver. People don't build against jetbrains-2024.1, they just update their stuff when JetBrains breaks something they use (which can happen at literally any time, just ask plugin devs)... because they're fundamentally unstable products and they don't care about actual stability, they just do an okay job and call it Done™ and developers on their APIs are forced to deal with it. Users don't care 99%+ of the time because the UI doesn't change and that is honestly good enough in nearly all cases.
That isn't following semver, which is why they don't follow semver. Which is fine (because they control their ecosystem with an iron fist). It's a completely different relationship with people looking at that number.
For applications, I totally agree. Year-number your releases, it's much more useful for your customers (end-users) who care about if their habits are going to be interrupted and possibly how old it is. But don't do it with libraries, it has next to nothing to do with library customers (developers) who are looking for mechanical stability.
But let me tell you something: Long-Term Support software mostly doesn't pay well, and it's not fun either. Meanwhile some Google clown is being paid 200k to fuck up Fitbit or rewrite Wallet for the 5th time in the newest language.
So yeah. I'd love to have stable, reliable dependencies while I'm mucking around with the newest language de jour. But you see how that doesn't work, right?
The engineers are at most just complicit. Those who aren't are laid off or they quit on their own accord.
No one is paying such salaries for mundane clerical job.
Who wants to continue maintaining C++03 code bases without all the C++11/14/17/20 features? Who wants to continue using .NET Framework, when all the advances are made in .NET? Who wants to be stuck with libraries full of vulnerabilities and who accepts the risk?
Not really addressed is the issue of developers switching jobs/projects every few years. Nobody is sticking around long enough to amass the knowledge needed to ensure maintenance of any larger code base.
Which is caused by or caused the companies to also not commit themselves for any longer period of times. If the company expects people to leave within two years and doesn't put in the monetary and non-monetary effort to retain people, why should devs consider anything longer than the current sprint?
With the exception that in this hypothetical world we'd get backported security updates (addressing that particular point), who'd want something like this would be the teams working on large codebases that:
- need to keep working in the future and still need to be maintained
- are too big or too time consuming to migrate to a newer tech stack (with breaking changes in the middle) with the available resources
- are complex in of themselves, where adding new features could be a detriment (e.g. different code styles, more things to think about etc.)
Realistically, that world probably doesn't exist and you'll be dragged kicking and screaming into the future, once your Spring version hits EOL (or worse yet, will work with unsupported old versions and watch the count of CVEs increase, hopefully very few will find themselves in this set of circumstances). Alternatively, you'll just go work somewhere else and it'll be someone else's problem, since there are plenty of places where you'll always try to keep things up to date as much as possible, so that the delta between any two versions of your dependencies will be manageable, as opposed to needing to do "the big rewrite" at some point.That said, enterprises already often opt for long EOL Linux distros like RHEL and there is a lot of software out there that is stuck on JDK 8 (just a very visible example) with no clear path of what to do once it reaches EOL, so it's not like issues around updates don't exist. Then again, not a lot of people out there need to think about these things, because the total lifetime of any given product, project, their tenure in the org or even the company itself might not be long enough for those issues to become that apparent.
Perl has been stable for a couple of decades.
Even Perl 5 is rapidly evolving. They just added try/catch. Added a new isa operator. Added a new __CLASS__ keyword. Added defer blocks.
EMC had a system called Target Code which was typically the last patch in the second-last family. But only after it had been in use for some months and/or percentage of customer install base. It was common sense and customers loved it. You don’t want your storage to go down for unexpected changes.
Dell tried to change that to “latest is target” and customers weren’t convinced. Account managers sheepishly carried on an imitation of the old better system. Somehow from a PR point of view, it’s easier to cause new problems than let the known ones occur.
Well that's the first issue: downright malpractice. Developers should learn how to know (and test) whether it is a major change or not.
The current situation is that developers mostly go "YOLO" with semantic versioning and then complain that it doesn't work. Of course it doesn't work if we do it wrong.
For example, I avoid graphical commercial OS, large, graphical web browsers. Especially mobile OS and "apps".
Avoidance does not have to be 100% to be useful. If it defeats reliance on such software then it pays for itself, so to speak. IMHO.
The notion of allowing RCE/OTA for "updates" might allegedly be motivated by the best of intentions.
But these companies are not known for their honesty. Nor for benevolence.
And let's be honest, allowing remote access to some company will not be utilised 100% for the computer owner's benefit. For the companies remotely installing and automatically running code on other peoples' computer, surveillance has commercial value. Allowing remote access makes surveillance easier. A cake walk.
Now, some software, they effectively do this risk mitigation for you. Windows, macOS, browsers all do this very effectively. Maybe only the most cautious enterprises delay these updates by a day.
But even billion dollar corporations don't do a great job of rolling out updates incrementally. This especially applies as tools exist to automatically scan for dependency updates, the list of these is too long to name - don't tell me about an update only a day old, that's too risky for my taste.
So for OS and libraries for my production software? I'm OK sitting a week or a month behind, let the hobbyists and the rest of the world test that for me. Just give me that option, please.
This is required for some components, like, e.g., glibc or openssh, to stay secure-ish.
Other distros have this as well (Thumbleweed, Void, etc.), and I really think most people should not be using recently-deployed software. A small community using them however helps testing so the rest of us can have more stability. Which is why I don't recommend using Arch (or Debian unstable) for general users, unless you specifically want to help testing and accept the risk.
Also randomizing update schedules by at least a few hours does seem very wise (I don't think even the most urgent updates would make or break in say 6 hours of randomization?)
The issue with changelogs is that they are an honor system, and they don't objectively assess the risk of the update.
Comparing changes in the symbol table and binary size could give a reasonable red/yellow/green indicator of the risk of an update. Over time, you could train a classifier to give even more granularity and confidence.
I believe the chances of having a bricked laptop because of a bad update are higher than the chances of getting malware because running one or two versions behind the latest one.
you can definitely do that with python today: assemble a large group of packages that conver a large fraction of what people need to do, and maintain that as the 1 or 2 big packages. nobody's stopping you.
it's already being done!
Most companies I've worked for have the attitude of the author, they treat updates as an evil that they're forced to do occasionally (for whatever reason) and, as a result, their updates are more painful than they need to be. It's a self-fulfilling prophecy.
Nothing good would happen if some machine running Windows XP in a hospital that's hooked up to an expensive piece of equipment that doesn't run with anything else suddenly got connected to the Internet. Nor does the idea of any IoT device reaching past the confines of the local network make me feel safe, given how you hear about various exploits that those have.
On one hand, you should get security patches whenever possible. On the other hand, it's not realistic to get just security patches with non-breaking changes only. Other times, pieces of hardware and software will just be abandoned (e.g. old Android phones) and then you're on your own, even if you'd want to keep them up to date.
Nowadays I work in medical software and hospitals are running outdated unpatched Windows computers everywhere.
Nobody cares about updates. Almost nobody. I never saw Windows 11. Windows 10 is popular, but there are plenty of Vistas. I'm outright declining supporting Windows XP and we lost some customers over this issue.
My development tools are somewhat outdated, because compilers love to drop old Windows versions and 32-bit architectures, so sometimes I just can't update the compiler. For example I'm stuck with Java 8 for the foreseeable future, because Vista users are too numerous and it's not an option to drop them.
Hacker News is like another world. Yes, I update my computer, but everyone else does not. Even my fellow developers often don't care and just use whatever they got.
That is generally recognizable as stupidity. And the ones that did so are now paying the price.
* compliance tactics are very prone to fads. just look at cookie banners.
If Nix language was replaced with something sensible I'd jump back in excitedly.
ALL OTHER OSes/distro are still stuck in the '80s in that sense and most people seems even unable to understand. On storage alone the famous "rampant layer violation" and the absurdity of btrfs and stratis "against" zfs are really good examples of blind tech reactionary behaviors by high skilled people and their outcomes a showcase of why we damn need to innovate instead of shooting ourselves in the feet switching from something obsolete to something even worse (like the now-almost-finished full stack virtualization on x86 and thereafter the paravirtualization/container mania still current) layering crap on crap with more and more unmanageable infra with enormous attack surfaces.
Another small and relatively known example: Home Assistant project: apart of their design, they choose to distribute a python application as a GNU/Linux entire distro because to the such move seems to be commercially sound and many others choose to follow them instead of simply pip-install HA in a local venv, wasting an immense amount of resources on their system for what?
Such kind of tech evolution must end or we will collapse soon digitally speaking.
Or big-brain it and use Nix to build your containers and get the best of both worlds.
Yeah, me too. I also would like a few million bucks in the bank.
It's naive to think that every project wouldn't want to set this goal, simply because it's so unrealistic.
- Docker Swarm and Docker Compose use different parsers for `docker-compose.yaml` files, which may lead to the same file working with Compose but not with Swarm ([1]).
- A Docker network only supports up to 128 joined containers (at least when using Swarm). This is due to the default address space for a Docker network using a /24 network (which the documentation only mentions in passing). But, Docker Swarm may not always show error message indicating that it's a network problem. Sometimes services would just stay in "New" state forever without any indictation what's wrong (see e.g. [2]).
- When looking a up a service name, Docker Swarm will use the IP from the first network (sorted lexically) where the service name exists. In a multi-tenant setup, where a lot of services are connected to an ingress network (i.e. Taefik), this may lead to a service connecting to a container from a different network than expected. The only solution is to always append the network name to the service name (e.g. service.customer-network; see [3]).
- Due to some reason I still wasn't able to figure out, the cluster will sometimes just break. The leader loses its connection to the other manager nodes, which in turn do NOT elect a new leader. The only solution is to force-recreate the whole cluster and then redeploy all workloads (see [4]).
Sure, our use case is somewhat special (running a cluster used by a lot of tenants), and we were able to find workarounds (some more dirty than others) to most of our issues with Docker Swarm. But what annoys me is that for almost all of the issues we had, there was a GitHub ticket that didn't get any official response for years. And in many cases, the reporters just give up waiting and migrate to K8s out of despair or frustration. Just a few quotes from the linked issues:
> We, too, started out with Docker Swarm and quickly saw all our production clusters crashing every few days because of this bug. […] This was well over two years (!) ago. This was when I made the hard decision to migrate to K3s. We never looked back.
> We recently entirely gave up on Docker Swarm. Our new cluster runs on Kubernetes, and we've written scripts and templates for ourselves to reduce the network-stack management complexities to a manageable level for us. […] In our opinion, Docker Swarm is not a production-ready containerization environment and never will be. […] Years of waiting and hoping have proved fruitless, and we finally had to go to something reliable (albeit harder to deal with).
> IMO, Docker Swarm is just not ready for prime-time as an enterprise-grade cluster/container approach. The fact that it is possible to trivially (through no apparent fault of your own) have your management cluster suddenly go brainless is an outrage. And "fixing" the problem by recreating your management cluster is NOT a FIX! It's a forced recreation of your entire enterprise almost from scratch. This should never need to happen. But if you run Docker Swarm long enough, it WILL happen to you. And you WILL plunge into a Hell the scope of which is precisely defined by the size and scope of your containerization empire. In our case, this was half a night in Hell. […] This event was the last straw for us. Moving to Kubernetes. Good luck to you hardy souls staying on Docker Swarm!
Sorry, if this seems like like Docker Swarm bashing. K8s has it's own issues, for sure! But at least there is a big community to turn to for help, if things to sideways.
[1]: https://github.com/docker/cli/issues/2527 [2]: https://github.com/moby/moby/issues/37338 [3]: https://github.com/docker/compose/issues/8561#issuecomment-1... [4]: https://github.com/moby/moby/issues/34384