Increased error rate on Gov.uk on 5 January 2022
insidegovuk.blog.gov.uk
insidegovuk.blog.gov.uk
Appreciate the honesty from central government, but it's not a great look.
People simply fear updates will "break" working software and simply "delay" them. It doesn't mmatter if there are security updates, some people have the updates == bad mentality.
Version 8.x or 22.y that you get won't change in their minor releases in terms of behavior. They save the breakage for N+1, such as 9.x and 23.y
They go through a ton of effort to backport the security fixes. This keeps things generally way more predictable.
Problems do happen (eg: misaligned symbols from partial updates/restarts), but they're not as common as one would assume
Arch (or any other mostly-unpatched rolling distro, I expect—don’t know how, say, CentOS Stream does on the former point) requires constant tending, but every weekly or monthly update is pretty much painless. Sometimes one thing breaks, but usually with an accompanying warning, and if you put it there you know how to unbreak it. Upgrading and especially dist-upgrading Debian IME felt like you were running around putting out fires for a couple of hours every time, even if those upgrades were much rarer.
(It’s like the old joke about regular paradise vs college paradise: in the regular one everything’s peachy except a man comes every evening and hammers a single nail up your arse, in the college one everything’s wonderful for months but then a man comes with a bucket of nails and tells you it’s the end of the semester.)
I don’t know which ends up requiring more effort in the integral sense, and if that’s even the measure that matters, in particular on servers (even “pet” ones), I only know that I dread updating on a single Arch machine much less than I do a Debian one. I’m sure this won’t scale, but keeping a herd of like a homogeneous dozen servers on a rolling distro with a reproducible setup (Arch+aconfmgr, or hell, even NixOS unstable) doesn’t sound like an inherently stupid idea.
(If you need ABI stability, e.g. if you’re running a proprietary database, none of this applies, but in such a case a rolling distro is probably not an option anyway for other reasons.)
dist-upgrades like you mention are where I would actually expect a little bit of that discomfort you mention. That's [if memory serves] where the 'major' upgrades come in (eg: 8 -> 9) where we'd expect some surprises.
I hold the somewhat controversial opinion that CentOS Stream should be fine for most people, as you'd generally be hard pressed to notice differences between minor revisions [outside of compliance/certification efforts]
I think the smaller upgrades are especially painful on Apt/deb based distributions. Updates tend to require some attention, if nothing else to answer this question: "Keep the modified configs, or accept the vendor defaults?"
Those based on RPM tend to handle that case a little bit better in my experience. If a config is modified it's trusted, but the vendor configs are available as '.rpmnew' files. This way, if you find a problem - you have a way to find yourself back to comfort.
While Apt can be configured similarly, it doesn't come that way [as far as I'm aware, on the big representations - Debian/Ubuntu]
I can't say for other families as I haven't had quite the experience managing fleets of those!
None of this is to say it's a 'magic bullet' of sorts. Some discipline is certainly still rewarded!
There's usually tiers of environments where you try it out with fairly low stakes [such as development/integration areas], then you work your way up the pipeline to upgrading the thing that matters.
Another approach is to simply replace the old with new. Depending on the elasticity of the environment you may favor one over the other.
I think that this point of view is an unfortunate reality due to the world that we live in. Everyone would like security updates, but only as long as they're guaranteed not to actually break anything else.
I've had regular Debian updates break GRUB and prevent a server from rebooting: https://blog.kronis.dev/everything%20is%20broken/debian-and-...
(admittedly, i think i had the full unattended upgrades enabled, not just the security ones, but that's still a pretty catastrophic failure)
Similarly, when i finally wanted to upgrade my install of GitLab to a newer version, it also failed spectacularly rendering the whole instance inoperable: https://blog.kronis.dev/everything%20is%20broken/gitlab-upda...
(there is no option to get "just" the security updates, you need a new major version once the old one is EOL)
I actually did a writeup since, where i migrated over to Gitea, Nexus and Drone, because even if future updates will break them, the scope will be much more manageable: https://blog.kronis.dev/articles/goodbye-gitlab-hello-gitea-...
My takeaway there was pretty similar to what you're alluding to:
> Also, something else that i just now realized was that my GitLab instance was so out of date, because updating it was a painful process - the data directories that i'd need to manually back up (in addition to automated backups which are done every other day) before an update would take a lot of space, the container images are large and download slowly, the changelogs are long and configuration changes plentiful (the Unicorn to Puma migration comes to mind) and as my other post shows, a lot can go wrong and it be a very frustrating process.
> In short: if you want your users to update more often and stay safe, make updating an easier and less scary process. Of course, in the case of GitLab, i don't think that's something that's necessarily easy to do, since their Omnibus install is about as close as you can get to that, but even that had problems, at least in my particular case. Sometimes you just need an alternative or a few other pieces of software to do more or less the same thing, and that's okay.
I've actually written about the difficulties with updates and how problematic they become when running a non-trivial amount of software before as well, in a slightly tongue in cheek manner: https://blog.kronis.dev/articles/never-update-anything
Apologies if that's quite a few links, it's just that i think that you're offering only a part of the argument and it should be expanded to better reflect the reality for some people. Updates are feared because in practice they absolutely do bring problems and you cannot expect that ahead of time and will have to deal with it sooner or later, so not only do you need to evaluate how to upgrade everything (including configuration and other possibly breaking changes), but also have your backups and failover systems ready ahead of time.
With how much software out there doesn't have proper test or staging environments and how often it's not easy to spin up a new environment, it's completely understandable that people are going to hold out for as long as they can, which leads to some pretty bad situations in practice. Of course, the need for any sort of a SLA and uptime only makes this much worse.
They could choose to not specify tags and run latest - then a power outage might cause an update by consequence of a fresh pull
This government really doesn't care about its image
If you push back and you get overridden, fine. But I doubt that most devs ever do that.
From a different domain, here is an excellent blog post from 2019 on how to write clear health information for people with varying levels of literacy. (And no, it doesn't mean "dumbing down" your writing.)
Pee and poo and the language of health: https://digital.nhs.uk/blog/transformation-blog/2019/pee-and...
https://technology.blog.gov.uk/2016/09/19/why-we-use-progres... https://news.ycombinator.com/item?id=12538144
Kudos to whoever was involved in making this outage report happen.
As a business + personal user over 15 years, I've noticed The UK's gov.uk websites have been improving at the fastest rate I've seen over the last 3-4 years in particular.
Does anyone know why this is? (Or perhaps you have had a different experience).
I did giggle a bit that they basically turned it off and on again though.