We were discharged at midnight by the doctor, the nurse didn't come into our exam room to tell us until 4am. I can't imagine the mess this has caused.
Or, we could build robust systems that can tolerate indefinite down time. Might cost more, might need more staff.
Pick one. I’ll always pick the one that saves human lives when systems go down.
2. Hospitals should not have executives.
3. Hospitals should be community funded with backstop by the federal government.
4. PE is a cancer - let the doctors treat it.
... and seconding all the best wishes for the mother involved. Do get well.-
Wishes for a speedy recovery to your mom!
I hope no one uses such single point of failure systems anymore. Especially CS. The same is applicable for Cloudflare as well! But at least, the systems will be functioning standalone and accessible in their case and could cause only netwide outage! (i.e., if the CF infra goes down!)
Anyways, who knows what is going to happen with such widespread vendor dependency?
The world gets reminded about the Supply Chain Attacks every year which is a good (but a scary) one that definitely needs some deep thinking...
Up for it?
That's an extra 4 hours of emergency room fees you ideally wouldn't have to pay for.
if it works, there's a lot more to be done to get the patient to stable.
Different medications may be pushed (injected into the patient) to help stabilize them. These medications are recorded via a bar code and added to the patients chart in Epic. Epic is the source of truth for the current state of the patient. So if that is suddenly unavailable that is a big problem.
Technologically it seems doable. Big enough order brings down the costs.
https://soldered.com/product/soldered-inkplate-5-5-2%e2%80%b...
Of course the real backup plan should be designed based on the actual needs, perhaps the whole system needs an "offline mode" switch. I assume they already run things locally, in case the big cable seeker machine arrives in the neighborhood.
Some hospitals require you to input this in order to even get physical access to the medications.
Although a crash cart would normally have common things necessary to save someone in an emergency, so I would think that if someone was truly dying they could get them what they needed. But of course there are going to be exceptions and a system being down will only make the process harder.
Yep, the business continuity boxes are basically minimally connected PDF archives of patient records "printed" multiple times a day.
Anyone involved in designing and/or deploying a system where an application outage threatens life safety, should be charged with criminal negligence.
A receipt printer in every patient room seems like a reasonable investment.
The grandparent indicated that the problem was that when all tow computers went down, they couldn’t look up what had already been done for the patient. I suggested a simple solution for that - receipt printers.
After the computers fail you tape the receipt to the wall and fall pack to pen and paper until the computers come back up.
I completely understand the scale of the outage today. I am saying that it was a stupid decision and possibly criminally negligent to make a life critical process dependent on the availability of a distributed IT application not specifically designed for life critical availability. I strongly stand by that POV.
But they switched over to EMR because the advantages of Pyxis[1] in getting the right medications to the right patients at the right time- and documenting all of that- are so large that for patient safety reasons alone it wins out over paper. You can fall back to paper, it's just a giant pain in the ass to do it, and then you have to do the data entry to get it all back into EMR's. Like my wife, who was working last night when everyone else in her department got Crowdstrike'd, she created a document to track what she did so it could be transferred into EMR's once everything comes back up. And the document was over 70 pages long! Just for one employee for one shift.
1: Workflow: Doctor writes prescription in EMR. Pharmacist reviews charts in EMR, approves prescription. Nurse comes to Pyxis cabinet and scans patient barcode. Correct drawer opens in cabinet so the proper medication- and only the proper medication- is immediately available to nurse (technicians restock cabinet when necessary). Nurse takes medication to patient's room, scans patient barcode and medication barcode, administers drug. This system has dramatically lowered the rates of wrong-drug administration, because the computers are watching over things and catch humans getting confused on whether this medication is supposed to go to room 12 or room 21 in hour 11 of their shift. It is a great thing that has made hospitals safer. But it requires a huge amount of computers and networks to support.
Why would a Pyxis cabinet run Windows? I realize Windows isn't even necessarily at fault here, but why on earth would such a device run Windows? Is the 90s form of mass incompetence in the industry still a thing where lots of stuff is written for Windows for no reason?
Why does everything need the internet?
Why would you ever need to move a patient from one hospital room containing one set of airgapped computers into another, containing another set of airgapped computers?
Why would you ever need to get information about a patient (a chart, a prescription, a scan, a bill, an X-Ray) to a person who is not physically present in the same room (or in the same building) as the patient?
And sending data out can be done quite securely. Then replies could be highly sanitized or kept on specific machines outside the air gap.
And now you've added an army of people running around moving USB sticks, or worse, printouts and feeding them into other computers.
It's madness, and nobody wants to do it.
In my eyes, there is a technical solution therr that keeps friction low for hospital staff: network stuff, on an internet, but not The Internet...
Edit: I've since been reading the other many many comment threads on this HN post which show the reasons why so much stuff in healthcare is connected to each other via good old internet, and I can see there's way more nuance and technicality I am not privy to which makes "just connect LANs together!" less useful. I wasn't appreciating just how much of medicine is telemedicine.
Yes there will be some pain, but the alternative is what we have right now.
> nobody wants to do it.
Tough luck. There's lots of things I don't want to do.
Just so I understand what you are saying you are proposing that we drown our hospital rooms in paper receipt constantly. In the off chance the computers go down very rarely?
Do you see any possible drawbacks with your proposed solution?
> possibly criminally negligent to make a life critical process dependent on the availability of a distributed IT application
What process is not “life critical” in a hospital? Do you suggest that we don’t use IT at all?
This much of patients state will be carried on their wrist. Maybe for complex cases you need two stickers. Have to be judicious in encoding data, maybe just last 48 hours.
Handheld qr readers, off line that read and display QR data strings.
I'm not even beyond two degrees of seperation here. I don't think a court'll have trouble navigating it.
Cloudstrike did not even have a duty of care to their customer, let alone their customer’s customer (speaking for my jurisdiction, of course).
Now, a tort is less of a stretch than a crime, but thank goodness I’m not a lawyer so I don’t have to figure out what circumstances apply and how much liability the TOS and EULAs are able to wash away.
If I go to Home Depot to buy rope for belaying at my rock climbing center and someone falls, breaks the rope and dies, then I am on the hook for manslaughter.
Not the rope manufacturer, who clearly labeled the packaging with "do not use in situations where safety can be endangered". Not the retailer, who left it in the packaging with the warning, and made no claim that it was suitable for a climbing safety line. But me, who used a product in a situation where it was unsuitable.
If I instead go to Sterling Rope and the same thing happens, fault is much more complicated, but if someone there was sufficiently negligent they could be liable for manslaughter.
In practice, to convict of manslaughter, you would need to show an individual was negligant. However, our entire industry is bad at our job, so no individual involved failed to perform their duties to a "reasonable" standard.
Software engineering is going to follow the path that all other disciplines of meatspace engineering did. We are going to kill a lot of people; and every so often, enough people will die that we add some basic rules for safety critical software, until eventually, this type of failure occuring without gross negligence becomes nearly unthinkable.
Then again, the management that put this in are probably also the same idiots that insist on a 7 day lead time CAB process to update a typo on a brochure ware website "because risk".
There are plenty of drugs that can only be given in certain quantities over a certain period of time, and if you go beyond that, it makes the patient worse not better. Similarly there are plenty of bad drug interactions where whether you take a given course of action now is directly dependent on which drugs that patient has already been given. And of course you need to monitor the patient's progress over time to know if the treatments have been working and how to adjust them, so if you suddenly lose the record of all dosages given and all records of their vital signs, you've lost all the information you need to treat them well. Imagine being dropped off in the middle of nowhere, randomly, without a GPS.
Will outages like this motivate a backup paper process? The automated process should save enough information on paper so a switch over to paper process at any time is feasible. Similar to elections.
there are backup paper processes, but they start fresh when the systems go down
If it was printing paper in case of downtime 24/7, it would be massive wasteage for the 99% of time system is up
Maybe a handheld device for scanning in drugs or entering procedure information that stores the data locally which can then be synced with a larger device with more storage somewhere that is also 100% local and immutable which then can sync to online systems if that is needed.
If medical data were synced to the cloud but also stored on the endpoint devices and local servers, you’d have more redundancy. Obviously much more complexity to it but that’s what it would take. Epic as single source of truth means everyone is screwed when it is down. This is the trade off that’s been made.
That's a recipe for a different kind of disaster. I actually used Google Keep some years ago for medical data at home — counted pills nightly, so mom could either ask me or check on her phone if she forgot to take one. Most of the time it worked fine, but the failure modes were fascinating. When it suddenly showed data from half a year ago, I gave up and switched to paper.
More seriously, we need better purpose build medical computing equipment, that runs on it's own OS, and only has outbound network connectivity for updating other systems.
I also think of things like the old school "check list boards" that used to be literally built into the yolk of the airplane they were made for.
That doesn't help when the system goes down and you lose the record of all medications administered prior to having to switch over to the Sharpie.
Reality is complex.
And that's in the OR, where vitals are automatically captured. There just aren't enough computers to do real-time electronic documentation, and even if there were there wouldn't be enough space.
Its easier, faster, and more accurate than writing in my experience. We have a page solely dedicated to codes and the most common interventions. Got IO? I press a button and its documented with timestamp. Pushing EPI, button press with timestamp. Dropping an I-Gel or Intubating, button press... you get the idea.
The details of the interventions can be documented later along with the narrative, but the bulk of the work was captured real-time. We can also sync with our monitors and show depth of compressions, rate of compressions and rhythms associated with the continuous chest compression style CPR we do for my agency.
Going back to paper for codes would be ludicrous for my department. The data would be shit for a start. Hand writing is often shit and made worse under the stress of screaming bystanders. Depending on whether we achieved ROSC or not would increase the likelihood of losing paper in the shuffle
This wisdom is echoed in some religious practices that avoid complete reliance on modern technology.
Okay, how does that monitor work? Genuinely curious.
Nurses hated it.
If they go and publish "According to hackernews user davycro ..." _then_ there's a problem.
It makes my blood boil to be honest that there is no liability for what software has become. It's just not acceptable.
Companies that produce software with the level of access that Crowdstrike has (for all effective purposes a remote root exploit vector) must be liable for the damages that this access can cause.
This would radically change how much attention they pay to quality control. Today they can just YOLO-push barely tested code that bricks large parts of the economy and face no consequences. (Oh, I'm sure there will be some congress testimony and associated circus, but they will not ever pay for the damages they caused today.)
If a person caused the level and quantity of damage Crowdstrike caused today they would be in jail for life. But a company like Crowdstrike will merrily go on doing more damage without paying any consequence.
What about companies that deploy software with the level of quality that Crowdstrike has? Or Microsoft 365 for that matter.
That seems to be the bigger issue here; after all Crowdstrike probably says it is not suitable for any critical systems in their terms of use. You shouldn't be able to just decide to deploy anything not running away fast enough on critical infrastructure.
On the other hand, Crowdstrike Falcon Sensor might be totally suitable for a non-critical systems, say entertainment systems like the Xbox One.
Local emergency services were basically nonfunctioning for better part of the day along with the heat wave and various events, seems like a number of deaths (locally at least, specific to what I know for my mid sized US city) will be indirectly attributable to this.
Despite this horrific outage, in the end it sounds like a much better and anti-fragile system than a government telling people how to do things.
They absolutely should be liable for the losses, in each case where they caused it.
(Which is most of them. Most companies install crowdstrike because their auditor want it and their insurance company says they must do whatever the auditor wants. Companies don't generally install crowdstrike out of their own desire.)
But of course they will not pay a single penny. Laws need to change for insurance companies, auditors and crowdstrike to be liable for all these damages. That will never happen.
I have worked at places which controlled the roll-out of new security updates (and windows updates) for this very reason. If you invest enough in IT is possible. But you have to have a lot of money to invest in IT to have people good enough to manage it. If you can get SwiftOnSecurity to manage your network, you can have that. But can every hospital, doctor's office, pharmacy, scan center, etc. get top tier talent like SwiftOnSecurity?
Healthcare IT is very important, because computers are good at record-keeping, retrieval and storage, and that's a huge part of healthcare.
When it came to audit time, the auditors were always impressed that our team had better timely updates than the corporate office side of things.
I never really thought we were doing anythin all that special (in fact, there were always many things I wanted to improve anout the process) but reading about this issue makes me think that maybe we really were just that much better than the average IT shop?
But did they also control the roll-out of virus/threat definition files? Because if not their goose would have been still cooked this time.
If, for example, they were doing slow rollouts for configs in addition to binaries, they could have caught the problem in their canary/test envs and not let it proceed to a full blackout.
The fact that ransomware is still a concern is an indication that we've failed to update our IT management and design appropriately to account for them. We took the cheap way out and hoped a single vendor could just paper over the issue. Never in history has this ever worked.
Also speaking of generators a large enough hospital should be running power failure test events periodically. Why isn't a "massive IT failure test event" ever part of the schedule? Probably because they know they have no reasonable options and any scale of catastrophe would be too disastrous to even think about testing.
It's a lesson on the failures of monoculture. We've taken the 1970s design as far as it can ago. We need a more organically inspired and rigorous approach to systems building now.
Granted, you are still vulnerable of physical attacks (i.e. the person coming with an USB stick) but I would say much more difficult, and if you put firewalls also between compartment of internal networks even difficult.
Also, I think the use of Windows in critical settings is not a good choice, and to me we had a demonstrations. For who says the same could have happened to Linux, yes but you could have mitigated it. For example, to me a Linux system used in critical settings shall have a root read-only root filesystem, on Windows you can't. Thus the worse you would had is to reboot the machine to restore it.
Segmenting your internal network is a good defence against lots of attacks, to limit the blast radius, but it's hard and expensive to do a lot of it in corporate environments.
2. Why would anyone trust a ransomware perpetrator to honor a deal to not reveal or exploit data upon receipt of a single ransom payment? Are organizations really going to let themselves be blackmailed for an indefinite period of time?
3. I'm unconvinced that crowdstrike will reliably prevent sensitive data exfiltration.
2. Appearently yes. Why do you think calls to ban payments exist?
3. At minimum it raises the bar for the hackers - sure, it's not like you can't bypass edr but it's much easier if you don't have to bypass it at all because it's not there
crowsdstrike is not a DLP solution. You can solve that problem (where necessary) by less intrusive means.
*Ok ok I know it's bypassable but one of the happy paths for an attack is to pivot to the machine that doesn't have edr and continue from there.
This is part of the problem too. These insurance/audit companies need to be made liable for the damage they themselves cause when they require insecure attack vectors (like Crowdstrike) to be installed on machines.
https://www.reddit.com/r/debian/comments/1c8db7l/linuximage6...
And even worse, possibly quite a few deaths as well.
I hope (although I will not be holding my breath), that this is the wake-up call we need to realise that we cannot have so much of our critical infrastructure rely on the bloated OS of company known for its buggy, privacy-intruding, crapware riddled software.
I'm old enough to remember the infamous blue-screen-of-death Windows 98 presentation. Bugs exist but that was hardly a glowing endorsement of high-quality software.. This was long ago, yet it is nigh on impossible to believe that the internal company culture has drastically improved since then, with regular high-profile screw-ups reminding us of what is hiding under the thin veneer of corporate of respectability.
Our emergency systems don't need windows, our telephone systems don't need windows, our flight management systems don't need windows, our shop equipment systems don't need windows, our HVAC systems don't need windows, and the list goes on, and on, and on.
Specialized, high-quality OSes with low attack surfaces are what we need to run our systems. Not a generic OS stuffed with legacy code from a time when those applications were not even envisaged.
Keep-it-simple-stupid -KISS-is what we need to go back to, our lives literally depend on it.
With the mutli-billion dollars screw-up that happened yesterday, and an as-of-yet unknown number of deaths, it's impossible to argue that the funds are unavailable to develop such systems. Plurality is what we need, built on top of strong standards for compatibility and interoperability.
Perhaps rather than an indictment on Windows, this is a call to re-evaluate microkernels, at least for critical systems and infrastructure.
And building something around microkernels would definitely not be a bad starting point.
What does this mean? Did the power go down? Is all the equipment connected? Or is it the insurance software that can't run do nothing gets done? Maybe you can't access patient files anymore but is that taking down the whole thing?
Have a family member in crit care who was getting a sepsis workup on a patient when this all happened. They somehow got plain film working offline after a bit of effort.
Or maybe I'm giving these institutions too much credit?
If the alternatives/plan b's were as good or better than the plan a's then they wouldn't be the alternatives. Nobody is going to have half a hospital's care capacity sit as backup when they could use that year round to better treat patients all the time, they just have plans of last resort to use when what they'd like to use isn't working.
(worked healthcare IT infrastructure for a decade)
Seems like a possible plan would be duplicate computer systems that are using last week's backup and not set to auto-update. Doesn't cover you if the databases and servers go down (unless you can have spares of those too), but if there is a bad update, a crypto-locker, or just a normal IT failure each department can switch to some backups and switch to a slightly stale computer instead of very stale paper.
Money spent on spares is not spent on cares.
The "downtime" computers were affected just like everything else because there was no network.
Phones are all IP-based now; they didn't work.
Couldn't check patient histories, couldn't review labs, etc. We could still get drugs, thankfully, since each dispensing machine can operate offline.
I worked for a company that sold and managed medical radiology imaging systems. One of our customers' admins called and said "Hey, new scans aren't being properly processed so radiologists can't bring them up in the viewer". I told him I'd take a look at it right away.
A few minutes later, he called back; one of their ERs had a patient dying of a gunshot wound and the surgeon needed to get the xray up so he could see where the bullet was lodged before the guy bled out on the table.
Long outages are terrifying, but it only takes a few minutes for someone to die because people didn't have the information they needed to make the right calls.
You don't need 4 years of specialized training to see a bullet on a scan.
X-Ray has limitations though - most of our emergencies aren't as easy to diagnose as bullets or pneumonia. CT, CTA, and to a lesser extent MRI are really critical in the emergency department, and you definitely need four years of training to interpret them, and a computer to let you view the scan layer-by-layer. For many smaller hospitals they may not have radiology on-site and instead use a remote radiology service that handles multiple hospitals. It's hard to get doctors who want to live near or commute to more rural hospitals, so easier for a radiologist to remotely support several.
One day, during a rapid pediatric patient intervention, a caregiver tried to log in to a PC to check a drug interaction. The computer took a long time to log in because of a VDI problem where someone had stored many images in a file that had to be copied on login. While the care team was waiting for the computer, an urgent decision was made to give the drug. But a drug interaction happened — one that would have been caught, had the VDI session initialized more quickly.
The patient died and the person whose VDI profile contained the images in the bad directory committed suicide. Two lives lost because files were in the wrong directory.
You come in for you next shift and are finishing charting from your prior shift. You open one of your partially finished charts and a little popup tells you "you are editing the chart for a deceased patient".
This is why I'm impressed by anyone who works in a hospital, especially the more urgent/intensive care
They could as well launch that app in OpenBSD.
I’m no fan of Windows or Microsoft but the commitment to backwards compatibility should not be underestimated.
Instead, it's a bunch of independent-ish, for-profit software & hardware companies. Each one trying to make it cheap & easy to develop their own product, and to maximize sales. Given the dominance of MS-DOS and Windows on cheap-ish & ubiquitous PC's, starting in the early-ish 1980's, the current situation was pretty much inevitable.
The big health products are built on windows because they are built by outsourced software shops and target the majority of builds which are basically the equivalent of bob's hardware store still running windows 95 on their point of sale box.
The major players that took over this space for the big players had to migrate from this, so they still targeted "wintel" platforms because the vast majority of healthcare servers are windows.
Its basically the tech equivalent of everything evolved from the width of oxen for railway.
What are the hard problems? I can think of a few, but I'm probably wrong.
Dependency chains: many pieces of kit either only have drivers on windows or work much better on Windows. You are at the mercy of the least OS diverse piece of kit. Label printers are notorious for this as an e.g.
Staffing: Many of your staff know how to do their jobs excellently, but will struggle with tech. You need them to be able assume a look and feel, because you dont want them fighting UX differences when every second counts. Their stress level is roughly equiv. to their worst 10 seconds of their day. And staff will quit or strike over UX. Even UI colour changes due to virtualization down scaling have triggered strife.
Change Mgmt: Hospitals are conservative and rarely push the envelope. We are seeing a major shift at the moment in key areas (EMR) but this still happening slowly. No one is interested in increasing their risk just because Linux exists and has Win64 compatability. There is literally no driver for change away from windows.
(Not including this colossal fuck up.)
Billing and insurance reimbursement process change all the time and is a headache to keep up to date. E.g. the actual dentist software is paint but with mainly the bucket and some way to quickly insert teeth objects to match your mouth. I.e. almost no medical skill in the software itself helping the user.
Question is: why half+ of Fortune 500 companies allowed Crowdstrike - Windows hackers - access and total control of their not-a-ms-windows business ? Obviously Crowdstrike do not do medicine or lifting cranes differentiation. "In the middle of the surgery" is not in their use case docs!
There was somewhere Mercedes pitstop image with wall of BSoD monitors :) But that is not Crowdstrike business either...
And all that via public internet and misc clouds. Banks have their own fibre lines, why hospitals can't?
Airports should disconnect from Internet too, selling tickets can be separate infra, synchronization between POSes and checkout don't need to be in real time.
There is only one sane way to prevent such events: EOD controlled by organization and this is sharply incompatible with 3rd party on-line EOD providers. But they can sell it in a box and do real time support when called.
Very easy for us to second guess today of course. But in another scenario a manager is being torn a new one because they fell victim to a ransomware attack via a zero day systems were left vulnerable to because Crowdstrike wasn’t updated in a timely manner.
Mostly, if you are reasonably timely about keeping updates applied, you're fine.
Sure. And Crowstrike releasing an update that bricks machines is also not the normal case. We're debating between two edges cases here, the answers aren’t simple. A zero day spreading like wildfire is not normal but if it were to happen it could be just as, if not more, destructive than what we’re seeing with Crowdstrike.
This is beyond hospital IT control. Clownstrike (sorry, Crowdstrike) unconditionally force-updates the hosts.
As to why they didn't catch this during tests or why they don't use perform gradual change rollouts to hosts, your guess is as good as mine. I hope we get a public postmortem for this.
[1]https://www.crowdstrike.com/blog/statement-on-falcon-content...
It says that if a system isn’t “affected”, meaning it doesn’t reboot in a loop, then the “protection” works and nothing needs to be done. That’s because the Crowdstrike central systems, on which rely the agents running on the clients’ systems, are working well.
The “sensor” is what the clients actually install and run on their machines in order to “use Crowdstrike”.
The crash happened in a file named csagent.sys which on my machine was something like a week old.
(1) Entire system is crashed.
(2) System is running AND protected from security threats by Falcon Sensor.
And to mean that this is not a possible state:
(3) System is running but isn't protected by Falcon Sensor.
In other words, I interpreted it to mean that they're trying to reassure people they don't need to worry about crashes and hacks, just crashes.
Even if they did switch, they'd then want to install all the equivalent monitoring crap. If such existed, it would likely be some custom kernel driver and it could bring a unix system to its knees when shit goes wrong too.
The fact that she was discharged without an overnight admit suggests to me that the MRI did not show a stroke, or perhaps she was outside the treatment window when she went to the hospital.
With that said, Microsoft could've done this with Defender just as easily, so be mindful of system diversity in your business continuity and disaster recovery plans and enterprise architecture. Heterogeneous systems can have inherent benefits.
https://apnews.com/article/tech-outage-crowdstrike-microsoft...
> “This is a function of the very homogenous technology that goes into the backbone of all of our IT infrastructure,” said Gregory Falco, an assistant professor of engineering at Cornell University. “What really causes this mess is that we rely on very few companies, and everybody uses the same folks, so everyone goes down at the same time.”
The irony is the NHS likely installed CrowdStrike as a direct reaction to WannaCry.
At large-scale, you don’t solve problems, you only replace them with smaller ones.
If you want actually good payout, your crypto locker has to either encrypt network filesystems, or infect crucial core systems (domain controllers, database servers, the filers directly, etc).
Ransomware getting smarter about sideways movement, and proper data exfiltration etc attacks, are part of what led to proliferation of requirements for EDRs like Crowdstrike, btw
But that's besides the point. Point is, attacks distributed over time and space ultimately make the overall system more resilient; an attack happening everywhere at once is what kills complex systems.
> Ransomware getting smarter about sideways movement, and proper data exfiltration etc attacks, are part of what led to proliferation of requirements for EDRs like Crowdstrike, btw
To use medical analogy, this is saying that the pathogens got smarter at moving around, the immune system got put on a hair trigger, leading to a cytokine storm caused by random chance, almost killing the patient. Well, hopefully our global infrastructure won't die. The ultimate problem here isn't pathogens (ransomware), but the oversensitive immune system (EDRs).
These security software vendors have found a wonderful tacit moat: they have managed to infect various questionnaire templates by being present in a short list of "pre-vetted and known" choices in a dropdown/radiobutton menu. If you select the sane option ("other"), you get to explain to technically inept bean counters why you did so.
Repeat that for every single regulator, client auditing team, insurance company, etc. ... and soon enough someone will decide it's easier and cheaper to pick an option that gets you through the blind-leading-the-blind question karaoke with less headaches.
Remember: vast majority of so-called security products are sold to people high up in the management chain, but they are inflicted upon their victims. The incentives are perverse, and the outcomes accordingly predictable.
Tell them it’s for preserving diversity in the field.
For anyone browsing the thread archive in the future: you can have that quip in your back pocket and use it verbally when having to discuss the bingo sheet results with someone competent. It's a good bit of extra material, but it can not[ß] be your sole reason. The term you do want to remember is "additional benefit".
The reasons you actually write down boil down to four things. High-level technical overview of your chosen solution. Threat model. Outcomes. And compensating controls. (As cringy as that sounds.)
If you can demonstrate that you UNDERSTAND the underlying problem, and consider each bingo sheet entry an attempt at tackling a symptom, you will be on firmer ground. Focusing on threat model and the desired outcomes helps to answer the question, "what exactly are you trying to protect yourself from, and why?"
ß: I face off with auditors and non-technical security people all the time. I used to face off with regulators in the past. In my experience, both groups respond to outcome-based risk modeling. But you have to be deeply technical to be able to dissect and explain their own questions back to them in terms that map to reality and the underlying technical details.
This should be the standard for any life sustaining or surgical systems, and any critical weapons systems.
Most of what I do is creating the tools to let the field reps go into hospitals and update capital equipment in a disconnected state (IE, the reps must be physically tethered to the device to interact with it). The fact that any critical equipment would get an auto-update, especially mid-surgery is incredibly bad practice.
Has to be wifi because the carts the nurses use roll around. Has to be networked so you can have EMR's that keep track of what your patients have gotten and the Pharmacists, doctors, and nurses can interface with the Pyxis machines correctly. The nurse scans a patients barcode at the Pyxis, the drawer opens to give them the drugs, and then they go into the patient's room and scan the drug barcode and the patients barcode before administering the drug. This system is to prevent the wrong drug from being administered, and has dramatically dropped the rates of mis-administering drugs. The network has to be everywhere on campus (often times across many buildings). Then the doctor needs to see the results of the tests and imaging- who is running around delivering all of these scans to the right doctors?
You don't know what you are talking about if you think this is easy.
If the system has to be networked with the outside world, who is responsible for physically updating all of these machines, so they don't get ransomware'd? Who has to go out and visit each individual machine and update it each month so the MRI machine doesn't get bricked by some teen ransomware gang? Remember that was the main threat hospitals faced 3-4 years ago, which is why Crowdstrike ended up on everyone's computer: because the ransomware insurance people forced them too.
There is a reason that I am a software engineer and not an IT person. I prefer solving more tractable problems, and I think proving p!=np would be easier than effectively protecting a large IT network for people who are not computing professionals.
One of my favorite examples: in October 2013 casino/media magnate and right wing billionaire Sheldon Adelson gave a speech about how the US and Israel should use nuclear weapons to stop Iran nuclear program. In February 2014 a 150 line VB macro was installed on the Sands casino network that replicated and deleted all HDDs, causing 150 million dollars of damage. That was to a casino, which spends a lot of money on computer security, and even employs some guys named Vito with tire irons. And it wasn't nearly enough.
The manufacturer does. As I mentioned in my OP I help build the software for our field reps to go into hospitals and clinics to update our devices in a disconnected state. Most of the critical equipment we manufacture has this as a requirement since it can't be connected to a network for security reasons.
As for discharge orders, etc, I can't speak to that, but that's also not what I would consider critical. I'm talking about things like surgical robots, which can not be connected to a network for obvious reasons, especially during a surgery.
Anyway: it is getting to the point that I cynically predict we may be required to add things to the system (such as embedding PCs), just so we can turn around and "secure" them to comply with the requirements that shouldn't be applied to these systems. Maybe this current outage event will be a wake up call to how misplaced the priorities are, but I doubt it.
... at which point you will lose battles to enemies who have successfully networked their command and control operations. (For extra laughs, just wait until this is also true of AI.)
Ultimately there are just too darned many advantages to connecting, automating, and eventually 'autonomizing' everything in sight. It sucks when things don't go right, or when a single point of failure causes a black-swan event like this one, but in an environment where you're competing against either time or external adversaries, the alternatives are all worse.
A better thing to do is do phased deployment, so you can see if an update will cause issues in your environment before pushing it to all systems. As this incident shows, you can’t trust a software vendor to have done that themselves.
I've mostly resigned myself today to deploying the configuration change and watching for anomalies in my monitoring for a number of hours or days afterward, but I acknowledge that I also have both a process supervisor that will happily let me crash loop my programs and deployment infrastructure that will nonetheless allow me to roll things back. Without either of those, I'm honestly at a loss as to how I'd safely operate this product.
# Update A
## config.ext
foo = false
## src.py
from config import config
if config('foo'):
work(2 / 0)
else:
work(10 / 5)
"Yep, we rigorously tested it." # Update B
## config.ext
foo = true
"It's just a config change, let's go live."The most insidious part of this is when there are entire swaths of infrastructure in place that circumvent the usual code review process in order to execute those configuration changes. Boolean flags like your `config('foo')` here are most common, but I've also seen nested dictionaries shoved through this way.
> Adama: It's an integrated computer network, and I will not have it aboard this ship.
> Roslin: I heard you're one of those people. You're actually afraid of computers.
> Adama: No, there are many computers on this ship. But they're not networked.
> Roslin: A computerized network would simply make it faster and easier for the teachers to be able to teach--
> Adama: Let me explain something to you. Many good men and women lost their lives aboard this ship because someone wanted a faster computer to make life easier. I'm sorry that I'm inconveniencing you or the teachers, but I will not allow a networked computerized system to be placed on this ship while I'm in command. Is that clear?
> Roslin: Yes, sir.
> Adama: Thank you. 'Scuse me.
In theory you could build an air-gapped network within a hospital, but then how do you transmit updates to the EMR's across different campuses of your hospital? How do you issue electronic prescriptions for patients to pick up at their home pharmacy? How do you handle off-site data backup?
Quite honestly, outside of defense applications I'm not aware of people building large air-gapped networks (and from experience, most defense networks aren't truly air-gapped any more, though I won't go into detail). Hospitals, power plants, dams, etc. all of them rely heavily on computers these days, and connect those over the regular internet.
1: My wife was the only pharmacist in her department last night whose computer was unaffected by Crowdstrike (for unknown reasons). She couldn't record her work in the normal ways, because the servers were Crowdstrike'd as well. So she spun up a document of her decisions and approvals, for later entry into the systems. It was over 70 pages long when she went off shift this morning. She's asleep right now.
2: https://www.bd.com/en-uk/products-and-solutions/products/pro...
TIP: many buildings can be part of one LAN! It is called VPN and Russia and China do not like it becouse it is good for peoples!
TIP: data can be easily exchanged when needed! Including LAN.
--
My wife is a hospital pharmacist. (1) When she gets a new prescription in, she needs to see the patients charts on the electronic medical records, and then if she approves the medication a drawer in the Pyxis cabinet (2) will open up when a nurse scans the patients barcode, allowing them to remove the medication, and then the nurse will scan the patient's barcode and the medication barcode in the patients room to record that it was delivered at a certain time. Computers are everywhere in healthcare, because they need records and computers are great at record-keeping. All of those need networks to connect them, mostly on wifi (so the nurses scanners can read things).
--
It was description of very local workflow...
It was description of data flow - no any reason it should be monopolized by unsecure by design os vendor that need to be mandatory secured by essentialy kernel rootkit aka os hacking. Which contradicts using that os in the first place!
And looks like Crowdstrike is just if you ask for price then you can't have it version of SELinux :>>> RH++ for two decades of making presentations of SELinux necessity.
But over all allowing automatic updates from 3rd party not having clue about medicine to hospital system, etc. is managers criminal negligence. Simple as that. Curent state of the art ? More negligence! Add (business) academia & co to chronic offenders. Call them what they truly are - sociopaths via craft training facilities.
>In theory you could build an air-gapped network within a hospital, but then how >do you transmit updates to the EMR's across different campuses of your hospital?
How do you transmit to other campuses of other hospitals ? EASY! Transfer mandatory data. Pleas notice I used words like "mandatory" and "data". I DID NOT SAY "use mandatory http stack to transfer data"! NO. NO, I'm far, faaar from even sugesting THAT ! :>
>How do you issue electronic prescriptions for patients to pick up at their home pharmacy?
Hard sold on that "air-gapped and in cage" meme, eh? Send them required data via secure and private method! Communications channels already "hacked" - monopolized - by FB? Obviously that should do not happend in first place. So resolve it as part of un-win-dosing critical civilian infra.
>How do you handle off-site data backup?
That one I do not get. You saying that cloud access is a only possibility to have backups??? And Internet is a must to do it?? Is medical staff brain dead? Ah, no... It's just managers... Again.
>Quite honestly, outside of defense applications I'm not aware of people building large air-gapped networks
And dhcp and "super glue" and tons of other things was invented by military, for a reason, but that things proliferated to civilians anyway. For good reasons. Air-gapping should be much more common when wifi signal allows tracking how you move in your own home. Not to mention GSM+ based "technologies"...
There is old saying: Computers maximize doing. And when somewhere is chaos then computers simply do their work.
If we adjusted our foreign policy slightly, I think we would dissuade that whole class of attacker.
I can't believe they pushed updates to 100% of Windows machines and somehow didn't notice a reboot loop. Epic gross negligence. Are their employees really this incompetent? It's unbelievable.
I wonder where MSFT and Crowdstrike are most vulnerable to lawsuits?
1. Problem A happens, it’s pretty bad
2. A fix is rushed out very quickly for problem A. It is not given the usual amount of scrutiny, because Problem A needs to be fixed urgently.
3. The fix for Problem A ends up causing Problem B, which is a much bigger problem.
tl;dr don’t rush your hotfixes through and cut corners in the process, this often leads to more pain
I am LMFAO at the entire situation. Somewhere, George Carlin is smiling.
Everything about it reeks of incompetence and gross negligence.
It’s the old story of the user and purchaser being different parties-the software needs to be only good enough to be sold to third parties who never neeed to use it.
It’s a half-baked rootkit part of performative cyberdefence theatrics.
That describes most of the space, IMO. In a similar vein, SOC2 compliance is bullshit. The auditors lack the technical acumen – or financial incentive – to actually validate your findings. Unless you’re blatantly missing something on their checklist, you’ll pass.
Any exception made to this checklist is reviewed by third parties that couldn't care less, bean counters, or those technically incapable of understanding the nuance, leaving only the large providers able to compete on the playing field they manufactured.
This is the result of giving away US jobs overseas at 1/10th the salary
Edit: Found it. https://www.youtube.com/watch?v=nS9nLvGMLH0&t=947s
“Loive from NPR news in Washington“
Usually when I write this devs get all defensive and ask me what the worst thing is that could happen.. I don't know.. Could you guarantee it doesn't involve people dying?
Dear colleagues, software is great because one persons work multiplies. But it is also a damn fucking huge responsibility to ensure you are not inserting bullshit into the multiplication.
We got a new requirement to give doctors access to print them on demand. Before this, doctors only read dot matrix-printed reports that had been vetted for decades. With our XSL-FO PDF generator, it was possible that a column could be pushed outside the print boundary, leading a doctor to see 0.9 as 0. I assume in a worst worst case scenario, this could lead to an misdiagnosis, intervention, and even a patient's death.
I was the only one in the company who cared about doing a ton more testing before we opened the reports to doctors. I had to fight hard for it, then I had to do all the work to come up with every possible lab report scenario and test it. I just couldn't stand the idea that someone might die or be seriously hurt by my software.
Imagine how many times one developer doesn't stand up in that scenario.
But it should not hinge on us convincing people.
If we can at least get that basis then we can start to define more things such as jobs that non Engineers can not legally do, and legal ramifications for things such as software bugs. If someone will lose their professional license and potentially their career over shipping a large enough bug, suddenly the problem of having 25,000 npm dependences and continuous deployment breaking things at any moment will magically cease to exist quite quickly.
I hope organisations start revisiting some of these insane decisions.
Just a few weeks ago I had an OpenBSD box render itself completely unbootable after nothing more than a routine clean shutdown. Turns out their paranoid-idiotic "we re-link the kernel on every boot" coupled with their house-of-cards file system corrupted the kernel, then overwrote the backup copy when I booted from emergency media - which doesn't create device nodes by default so can't even mount the internal disks without more cryptic commands.
Give me the Windows box, please.
I’ve had hardware stop working because I updated the kernel without checking if it removed support, but a. that’s easily reversible b. Linux kept working fine, as expected.
I’ll also point out, as I’m sure you know, that the BSDs are not Linux.
I switched to opensuse afterwards
They ended up giving MS a substantial amount of money to extend support for their use case for some number of years. I can't remember the number he told me but it was extremely large.
Why would Windows systems be anywhere near critical infra ?
Heart attacks and 911 are not things you build with Windows based systems.
We understood this 25 years ago.
... Except Apple pretty much pushes you to run such tools just to get reasonable management key alone things like real-time integrity monitoring of important files (Crowdstrike in $DAYJOB[-1] is how security knew to ask whether it was me or something else that edited PAM config for sudo on corporate Mac)
The best corporate dev platform at moment is WSL2 - most of the activity inside the WSL2 vm isn't monitored by the windows tooling so performance is fast. Eventually security will start to mandate agents inside the WSL2 instance, but at the moment most orgs dont.
Goodluck teaching administrators an entirely new ecosystem, goodluck finding software off the shelf for Linux.
Bespoke is expensive, expertise is rare, Linux is sadly niche.
I recently did a project with a company that wanted to move their app to Azure from AWS — not for any good technical reason but just because “we already use Microsoft everywhere else.”
Completely stupid. S3 and Azure Blob don’t work the same way. MCS and AWS SES also don’t work the same way — but we made the switch not even for reasons of money, but because some Microsoft salesman convinced the CIO that their solution was better. Similar to why many Jira orgs force Bitbucket on developers — they listen to vendors rather than the people that have to use this stuff.
Tbf, you are giving up a clustering index in that trade. May or may not matter for your workload, but it’s a remarkably different storage strategy that can result in massive performance differences. But also, you could have the same by shifting to MySQL, sooooo…
Teach a 60 year old industrial powertrain salesman to use Linux and to redevelop their 20 year old business software for a different platform.
Also explain why it’s worth spending food, house, and truck money on it.
Finally, local IT companies are often incompetent. You get entire towns worth of government and business managed by a handful of complacent, incompetent local IT companies. This is a ridiculously common scenario. It totally sucks, and it’s just how it is.
Windows servers are “niche” compared to Linux servers. Command line knowledge is not “uncommon expertise,” it’s imo the bare minimum for working in tech.
I’m not wildly opinionated here, I should clarify. I’d love a more Linux-y world. I’m just saying that a lot of small-medium towns, and small-medium businesses are really just getting by with what they know. And really, Windows can be fine. Usually, however, you get people who don’t understand tech, who can barely use a Windows PC, nevermind Linux, and don’t really have the budget to rebuild their entire tech ecosystem or the knowledge to inform that decision. It sucks, but it’s how it is.
Also, Open Office blows chunks. Business users use Windows. M365 is easy to get going, email is relatively hands-off, deliverability is abstracted. Also, a LOT of business software is Windows exclusive. And that also blows chunks.
I would LOVE a more open source, security minded, bespoke world! It’s just not the way it is right now.
My entire career has been spent building, and maintaining, critical infra.[1]
Further, in my volunteer time, I come into contact with medical, dispatch and life-safety systems and equipment built on Windows and my question remains the same:
Why is Windows anywhere near critical infra ?
Just because it is common doesn't mean it's any less shameful and inadequate.
I repeat: We've fully understood these risks and frailties for 25 years.
[1] As a craft, and a passion - not because of "exciting career opportunities in IT".
> As a craft, and a passion
I believe you’ve nailed the core problem. Many people in tech are not in it because they genuinely love it, do it in their off time, and so on. Companies, doubly so. I get it, you have to make money, but IME, there is a WORLD of difference in ability and self-solving ability between those who love this shit, and those who just do it for the money.
What’s worse is that actual fundamental knowledge is being lost. I’ve tried at multiple companies to shift DBs off of RDS / Aurora and onto at the very least, EC2s.
“We don’t have the personnel to support that.”
“Me. I do this at home, for fun. I have a rack. I run ZFS. Literally everything in this RFC, I know how to do.”
“Well, we don’t have anyone else.”
And that’s the damn tragedy. I can count on one hand the number of people I know with a homelab who are doing anything other than storing media. But you try telling people that they should know how to administer Linux before they know how to administer a K8s cluster, and they look at you like you’re an idiot.
There is tremendous demand for technology that works well and works reliably. Sure, setting up a database running on an EC2 instance is easy. But do you know all of the settings to make the db safe to access? Do you maintain it well, patch it, replicate it, etc? This can all be done by one of the old school sysadmins. But they are rare to find, and not easy to replace. It's hard to judge from the outside, even if you are an expert in the field.
So when the job market doesn't have the amount of sysadmins/devops engineers available, then the cloud offers a good replacement. Even if you as an individual company can solve it by offering more money and having a tougher selection process, this doesn't scale over the entire field, as at that point the whole number of available experts comes in.
Aurora is definitely expensive, but there is cheaper alternatives to it. Full disclosure, I'm employed by one of these alternative vendors (Neon). You don't have to use it, but many people do and it makes their life easier. The market is expected to grow a lot. Clouds seem to be one of the ways our industry is standardizing.
> But do you know all of the settings to make the db safe to access? Do you maintain it well, patch it, replicate it, etc?
Yes, but to be fair, I’m a DBRE (and SRE before that). I’m not advocating that someone without fairly deep knowledge attempt to do this in prod at a company of decent size. But your tiny startup? Absolutely; chuck a default install of Postgres or MySQL onto Debian, and optionally tune 2 – 3 settings (shared_buffers, effective_cache_size, and random_page_cost for Postgres; (innodb_buffer_pool_* and sync_array_size for MySQL – the latter isn’t necessary until you have high concurrency, but it also can’t be changed without a restart so may as well). Pick any major backup solution for your DB (Barman for Postgres, XtraBackup for MySQL, etc.), and TEST YOUR BACKUPS. That’s about it. Apply any security patches (or use unattended-upgrades, just be careful) as they’re released, and don’t do anything outside of your distro’s package management. You’ll be fine.
Re: Neon, I’ve not used it, but I’ve read your docs extensively. It’s the most interesting Postgres-aaS product I’ve seen, alongside postgres.ai, but you’re (I think) targeting slightly different audiences. I wish you luck!
This is always great feedback to hear, thank you!
1. more Windows programmers than Linux so they're cheaper.
2. more third-party software for e.g. reporting, graphing to integrate with
3. no one got fired for buying Microsoft
4. any PC can run Windows; IT departments like that.
Well, that's a pretty big problem. I don't know how we ended up in a situation where everybody is okay with the most important software being the most insecure, but the money needed to keep critical infra totally secure is clearly less than the money (and lives!) lost when the infra crashes.
Why? "Hardening" the OS is exactly what Crowdstrike sells and bricked the machines with.
Centralization is the root cause here. There should be no by design way for this to happen. That also rules out Microsoft's auto updates. Only the IT department should be able to brick the hospitals machines.
Why would computers be anywhere near critical infra? This sounds like something that should failsafe, the control system goes down but the thing keeps running. If power goes down, hospitals have generator backups, it seems weird that computers would not be in the same situation
This is just a guess, but maybe the client machines are windows. So maybe there are servers connected to phone lines or medical equipment, but the doctors and EMS are looking at the data on windows machines.
-aaS by definition requires you to open yourself to someone else to let them do the work for you. It doesn't empower you, it empowers them.
Smarter people than us have already thought through this and the cost-benefit analysis said "connect it to a server"
If modern medicine is dangerous and fragile because of network connected equipment then that should be fixed even if the way it currently works doesn’t allow it.
If you don't understand why it has to be networked with extremely bad fallback to paper, then I suggest working in healthcare for a bit before pontificating on how everything should just go back to the stone age.
The question is not whether or not hospitals need internet at all or to go back into printing things in paper or whatever nobody ever said. The question is whether everything in the hospital should be connected to the internet. Again the example used was simple. Having the computer processing and exporting the data from an MRI machine connected online in order to transfer the data, vs using a separate computer to transfer the data and the first computer is offline. This is how we are supposed to transfer similar data at my work for security reasons. I am not sure why it cannot happen in there. If you cannot transfer data through that computer, there could be an emergency backup plan. But you need to solve only the transfering data part. Not everything.
Same goes for electronic medical records. There are people who assign ICD-10 codes (insurance billing codes) to patient encounters. Often this is a second job for them and they work remote and typically at odd hours.
A modern hospital cannot operate without internet access. Even a medical practice with a single doctor needs it these days so they can file insurance claims, access medical records from referred patients and all the other myriad reasons we use the internet today.
This stuff isn't impossible to solve. Rather, the incentives just aren’t there. People would rather build an apparatus for blame-shifting than actually just building a better solution.
Why not an internal only network for all the terminals to talk to a central server, then disable any other networking for the terminals? Why do those terminals need a browser where pretty much any malware is going to enter from? If hospitals are paying out the ass for their management software from epic/etc, they should be getting something with a secure design. If the central server is the only thing that can be compromised then when edr takes it down you at least still have all your other systems, presumably with cached data to work from
Its just laziness, and to be honest, an outage like this has no impact on their management reputation as a lot of other poorly run companies and institutions were also impacted, so the focus is on crowdstrike and azure, not them.
No, it doesn't.
Some have chosen - for reasons of efficiency and scale and cost - to place it online.
However, this is a trade-off for fragility.
It's not insane to make this trade-off ...
... but it is insane to not realize one is making it.
maybe Heartbleed or the xzUtils debacles convinced them to switch.
I mean, if the problem is that hospitals can't function anymore, money is hardly the biggest problem
Regardless, thanks for your report; seeing it was very sobering. I hope you can get some rest, and that things will soon return to normalcy.
We were back in the 1960's with paper and pen for everything, no updates on nature of call, no address information, nothing... find out when you show up and hope the scene is secure. It was wild as it was coupled to a relatively intense Monsoon storm.
I've told my testers for years their efficacy at their jobs would be measured in unnecessary deaths prevented. Nothing less. Exactly this outcome was something I've made unequivocally clear was possible, and came bundled with a cost in lives. Yet the "Management and bean counter types" insist "Oh, nope. Only the greenbacks matter. It's the only measure."
Bull. Shit. If we weren't so obsessed with imaginary value attached to little green strips of paper, maybe we'd have the systems we need so things like this wouldn't happen. You may not be able to enumerate every, but you damn well can enumerate enough. Y'all just don't want to because then work starts looking like work.
That doesn’t count serious bodily injury, suffering, people who were victimized, people who had their lives set back for decades due to a missed opportunity, a person who missed the last chance to visit a loved one, etc.
There are uncountable different impacts that happen when you’re talking about events on the scale of an economy. Which is why economists use dollars. The proxy isn’t useful because it is more important than life, it it useful because the diversity of human experience is innumerable.
At least putting a number to life is an genuine attempt even though it may be distasteful.
The fact is that there already is a number on it, which one can derive entirely descriptively without making moral judgements. Insurance companies and government social security offices already attempt to determine the number.
The number is not infinite or we'd have no cars.
Not questioning that it happened, but this was a boot loop after a content update. So if the computers were off and didn't get the update, and you booted them, they would be fine. And if they were on and you were using them, they wouldn't be rebooting, and it would be fine.
How did it happen that you were rebooting in the middle of treating a heart attack? [Edit: BSOD -> auto reboot]
> And if they were on and you were using them, they wouldn't be rebooting, and it would be fine.
Windows has been notorious for forcing updates down your throat, and rebooting at the least appropriate moments (like during time-sensitive presentations, because that's when you stepped away from the keyboard for 5 minutes to set up the projector). And that's in private setting. Corporate setting, the IT department is likely setting up even more aggressive and less workaround-able reboot schedule.
Things like this is exactly why people hate auto-updates.
The assumption is that "pros" and "enterprise" either know how to use provided controls or have WSUS server setup which takes over all of scheduling updates.
Especially if an infected machine can attack others?
Users/IT regularly would never update or deploy patches which has its own consequences. There’s no perfect solution—but rather there to accept the pain.
It’s a lot like herd immunity in vaccines.
Yes. But you don't deploy experimental vaccines simultaneously across the entire population all at once. Inoculating an entire country takes months; the logistics incidentally provide protection against unforeseen immediate-term dangerous side effects. Without that delay, well, every now and then you'd kill half the population with a bad vaccine. The equivalent of what's happening now with CrowdStrike.
in the same way cars are notorious for forcing you to run out of gas while you're driving them and leaving you stranded... because you didn't make time to refill them before it became a problem.
> "Things like this is exactly why people hate auto-updates."
And people also hate making time for routine maintenance, and hate getting malware from exploits they didn't patch, and companies hate getting DDoS'd by compromised Windows PCs the owners didn't patch, and companies hate downtime from attackers taking them offline. There isn't an answer which will please everyone.
> "This prevention of functionality during a critical period while forcing an update would be like if a modern car refused to drive during an emergency"
Machines don't know if there's an emergency going on; if you don't do maintenance, knowing that the thing will fail if you don't, then you're rolling the dice on whether it fails right when you need it. It's akin to not renewing an SSL certificate - you knew it was coming, you didn't deal with it, now it's broken - despite all reasonable arguments that the connection is approximately as safe 1 minute after midnight as it was 1 minute before, if the smartphone app (or whatever) doesn't give you any expired cert override then complaining does nothing. Windows updates are released the same day every month, and have been mandatory for eight years: https://www.forbes.com/sites/amitchowdhry/2015/07/20/windows...
And we all know why - because Windows had a reputation of being horribly insecure, and when Microsoft patched things, nobody installed the patches. So now people have to install the patches. Complaining "I want to do it myself" leads to the very simple reply: you can - why didn't you do it yourself before it caused you a problem?
If you're still stubbornly refusing to install them, refusing to disable them, refusing to move to macOS or Linux, and then complaining that they forced you to update at an inconvenient time, you should expect people to point out how ridiculous (and off-topic) you're being.
> It's akin to not renewing an SSL certificate.
Your choice of analogies is a good one. I have done SSL type stuff since 1997.
Doesn't matter: I would have to work a few hours very carefully before modifying my web server config. And test it.
I am terrified by scale of deployment involved in this CloudStrike update.
In more professional settings than private small car ownership, you often will both have regular maintenance updates provided and mandates to follow them. Sometimes they are optional because your environment doesn't depend on them, sometimes they are mandatory fixes, sometimes they change from optional to mandatory overnight when previous assumptions no longer apply.
Several years ago a bit over 100 people and uncounted amount of possible more had their lives endangered because an extra airflow directing piece of metal was optional, and after the incident it was quickly made mandatory, with hundreds of aircraft being stopped to have the fix applied (which previously was only required for hot locations - climate change really bit it).
Similarly, when you drive your car and it fails to operate, that's just you. When it's a more critical service, you're either facing corporate, or in worst case, governmental questions.
If said update pushes you into bsod where automatic watchdog (by default set enabled in windows) reboots...well, here you have a bootloop
I can see very well how one computer could have screwed all others. It's really not hard to imagine.