But it is something to be very aware of for those of us who develop software run in e.g. hospitals and airlines, and should receive more attention, instead of only bringing up financial losses which is what usually happens. I noticed the same with the big ransomware attacks.
CrowdStrike knows that their software runs on computers that are in fricken hospitals and airports, they know that a mistake can potentially cause a human death. They also know how to properly test software, and they know how to do staggered releases.
Given what we know now, it seems pretty likely that to any reasonable person, the amount of risk they took when deploying changes to clients was in no way reasonable. People absolutely should go to jail for this.
This more or less originated with the unfortunately named MS Herald of Free Enterprise sinking (https://en.wikipedia.org/wiki/MS_Herald_of_Free_Enterprise) - after that incident, regulators decided that maybe they didn't want enterprise quite as free as all that, and cracked down significantly on shipping operators (though the attempt to prosecute its execs for corporate manslaughter did fail).
The reason this reminds me of that, assuming that I remember right, is that I think they had even taken the decision that the cost of paying lawsuits for those injuries was lower than the increase in revenue for being able to say "we have the hottest coffee"… and that was why they were deemed so severely liable.
They were definitely shown to have known it was resulting in injuries from other settlements:
https://en.wikipedia.org/wiki/Liebeck_v._McDonald%27s_Restau...
Why don't orgs test their updates? Every decent IT management/governance under the sun demands that you test your updates. How the hell did so many orgs that are ISO 2700x, COBIT, PCI-DSS, NIST CSF, etc. certified failed so hard??
(ToS/contracts will probably get you out of any damages.)
Because historically orgs have been really bad with applying updates: either no updates or delayed updates resulting in botnets taking over unpatched PC's. Microsoft's solution was to force the updates unconditionally upon everybody with very few opportunities to opt out (for large enterprise customers only).
Another complication comes from the fact that operating system updates are not essential for running a business and especially for small businesses – as long as the main business app runs, the business runs. And most businesses are too far removed from IT to even know what a update is and why it is important. Hence the dilemma of fully automated vs manually applied and tested updates.
Not a Microsoft's fan, but this is not true. Everyone who has Windows Server somewhere, with some spare disk space for the updates, has this ability. Just install and run WSUS (included in Windows Server) and you can accept/reject/hold indefinitely any update you want.
1) the prevailing majority of laptop and desktop PC installations (home, business and enterprise) are not Windows Server;
2) kiosk style installs (POS terminals, airport check-in stands etc) are fully managed, unsupervised installations (the ones that ground to a complete halt today) and do not offer any sort of user interaction by design;
3) most Windows Server installations are also unsupervised.
They are not, but the point is elsewhere: that Windows Server is going to provide the WSUS service to your network, so your laptop and desktop installations (in business and enterprise) are going to be handled by this.
Homes, on the other hand, do not have any Windows Server on their network, that's true.
As a hack to disable Windows updates, it is possible to point it to a non-existing WSUS server (so that can be done at home too). The client will then never receive any approval to update. It won't receive any info wrt available updates either.
> 2) kiosk style installs (POS terminals, airport check-in stands etc) are fully managed, unsupervised installations (the ones that ground to a complete halt today) and do not offer any sort of user interaction by design;
That's fine; this is fully-configurable via GPO.
> 3) most Windows Server installations are also unsupervised.
See 2.
What I take from this is that vendors need a LOT more investment in that work. They have both the money and are best positioned to do that testing since the incentives are aligned better for them than anyone else.
I’m also reminded of all of the nerd-rage over the years about Apple locking down kernel interfaces, or restricting FDE to their implementation, but it seems like anyone who wants to play at the system level needs a well-audited commitment to that level of rigorous testing. If the rumors of Crowdstrike blowing through their staging process are true, for example, that needs to be treated as seriously as browsers would treat a CA for failing to validate signing requests or storing the root keys on some developer’s workstation.
These are infra companies. Their incompetence can literally kill people.
So they drop half the 'safety' procedures once the auditor goes away? WTF! (I am semi-angry because there are so many easy solutions and workarounds to not fall for this!! (inside screaming).
How irresponsible must someone be to roll out something to 1k-5k-10k machines without testing it first??
Hubris-Atis-Nemesis-Tisis!!!!
https://www.greecehighdefinition.com/blog/hubris-atis-nemesi...
I'm not trying to enforce certifications because as a dev certifications always raise a bitter taste in my mouth. But those companies need certified processes that get re-certified every year. Sometimes even a cursory review from outsiders can find a lot of issues.
Do it to Boeing, sure.
Why not dilute the shareholder pool by a serious amount? There's no need for a statization to formally happen, the government can sell the shares back over time without actually exercising control.
Also fire execs and ban them from holding office on publicly traded companies for the foreseeable future.
Seizing shares doesn't impact the cash flow of the company directly, thus shouldn't cause job losses, but shareholders (who should put pressure on executives and the board to act with prudence to avoid these kinds of disasters) are adequately punished.
That could be amazing: "Ooopsie, in punishing Crowdstrike they've ended up folding and now there's a second global outage."
So if I own some Vanguard mutual fund as part of a retirement account, it’s now on me to put pressure on 500+ corporations?
Perhaps it’s on Vanguard to do so…but Vanguard isn’t going to just eat the cost of increased due diligence requirements. My fees will increase.
How does that increased due diligence even work? It’s not like I or Vanguard can see internal processes to verify that a company has adequate testing or backups or training to prevent cases like today’s failure.
When, on average, X number of those 500 companies in my mutual fund face this share seizure penalty per year…am I just supposed to eat the loss when those shares disappear? Does Vanguard start insuring against such losses? Who pays for that insurance in the end?
This doesn’t even really hurt the shareholders who are best placed to possibly pressure a company. This doesn’t hurt “billionaire executive who owns 40% of the outstanding shares”. I mean, sure, it will hurt that little part of their brain that keeps track of their monetary worth and just wants to see “huge number get huger”…but it doesn’t actually hurt them. It just hurts regular folks, as usual.
Consequently you don't put pressure on the 500 companies, you put pressure on the mutual fund and the mutual fund in turn puts pressure on the companies it invests in and exercises additional discretion in which companies it invests in.
>Perhaps it’s on Vanguard to do so…but Vanguard isn’t going to just eat the cost of increased due diligence requirements.
Yes they do, because mutual funds do compete with one another and a mutual fund that does the due diligence to avoid investing in companies that are held liable for these kinds of incidents will outperform the mutual funds that don't do this kind of due diligence.
> It’s not like I or Vanguard can see internal processes to verify that a company has adequate testing or backups or training to prevent cases like today’s failure.
I don't know specifically about Vanguard, but mutual funds in general do employ the services of firms like PwC, Deloitte, and KPMG to perform technical due diligence that assesses the target company's technology, product quality, development processes, and compliance with industry standards. VC firms like Sequoia Capital and Andressen Horowitz do their own technical due diligence.
You probably want, in addition to your proposal, executive stock-based compensation to be awarded in a different share class, used to finance penalties in such cases where the impact is deemed to be the result of gross negligence at the management level.
That gives shareholders of other companies good reason to care going forward.
Same thing with AI. You can't punish an AI, it has no body.
We don't have a good idea how to make software that is flawless, at least, not at scale for a cost that is acceptable. This is changing a little bit now with the drive by governments to use memory-safe languages, but that only covers a small part of the possible spectrum of bugs in software and hardware.
In this case it seems most of the software which is failing is dull back office stuff running on Windows - billing systems, train signage, baggage handling - which no one thought was critical, and there's no way on earth we could afford to rewrite it in the same way as we do aircraft systems.
Point of sale in a records store, less important. Point of sale in a pharmacy, could be problematic. Web shop customer call center, less important. Emergency services call center, could be problematic.
So if you develop your compression library it can't be used by anyone running critical infra unless you stamp it "critical certified", which in turn will make you liable for some quality issues with your software.
That's very realistic and already happens by requiring certain standards from the resulting product. For example, there are security standards and auditing requirements for medical systems, payment systems, cars, planes, etc.
Outside of regulated industries it's the context in which software is used which determines how critical it is. (As you say.)
So what you seem to be suggesting (effectively) is that use of software be regulated to a greater/lesser extent for all industries... and that just seems completely unworkable.
In physical world, you can specify a tolerance of 0.0005 in but the part is going to cost $25k a piece. It is trivially easy to specify tolerance, very hard to engineer a whole system that doesn't blow the cost and impossible to fund.
Great software architectures are the ones that operate cheaply, but are bulletproof when software fails. https://en.wikipedia.org/wiki/Chaos_engineering
Boeing disagrees.
We're talking about critical software. If we can't afford to reach the level of safety needed because it's too expensive, well so be it.
Besides, the enormously expensive flight systems don't seem to make my plane ticket expensive at all...
What failed today is a bunch of Windows stuff, of which there is a vast amount of software produced by huge numbers of companies, all of very variable quality and age.
Now, that it was not known previously to be critical, that may be. Whether we should have realised its criticality or not, is debatable. But going forward we should learn something from this. So maybe think more about cascading failures and classify more things as critical.
I have to wonder how the failure of billing and baggage handling has resulted in 911 being inoperative. I think maybe there's more to it than you mention here.
It's also common to rollout changes regionally to prevent global impact.
To me it seems Crowdstrike does not have a very good release process.
Here's something I don't understand: those jobs pay chump change compared to places like FB and (afaik) social networks don't have the same life-or-death context
That isn't an easy thing to do, but it should be possible.
There is no 'AI', that is always only hype. There is machine learning, which is a very powerful technology but I doubt MSFT will be leading that revolution. As for LLMs, MSFT might have some competitiveness there but I doubt it's going to be a very lucrative market. MSFT is highly overvalued.
I agree with you on the technical aspect, but the distinction makes regular people eyes glaze over within 5 seconds of that explanation. AI as a label for this is here to stay the same way cyber stopped meaning text sex of IRC. The people have spoken.
<< MSFT is highly overvalued.
Yes, but so is NVDA, the entire stock exchange and US real estate market. We are obviously due for a major correction and have been for a while. As in, I actually moved stuff around in my 401k to soften the blow in that event 2 years ago now. edit: yes, I am a little miffed I missed out on that ride.
So far, everything was done to prevent hard crash and in the election year, that is unlikely to change. Now after the election, that is another story altogether.
<< I doubt MSFT will be leading that revolution.
I think I agree. I remain mildly hopeful that the open model approach is the way.
Agree. First half of 2025 could be pretty spectacular (if/when we get through 2024).
I suspect there might be some pretty radical plans for US debt monetisation being drawn up, to be implemented early in the new presidential term.
You should stop trying to predict the next crash. According to the study, most people (including institutional investors) consistently believe there is a >10% chance the market will crash in the next 6 months when historically the probability is only 1%
Hmm? No. I will attempt to secure my own financial interest.
<< According to the study, most people (including institutional investors) consistently believe there is a >10% chance the market will crash in the next 6 months when historically the probability is only 1%
Historically is doing a fair amount of work here. I would argue there is little historical value to the data we face. Over the past few decades we went through through several mini revolutions ( industrial, information and whatever they end up calling now ) in terms of how we work, eat, communicate and, well, live.
All of these have upended how humans interact with the world effectively changing the calculus on the data that preceding it if not nullifying it altogether in some ways.
Your argument is to stop worrying since you are likely wrong anyway, by a factor of 10. I am saying is 1935 people also thought they have time to ride the wave.
edit: ok, need coffee. too many edits
If someone is injured or dies because the hospital has inadequate backup processes in the event of a Windows outage, some or maybe all liability for negligence falls on those who designed the hospital that way, not the IT supplier who didn't agree to it.
The tool isn't fit for purpose
Is that even possible any more? (That said, "operate" isn't a boolean, it's a continuum between perfect service and none, with various levels of degraded service between, even if you mean "operate" in the sense of "perform a surgical operation" rather than "any treatment or care of any kind").
All medical notes being printed in hard-copy could be done, that's the relatively easy part. But there's a lot of stuff which is inherently IT these days, gene sequencing, CT scans, etc., there's a lot that computers add which humans can't do ourselves — even video consultation (let alone remote surgery) with experts from a different hospital, which does involve a human, that human can't be everywhere at once: https://en.wikipedia.org/wiki/Telehealth
> Nothing critical / life-or-death / personal injury should rely on Windows / IT systems.
If you think that's bad, you may want to ensure you're seated before reading this about the UK nuclear deterrent: https://en.wikipedia.org/wiki/Submarine_Command_System
Many companies paying lip service to quality/reliability but internal incentives almost always go against maintenance and quality of service work (and instead reward new projects, features e. t. c.).
Of course it would. Restaurants are held liable for food poisoning, but they still operate just fine. They just - y’know - take care that they don’t poison their customers.
If computer systems were held liable, software would be a lot more expensive. There would be less of it. And it would also be better.
I think I can get behind that future.
Write me software that coordinates all flights to and from airports, capturing all edge-cases, that's bug free. Then tell me the number you estimate and the number of years to roll this out.
The right way to write code like that is to start simple and small - we're going to service airports X, Y and Z. Those airports handle Q planes per day. The software will be used by (this user group) and have (some set of responsibilities). The software engineers will work with the teams on the ground during and after deployment to make sure the software is fit for purpose. Someone will sign off on using it and trusting its decisions. And lets also do a risk assessment where we lay out all the ways defects in the software could cost money and lives, so we can figure out how risk averse we need to be.
Give me scope like that, and sure - I'll put a team together to write that code. It'll be expensive, but not impossible. And once its working well, I'd happily roll it out to more airports in a controlled and predictable manner.
Disclaimer. Neither Microsoft, nor the device manufacturer or installer, gives any other express warranties, guarantees, or conditions. Microsoft and the devicemanufacturer and installerexclude all implied warranties and conditions, including those of merchantability, fitness for a particular purpose, and non-infringement. If your local law does not allow the exclusion of implied warranties, then any implied warranties, guarantees, or conditions last only during the term of the limited warranty and are limited as much as your local law allows. If your local law requires a longer limited warranty term, despite this agreement, then that longer term will apply, but you can recover only the remedies this agreement allows.
So they might get sued on a local level?
I'm not sure how i'm going to explain the productivity loss and retraining costs to the board if im honest.
You can switch away from CrowdStrike but I doubt you'll be able to convince whoever mandated CS to be installed to not install an alternative that carries exactly the same risks.
In fact there was a recent CrowdStrike-related crash in RHEL:
https://old.reddit.com/r/crowdstrike/comments/1cluxzz/crowds...
Crowdstrike only exists because Windows and other Microsoft products are so insecure their default configuration.
Every State Farm insurance office in the country is still using a DOS App from the 1980's to run their office.
These things should have gone from mainframes of yore to various unix systems, ideally a mix of different unix systems in hot failover.
Without running uncontrolled "agent" software of course.
If it was, it would also make it impractical for a small business to contract with a large one because of risk.
Do these clients have SLAs? If so, they're definitely on the hook for something. You could probably get a few businesses together for a decent class-action against Crowdstrike. You're then expecting a lawyer to be able to convince a dozen semi-random people with varying degrees of computer knowledge that Crowdstrike's software was negligently designed, developed, and deployed in a way that caused financial or life losses for customers.
So, really, it's a coin flip.
Enough of this limited liability nonsense, there need to be serious, severe, life-changing consequences.
Wouldn't that be the desirable outcome, though? Given the amount of damage they have caused, they should cease to exist.
We need this so that every company board is always asking "are we investing enough to make sure this never happens to us?"
It takes 2 year for the legal system to catch up, at which point he starts a new company, bankrupts the old one, sells all his tools cheaply to the new company, and fires and rehires his workers. I've seen this game going on for 14 years now.
I think Crowdstrike would do the same: Start a new one, sell the software, fire and rehire the workers, then go on as if nothing happened
Killing the company because they made a mistake doesn't just throw away a ton of learned lessons (because the devs will probably be scattered around the industry where their newly acquired domain knowledge will be less valuable) but also forces a lot of companies to spend resources changing their antivirus scanners. For all we know, Crowdstrike might never fuck up again after this and forcing that change would burn hundreds of millions for basically no reason.
I don't think that's right, since it ignores externalities.
You want to create a system where every company is incentivized to make positive security decisions. If your response to a fuckup of unprecedented scale is just "they learned their lesson, they probably won't do that again", then the message these companies receive is that it is okay to neglect proper security procedures, because you get one global economic meltdown for free.
Haven't heard that one before but I love everything about that!
Was it really a botched update? Or was it a test run for holding the world hostage prior to a coup?
Zero. Exactly Zero.
Clearly you have never been involved in buying insurance or writing contracts for IT products/services.
Loss of contracts, profits, goodwill, economic loss, loss of data and all that jazz is excluded in whole or limited to a fixed monetary value.
It is known as indirect, consequential or special loss, damage or liability.
No lawyer worth their salt will let an IT product/service company draft a contract that does not have the above type of clause..
And good luck finding an insurance contract that will pay out for such losses, indeed most of them have conditions that state your contracts with customers must exclude or limit such losses.
Most software also has clauses excluding use in safety critical environments.