CrowdStrike global outage to cost US Fortune 500 companies $5.4B
theguardian.com
theguardian.com
https://www.washingtonpost.com/transportation/2024/07/23/del...
Even still, Delta had a really bad time recovering, and is still cancelling flights days later. It's not just "a day or two". At a billion a week, with 30-40% cancellation rates, that's 300-400M just for this one customer. And that's just lost revenue. Imagine the extra costs: customer service complaints, hardware / IT restoration, extra wages for flight attendants working double duty to keep the remaining flights going. Madness.
Even just in travel segment, how many hotels, car rentals, uber/lyft rides, etc were cancelled b/c of missed flights? How much do you think they paid on top of the lost revenue to handle customer complaints, IT restoration, etc?
The repair costs alone at a given airport must be staggering, as every terminal screen is BSOD and needs a tech to manually restore from bitlocker.
https://www.cbsnews.com/news/delta-flight-cancellations-toda...
So, just one of the affected companies brings the total to $1B, wouldn't you say $5B is actually a low estimate?
Of course in our real world, they are unlikely to pay anything at all and just continue operating as-is.
With modern tech stacks, we’re talking hundreds of updates a quarter. If not more.
We moved away from that for a reason. And the real fix is to make updates not take down the whole system. When was the last time every Linux system in the world was unable to boot all at the same time?
Even the crowdstrike issue only took down a very limited subset of installs and the blast radius is far more minimal due to the linux kernel being a lot more sane and cautious.
My phone does auto update, and nearly every auto update makes it worse and slower. That's been true for nearly every Android I've had.
One can do the same thing with computers. Build small systems that operate independently, with strictly defined inputs and outputs. Make them un-brickable. Distribute software on read-only media, or outright replace the CPU for each update. You roll back by having the staff unplug the new thing and plug in the old thing.
A high-spec CM4 is $90. A small USB stick costs basically nothing. A bespoke CM4-like machine with only ROM could be made in volume for relatively little money. Most of the problem with an approach like this is that the industry doesn’t think this way.
Maybe a big buyer should step up, e.g. a major military. Non-CrowdStrike-able tech for, say, the US military seems like it would be extremely valuable.
I feel like the answer should be obvious: if you need reliable systems then you don't create a single point of dependency on such fragile systems.
The latest blog from Bruce Schneier is a good writeup on this brittleness:
https://www.schneier.com/blog/archives/2024/07/the-crowdstri...
That wouldn't make any sense.
The blame realistically lies within each company who allowed a critical point of failure just so they get checkbox software and not actually have to expend the effort of making sure the company is able to function with their chosen infrastructure.
If airlines skimp on maintenance and planes crash, well, all those dead passengers should have picked an airline that takes maintenance more seriously. Next time they'll know better.
No next time, they're dead, so not a good selection criteria.
Agreed that both sides share some liability, but a vendor that says they create security software with root level code injection access must be held to the highest possible standard, and liability.
If someone puts up a shack and sells hamburgers and hundreds of people die from the rotten meat, they most certainly are liable.
But apparently crowsdstrike can brick computers remotely, causing enormous damage including some deaths, and just say oopsie and walk away without any liability.
Or a world where the regulator that required the checkbox — even if the firm can demonstrate a superior way of achieving the same objective — should pay for it.
I've been involved in procurement at a big corporation and one thing we always modified in contracts was making the vendor 100% liable for any damages caused by their outages, but many vendors wouldn't make that modification.
Which in this case is probably a lot less than what these companies are paying in clean up costs.
Building bridges is a very overused analogy, but it's still a good one. Structural engineers still design bridges and they get built.
What they don't do it throw some drawing together and YOLO it to the builder and let's see what happens, the way software gets built.
If a bridge has maintenance issues causing it to be unusable while it is being fixed, the bridge makers are not on the hook for the total downstream economic loss.
Realistically, this doesn’t happen because the returns don’t make sense. If vendors are liable for your business, they need to account for your business in their costs. That drives their costs (and prices) up.
That is not a bad thing. Today all these shaky companies exist because they can take all the profit and externalize all the downsides, laughing all the way to the bank.
If they were made liable, one of two things might happen, both of which are positive outcomes.
One, they might upgrade their engineering and quality processes to the point where they can guarantee they won't be the root cause for taking down large parts of the industry. Sure it'll cost more but if they can do it and still make a profit (just less), all is well.
Two, maybe it simply can't be profitable to build these root backdoor systems with enough layers of safety, in which case the companies will disappear. This is also good; if the product can't be safe, it should not exist.
Is there some good reason for this approach (need to get config updates into the wild as quickly as possible to combat zero-days or zero-hours?) or was this just a massive oversight?
Side rant... their postmortem took forever to get to the point, first explaining all their jargon and product names. Makes me really appreciate the Cloudflare ones.
It is so straightforward and it always works.
Incentives, outcomes, the usual.
But you don't release until QA is done, esp. if you're touching safety-critical systems. Turns out CS didn't realize they ran safety critical systems in airlines and hospitals.
How many times do you think someone proposed:
We should harden our code / delivery in case the system is important
Only to hear We can add robustness later, we will patch it in Terms of Service "for now"They very likely have automated tests. However, what if bug only triggers 90% of the time and you hit the lucky 10% during automated tests? Of course you can run tests 100 times but... is this a common practice? Moreover, we have both code and anecdotal evidences that the bug may indeed happen randomly. Tavis Ormandy posted a rough analysis of the crash context: https://x.com/taviso/status/1814762302337654829. It looks like the crash is caused by first checking if an uninitialized pointer is NULL, and if not, dereferencing it. If the uninitialized leftover data just happened to be zero, no crash happens.
And anecdotally, we saw people reporting that repeatedly rebooting their machines for 15+ times fixed the problem for them - because eventually you got lucky and in a boot it didn't happen and CrowdStrike managed to update itself to not crash.
> not have gradual/staggered rollouts for their deployments
No idea. Maybe their poor reliability guy got overrided by another team, like "how dare you delaying our important definition update? we're racing with threat actors!". I hope they learned their lesson.
I'm curious of this too. Has there ever been a scenario where zero-days could have caused so much damage that it'd warrant this speediness in patching? In a cost-benefit analysis would it justify x% of patches like this happening in preventing whatever security issues could occur without this type of infrastructure?
I don't think many enterprises will switch because of the effort required and, instead, they'll just yell at the account reps for a while and then go back to paying the invoice. However, a big lawsuit is something i didn't think of.
Section 6.1:
THERE IS NO WARRANTY THAT THE SOFTWARE OR ANY OTHER CROWDSTRIKE OFFERINGS WILL BE ERROR FREE, OR THAT THEY WILL OPERATE WITHOUT INTERRUPTION OR WILL FULFILL ANY OF SOFTWARE USER’S PARTICULAR PURPOSES OR NEEDS. THE SOFTWARE AND ALL OTHER CROWDSTRIKE OFFERINGS ARE NOT FAULT-TOLERANT AND ARE NOT DESIGNED OR INTENDED FOR USE IN ANY HAZARDOUS ENVIRONMENT REQUIRING FAIL-SAFE PERFORMANCE OR OPERATION. NEITHER THE SOFTWARE OR ANY OTHER CROWDSTRIKE OFFERINGS ARE FOR USE IN THE OPERATION OF AIRCRAFT NAVIGATION, NUCLEAR FACILITIES, COMMUNICATION SYSTEMS, WEAPONS SYSTEMS, DIRECT OR INDIRECT LIFE-SUPPORT SYSTEMS, AIR TRAFFIC CONTROL, OR ANY APPLICATION OR INSTALLATION WHERE FAILURE COULD RESULT IN DEATH, SEVERE PHYSICAL INJURY, OR PROPERTY DAMAGE. SOFTWARE USER AGREES THAT IT IS SOFTWARE USER’S RESPONSIBILITY TO ENSURE SAFE USE OF SOFTWARE AND ANY OTHER CROWDSTRIKE OFFERING IN SUCH APPLICATIONS AND INSTALLATIONS.
They essentially got the customer to accept a contract that says the software isn't designed for use in systems where failure could cause death, and that the customer accepts responsibility for using it appropriately.
I agree this whole incident was a massive blunder by CrowdStrike, but I'm not sure it makes sense to hold them liable for damage caused by customers using the product in a places they explicitly agreed not to use it in. In those cases, I think the organization that installed CrowdStrike's software in inappropriate places bears a lot of responsibility for the outcome, and their failure to understand the TOS they agreed to doesn't mean it's not a legally binding contract.
It'll be interesting to see how it all plays out.
If the rope climber is the same person who purchased the rope, then they get a Darwin award!
Otherwise need more detail: is the rope on loan? What's the licensing structure of the rope? Is the license still attached to the rope somehow?
Is the position that CrowdStrike should not be used on anything important because their software can not be trusted? I mean, that's where I'm at now, and I bet many others feel the same.
If damages can be demonstrated, what are the chances of airlines successfully claiming compensation? Or, in practice, do such cases usually result in significant discounts during the next contract renewal rather than actual damages paid out?
who is 'we'?
Most contracts have indemnity clauses that protect from, or cap damages, due to vendor issues. You can get a court to overrule that if you can prove something like gross negligence, or that such provisions don't apply to something like safety-critical airline systems.
CS could push back saying they just offer endpoint protection, and it's on your org for where you put it. Kind like Ikea saying "hey man we just make end tables" when someone decides to put Lack tables on every airplane, and they turn out to be super flammable.
https://techcrunch.com/2024/07/24/crowdstrike-offers-a-10-ap...
That's why they have so many enterprise customers. They're the only game in town that won't slow down your servers arbitrarily while still convincing an auditor that you do have an antivirus.
Too bad they also crash your whole system every now and then.
Unless the hardware was already near failure I don’t see how this could cause hardware failure. The worst case scenario was the machine just constantly rebooting but after 3 (I think, somewhere around that number) it should have launched into WindowsRE.
https://www.theregister.com/2024/07/21/crowdstrike_linux_cra...
If you're writing a kernel driver that is deployed throughout a great portion of Fortune 500, with the money that that entails, then you should definitely be able to afford to pay people to write defensive code and have proper pipelines in place.
Soo. $5.4B - $10
https://www.crowdstrike.com/falcon-content-update-remediatio...
It boils down to the "Content Validator" had a bug and gave a false positive.
It's kind of crazy that the 'rapid response content' update was then free to go out direct to production machines with zero actual live testing.
That's either due to c-suite excel cost-cutting/maximize profit or silicon valley yolo.
but to the parent poster's point... maybe. sometimes you gotta throw a wrench in it and see what happens.
According to the discussion in the thread, you're correct. Also, it was a $10 giftcard for .. uber eats. Where you can't get anything for less than ten bucks.
I wonder if Uber Eats and DoorDash give these out for free to companies as promotion. I bet most people who use the cards spend another $20 or more.
Also, does the IT manager get that giftcard? Do they share it with the rest of the team? Does the CTO get the card and shares it with the rest of the C-suite. What's the proper way of handling that other than reject with a harsh laugh in their face at the offer?
1) These types of products cause incidents all the time, this was just a very high impact one that happened to affect everybody all at once.
2) Their product is very good compared to the competitors.
All products in this space are black boxes, but CS is one of the least black-boxy, the alerts it produces are decent, the tooling is comes with is especially good (from an operator perspective), and the reporting it produces is exactly the sort of thing decision makers in big enterprise love to see.
I doubt there’s going to be much churn from this, definitely not an existential amount. As much as I personally can’t stand the organisation, I think they absorbed most of the bad press on behalf of all the service providers they took offline.
They’re more likely to investigate running two platforms simultaneously or, more likely, talk to competitors purely as a bargaining chip to negotiate a bigger discount for the next renewal.
The end result is still the same though, CrowdStrike will lose a lot of income and confidence but they’re not going out of business.
Frankly, even if they were to, I’m certain they’d end up getting bailed out anyway. But I can’t see it getting to the point to begin with.
Their biggest problem is obviously going to be with customer retention, but there are huge technical and regulatory hurdles their customers would have to go through to switch to a competitor. I'm sure many of their customers will accept this as an isolated incident and be quick to accept CrowdStrike's assurances that this won't happen again.
I think Delta Airlines is probably in more trouble right now than CrowdStrike.
Ideally the stock price should be pushed up to attract better engineers to go in and fix shit. If stockholders agree to bid up the stock 100% YoY, hell, even I'd look for a job there, and help fix shit in return for some juicy RSUs.
If you take away their funding, you can only expect worse in the future.
It crashed in the last week. That's what matters.
"We don't have enough time to write tests"
"Developers should be able to test their own code"
Yeah, I know using free software isn’t a panacea. Still it would be a step in the right direction, plus I could not refrain from the cheap shot at M$ Windows.
Crowdstrike shipped a driver that they marked as a mandatory boot driver. The Windows OS could have had more recovery options otherwise.
If half your tills are windows/defender and half linux/crowdstrike then half your tills are going to be working.
For EMS, hospitals, Windows makes sense on the server because they don't know any better. For anyone remotely technologically competent, Windows shouldn't even be considered an option other than as workstations. Linux on the server is the only way and no one can convince me otherwise.
>Linux on the server is the only way and no one can convince me otherwise.
Now meet the sysadmin that thinks the same, but for windows for clients. At the risk of overgeneralizing, people are only for "diversity" when it means supporting their preferred underdog platform (eg. linux desktop). When they're the dominant incumbent it's suddenly "dark pattern", "they don't know any better" and "no one can convince me otherwise".
If the results match: Everything is largely proven to be working as-designed, and the output is assumed to be valid. This is an advantage.
If one breaks: Nothing is proven to be working, but that's no worse than we have today with just one system. This is not a disadvantage.
Let's blame bullshit compliance?
While the driver for purchase is almost always to pass audits, it's still a good product.
It probably does provide benefits against some clueless intern in accounting downloading a macro-enabled excel file that has a malware enabled.
- Wrote code responsible for $5.4 billion
Probably a refund is all they’ll be on the hook for.
Sadly, damage done like this is just chalked up to an accident, and swept under the rug.
By cashing in this $10 Uber Eats coupon you agree to hold harmless...
- https://news.ycombinator.com/item?id=41058261
- https://techcrunch.com/2024/07/24/crowdstrike-offers-a-10-ap...
We put absolutely no critical thought into whether this was a likely thing, and we completely ignored the many government and media reports that are credibly sourced which state that there are known phishing scams and other threat actors trying to capitalize on this incident.”
I highly doubt this is something that Crowdstrike actually did.
Edit: Amazingly they did, the article has been updated with a statement. Amazingly stupid all around.
From my point of view, one of the greatest problem for them is that they bypassed customers deployment policies.
Caveat emptor. Falcon and other similar security products often push updates at-will, and they're fully transparent about this if you actually read the contract terms and understand the vendor's approach to operations. I have worked with many clients that elect not to use such tools in certain sensitive environments, specifically to mitigate the risk of being impacted by something like CrowdStrike's 7/19 event.
Do you really want to wait until for the weekly/monthly/quarterly deployment window to deploy a detection update for a 0day, or a new type of malware?