Microsoft says 8.5M Windows devices were affected by CrowdStrike outage
techcrunch.com
techcrunch.com
Edit: I did some back of napkin math. ~30 million work for a fortune 500. Let's say 2/3rds of those have a Windows desktop provided by employer, so ~20M. I think I read crowdstrike has about ~25% market share, so that's 5 mil just in fortune 500. No way it's just 8.5M
Its not only big corps, but hospitals, governments, mom and pops.
Crowdstrike was baaaaaaad.
Front end or back end?
Because the backend hasn't been windows in most places for a very long while.
I do think there are a mountain of AD servers out there though. Not sure I care to quantify exactly how many makes a mountain. but I'd think 2 commas for sure. more than a million? less than 100 million? that seems like the right ballpark.
To compensate and keep the focus on them, as masters of all outages...They will take at least until Tuesday, (according to their own info...) to fix the current ongoing issue with Teams scheduling: https://portal.office.com/servicestatus
https://blogs.microsoft.com/blog/2024/07/20/helping-our-cust...
"CrowdStrike has helped us develop a scalable solution that will help Microsoft’s Azure infrastructure accelerate a fix"
What if I told you that such Mag7 speak is not to be trusted, at all, even?
This is unacceptable practice. I understand non tech media not getting it, but this lack of awareness from tech news is sad.
Yesterday was a catastrophe and you are still stuck with such naive and simplistic view: you want your antivirus to be auto updating.
The options aren't "everyone auto updates or no updates for weeks", there's a balance point. It's very clear what choice most critical companies this week did though.
I will give you that I highly doubt that a large number of these machines are anywhere near that critical nature, but there are some that will fall within that much risk.
What do you do, just not update to handle new risks? A lot of systems going down is really bad, don't get me wrong. But is it worse that you could be breached depending on the data (and other services) those systems may have access too?
To me this is a flaw in Crowdstrike but also Windows that this could happen in the first place, and a serious flaw on Crowdstrike's side that this somehow got out.
And yes I do acknowledge that much of this is security theatre, but I also would not be surprised if it does sometimes work.
Some, minor, blame falls on Windows due to its ability to BSOD as easily as it does.
As far as the companies, it is a tricky situation. Many of the companies have Crowdstrike enabled and automatic updates turned on to check some audit box. They have to keep the updates going out regularly.
We are well past the point in tech that a company is solely responsible for their systems with external dependencies being the norm. Either with the shared security model with cloud services like AWS or a reliance on external API's and servers. You have to trust the vendor you are working with for whatever critically important system is going to do their job. Could you look back and say that maybe you chose the wrong vendor for a specific piece of software, but this could have happened to other vendors.
Something that I am not entirely sure of is for those audit, compliance, etc requirements can they use an alternative update method. And this is something that would be different based on each compliance, but to the best of my knowledge for security software most want you to have automatic updates.
If this was the case of all of these servers going down because of a major AWS outage would you really be saying the companies are to blame?
This is an absurd take, specially after an outage who took down 911 response centers, hospitals and has millions of passengers still stranded.
You trust no vendor and assume everything fails all the time.
There might be smaller parts of your system you could say this, but unless your system is 100% airgapped, and all of the wiring, servers, etc are all put down by you and you are working with a LAN.
There are not many systems that fall within that definition. As soon as you hit using the internet for communication you are reliant on your ISP working. Maybe you can have a redundant connection, but then you have to assume both of those will do their job and that they don't have a dependency that could bring them both down.
So no, it's not absurd unless you are never going to the internet. You have to make the decisions on what your system relies on and what it can handle.
I fully understand what this brought down, but again there are plenty of other instances where you assume an outside company is going to do their job.
Looking back and saying, well maybe this was a bad idea because its an external dependency isn't helpful when we can point to any number of other external dependencies that may not have brought down as many systems but can just as easily bring down critical systems.
- You need more than one ISP
- You need diverse Operating Systems and Databases
- You deploy in phases with canary releases
- You don't deploy on Fridays....
How difficult can it be?
I addressed this in my previous response. It is still an external trust, even if you have redundancy.
> - You need diverse Operating Systems and Databases.
I have never ever seen a company run the same server side software deployed to multiple different operating systems.
> - You deploy in phases with canary releases.
As I mentioned in a previous post, there are going to be critical enough systems that may be under a serious threat of breach that any wait is not worth the risk.
Also as I have already mentioned, in many cases automatic updates is turned on for compliance reasons that may not allow what we think is common sense for the vast majority of software.
> - You don't deploy on Fridays....
I agree but to the best of my knowledge this was essentially a security definition updates not a code update. That is the kind of thing that you would push out when you have it otherwise your systems could be vulnerable over the weekend.
Disagree strongly. You are analyzing risk the wrong way. That is what I call: "Security by being on the latest patch"
Zero days occur every day and many are ongoing right now. Your antivirus vendor or OS vendor, needs hours to days, to weeks, to detected them, understand the attack, come up with a defense, test (hopefully...) the defense patch, deploy in phases (hopefully). So you are always many hours to days behind the latest threats and before getting such a protection.
The core idea here is "Critical System"
If the system is critical, it's security and robustness needs to rely on it's security architecture. Not "being on the latest patch". You will always be catching up to any threats.
Also you are still ignoring, that for many of these companies they have not have a choice due to compliance requirements.
That being said, so great maybe we can avoid this issue. But instead maybe next time instead it will be. "Well, you run security software X and when you were breached they had a protection out for this, why were you not up to date?"
The fact remains that what happened yesterday was an extraordinary situation that I highly doubt anyone seriously thought it was a serious risk. Since most people would safely assume that a vendor pushing security updates would do basic testing.
Also you are focusing on security when there are other dependencies that could bring down your system. That is my point here. We are focusing so much on how this one thing should have been done differently and that the companies are somehow to blame when this could have been any number of other things that would not have been as global of an impact but could still bring down major systems.
> Also you are still ignoring, that for many of these companies they have not have a choice due to compliance requirements.
They have a choice. They could run their system properly. You are arguing for reasons of compliance...When this incident is the clear demonstration being compliant has nothing to do with being secure and robust.
Its all PaaS/SaaS now, old-school properly engineered isolated solutions require too much expensive staffing.
I'm waiting for a vendor like zscaler to be hacked - what could go wrong with having thousands of companies do MITM SSL interception via a single vendor.
That's a nice juicy target for hackers if I ever saw one...
This is entirely on cloudstrike, or perhaps clown strike is more appropriate.
While many companies probably do that, it's usually not required if you can argue for an alternative approach and how it fits your risk appetite better (e.g. progressive updates on a routine schedule).
Would that be better?
I would pick automated testing and spread fleet deploys. There's no reason in any enterprise this should take more than 1-2 hours, which is a perfectly acceptable window of risk.
Businesses are under a constant barrage of cyber attacks, with goals to steal the data, encrypt it and then blackmail or sell all the data. Ransomware payouts exceeded $1 bil last year. And that doesn't include all the damage done besides the payouts.
Edit: Supposedly global cost of cybercrime is expected to reach $20 trillion+ by 2027.
I understand cybercrime is real, however I highly doubt the amount of real time RCE exploits leaked into the wild executed within 2 hours is > 0.01% of the updates pushed by CrowdStrike.
If there's a new pattern of social engineering/phishing attack it might be a question of hours to be able to respond to that and identify those specific patterns. Or just every minute will mean that more companies and machines will be compromised if there's a mass phishing campaign going on.
A typical solution would be to have two machines, one with the automatic updates and a second one without automatic updates that jumps in in case the first one breaks down.
Great, now the other one is still vulnerable and hackers can still steal information from it.
However that isn't popular and most orgs would prefer a day of downtime from this type of outage vs the hassle and cost of doing it right.
Hospitals should not loose their ability to provide care to sick people, just because of an misconfiguration of an antivirus. That is as bad as airplanes crashing because of a lack of redundancy and management of risks.
Only stops script kiddies, at best.
crowdstrike said the update was a "configuration file".
https://www.crowdstrike.com/blog/technical-details-on-todays...
Either way, it's unacceptable for critical services to be beholden to the validation of a third party upstream. The companies in question are responsible for that negligent handing off of ownership.
I don't think anyone is going to disagree that's engineering best practice and should theoretically be done, but how is microsoft going to enforce this? Do you want to force developers wanting to publish software for windows to undergo annual audits (soc-2 style) to confirm that all the engineering best practices are indeed being followed? Not even Apple is that strict.
This testing is the responsibility of the company whose computer fleet it is. They have many upstream software vendors - often 10s or 100s - and should be doing this testing every time. You should never rely on the vendor to test (evidently). I'd go so far as to assume all vendor updates are hostile and build your test model against that.
Automated testing of new software in companies with 10k+ desktops (which covers most affected companies here) should be as common as password policies or email attachment policies.
If the vendor implements things in a way that doesn't allow this style of testing, they don't meet security requirements and another should be found.
That is also pretty much the Apple way nowadays.
"how is microsoft going to enforce this? Do you want to force developers wanting to publish software for windows to undergo annual audits (soc-2 style) to confirm that all the engineering best practices are indeed being followed? Not even Apple is that strict."
Or are you saying that they should ban EDR vendors from installing drivers at all? How are you going to implement the invasive monitoring needed for EDR to work?
And you think Crowdstrike's driver isn't signed? Given all drivers have to be signed for windows to load it, I highly doubt that's the case. Moreover, I doubt WHQL's testing covers logic bugs. Graphics drivers crash all the time for instance, and they're definitely WHQL certified. You could inspect the code even harder, but that just goes into my previous question.
That's the irony of the situation. The criticality of the systems (arguably) necessitates real-time updates, otherwise they'd be vulnerable to threat actors.
This is an Oxymoron
The first time I ever rolled out Falcon, the sales engineer said, “if you want to be on the latest when it releases, choose this policy. Generally customers like to be one release (N -1) behind. This is the safest option in my experience. We rarely have issues but this is the way to prevent issues if we do ship something bad.”
I’ve been telling other admins this is the safest option moving forward. I don’t see a need for my org to run bleeding edge releases of newer products. This also applies to OS updates unless it’s a zero day. Major OS releases I wait for the first .1 update to release. Currently doing this with Ubuntu Desktop 24 LTS as it shipped with missing features from 22 and a broken autosetup functionality. August is the first update to 24 LTS and we’ll test and determine if the bugs have been squashed.
I can’t think of any way to always be on the latest upgrade of anything critical. All of these companies were on the bleeding edge release of CrowdStrike and it brought a lot down globally.
https://news.ycombinator.com/item?id=41015038
> "b) Since n, n-1 and n-2 versions of the sensor all died equally spectacularly, that bug as been around for at least three versions of csagent.sys."
There's so much misinformation around this Crowdstrike issue. The change deployed was in what is referred to as a "channel file" which isn't part of the software update mechancism (what you call N-X) but part of the intra-day frequent signature/channel updates it gets (that we all have no control over).
Crowdstrike are calling it an unfortunate "logic error" but they and few others are talking about the how a binary payload could get released to the public without seemingly any pre-release testing of the payload. If the content that was made available to the public had ran on a test endpoint, they would have discovered this "logic error" before taking down a high number of the world's systems simultaneously.
Another possible source is Crowdstike itself who definitely has the data.
Hospitals - physicians/doctors/nurses lost access to critical equipment. Patients may have suffered degraded care as well. Reports of this outage impacting active surgeries. Patients forced to reschedule appointments around ClownStrike
Airlines - many flights grounded. Delays, delays, delays. Wasted fuel, time. Loss of revenue due to rescheduled flights, refunding customers. Local airports flooded with grounded flights, increased personnel to deal with it. FAA stressed.
Banks - many people lost access to money. Frustration for people trying to get access to pay bills, or get paid themselves.
"This is the 2nd time CrowdStrike CEO George Kurtz has been at the center of a global tech failure" - https://www.businessinsider.com/crowdstrike-ceo-george-kurtz...
We need jail time for these executives otherwise nothing changes
I've heard of 250k employee companies where people got a snow day off this.
That is about -.1% of all the MS machines.
As a linux user, I dont understand the big deal, the effects of this.
As a linux user, I don't see how you don't see this was a huge deal.