What everyone seems to be missing in the CrowdStrike incident
blog.axantum.com
blog.axantum.com
But in a recent update of https://www.crowdstrike.com/falcon-content-update-remediatio... they have a long explanation of of things are supposed to work, with a lot of nice words (sounds almost AI-written...) and quite a few implications that it's really the customers fault who have not configured their systems to for example stay one version behind the latest and still a very short explanation of what went wrong.
But... Lo and behold, what are CrowdStrike going to do to avoid this happening in the future?
"Implement a staggered deployment strategy for Rapid Response Content in which updates are gradually deployed to larger portions of the sensor base, starting with a canary deployment." About time...
My main point, and the reason for the title, is that this has not been the major takeaway in main stream media analyses. Of course not "everyone" has missed this, but pretty much all media articles about the incident do appear to miss this.
Most people didn't even know what Crowdstrike was, let alone understand the concept of testing updates and staggering them.
Lastly, the media are at risk of reporting the wrong thing and being a target of litigation. Therefore they report often in hyperbole and without much factual information until the facts are determined.
>I've participated in many roll outs, and never would I allow a big-bang roll out like this. CrowdStrike should be charged with negligence for having this type of process. It's just plain irresponsible.
Agree with all of this. Related to deployment process or lack of one, the hour of the deployment has always struck me from the beginning. The largest impact was in the United States yet an update was pushed very early hours US time.
Presumably the off-hours deployment wasn't because of lowering the potential impact as they sent it to everyone.
This isn't some UI makeover that they can push until next Tuesday. They're pushing updates to the detection logic for what could be an evolving threat, so odd timing of the update is at least somewhat justified. Do you really want a botnet to rip through corporate networks over the weekend while you wait for a Tuesday deploy?
This isn’t a indie game dev or some group of volunteers
CS actively sells and uses fear to gain customers to a bad product. Their whole 3B business is based on unearned trust
IMO that constitutes fraud if it’s this magnitude. Someone needs to go to jail and investors need to lose all their money and liquidate the company
Looking at all the responses to the incident advocating for more centralised control, it almost seems like it was a deliberate provoking of the acceleration towards digital totalitarianism. "The only thing we have to fear, is fear itself."
If it's not - all an attacker would have to do is to deposit a file in %WINDIR%\System32\drivers\CrowdStrike with a name such as C-00000291.sys containing zeros - and the system becomes unbootable without manual intervention!*
An attacker who has already gained enough permissions to do that can just "delete system32" instead, or worse.
That seems a bit dramatic. I don’t do big corporate IT but I thought a lot of corporate IT shops have the ability with Microsoft to choose what updates are pushed out to computers on their domain. If so, then something like that could have prevented it, presuming they have the ability to allow a single computer or small group to receive the update to confirm it works successfully.
I'm getting to the point where I'm not willing to implement these types of systems for anyone anymore. Not after seeing the breadth of data hoovering and consolidation being pursued.
At some point I just realized the only thing preventing these systems being used in the ways I dread is y'all being decent.
...I'm not willing to cut that check anymore. Seen too much.
CrowdStrike has stated that no the crash was not related to the file of zeros.
> CrowdStrike states "Boot into safe mode. Delete C-00000291*.sys." That's the file(s) with the zeroes
That's potentially multiple files, but do we know only one comprises just zeros?
That's not quite the same as saying "This is not related to Channel File 291 containing all nul bytes."...
I don't have first to hand knowledge here, but rely on Dave Plummer's statement.
Regardless of zeroes or single files or not, the fact is that bad data in C-00000291.sys in combination with bad validition in the driver causes it to crash. Deleting C-00000291.sys causes the driver to stop crashing.
Anyway, my main point isn't really about this. It's about the big bang global roll out simultaneously to at least 8.5 million systems in one go that's irresponsible.
The driver architecture is the lesser evil here, although it's bad enough!
I think we've seen no evidence that data is to blame.
> Deleting C-00000291.sys causes the driver to stop crashing.
So perhaps just its existence is to blame.
> The driver architecture is the lesser evil here
Except if the crash had been limited to the driver, it would have left the machine running unprotected which is far greater an evil.
Why is it so hard for manufacturers to just go ahead and explain what really went wrong, without a lot of corporate b..t? Probably, if they do really say what happened in so many words they might open themselves for negligence lawsuits. Hopefully somebody files one anyway. The industry needs to learn to be better, and the only thing that talks loudly enough is probably money. Lost revenue, liability damages, and share holder value loss.
This is, in fact, not a fact. We really don't know yet.
CrowdStrike blue screened one of my laptops twice right as the incident was getting started, before a fix was available. There was no boot loop in my case. I was back up and in the middle of an episode of Breaking Bad the second time it got me, 30 minutes after the first. Did the agent wait that long to load a content update it had already loaded before? Maybe, but it's at least as likely that the content was loaded the whole time, and that some activity pattern set it off. Thus, I'm skeptical of the problem being simple content validation.
Rules in many industries are often written in blood, whilst not the case here (I would hope), these incidents are what sparks change in many cases. We now have large companies all around the world bearing the brunt of not testing an update before putting it into production. Whether it was because of them or reliance on a third party, executives en masse are now aware of what happens if you don't do these things.
When we talk about promoting changes into production before testing, we will all come back to this moment
Sure there's the defense in depth argument, but this is too in depth, and as proven today not without its risks.
This seems to be saying Microsoft is letting Crowdstrike to bypass Microsoft's own security measures for Windows.
That rings true, too many times in this industry, easy wins over security. I think the only place this does not happen is on OpenBSD.
Aside the biggest issue is still having in 2024 non-declarative, non-rollbackable systems in productions, no LOM, no easy massive automation.
Whether or not "Rapid Response Content" and "Template Instances" are Turing complete is unclear, but the fact of the matter is that according to CrowdStrike "problematic content in Channel File 291 resulted in an out-of-bounds memory read triggering an exception", so the interpretation of the content is at least fairly complex. CrowdStrike also states that "Each Template Instance maps to specific behaviors for the sensor to observe, detect or prevent", and mentions a "Content Interpreter". Whether it's code or configuration data is not really relevant though, the point is that it's interpreted by a kernel mode driver which did not have sufficient validation of it's "Content" to prevent the crash.
What could go wrong…
What is definitely known is that a WHQL kernel mode driver from CrowdStrike crashes, and removing a single file external to the driver causes it to stop crashing. Some pretty sure conclusions can be drawn from that. No "probably" required.
What I dont understand is why also didn't these companies just role back to the previous known working image when servers failed to reboot? please say they are not all allowing auto updates without testing and no backups.
Here's the section where Dave talks about it: https://youtu.be/wAzEJxOo1ts?si=aCX8pOTP0D_IRNAx&t=670