Bugs in code I can almost always understand and forgive, even the ones that seem like they’d be obvious with hindsight. But this is just an egregious lack of the most basic rollout standards.
Bugs in code I can almost always understand and forgive, even the ones that seem like they’d be obvious with hindsight. But this is just an egregious lack of the most basic rollout standards.
If any one of the five points above hadn’t happened, this event would have been avoided. However, if number 1 had been addressed - any of the others could have happened (or all at the same time) and it would have been fine.
I understand that we should assume that bugs will be present anywhere, which is why staggered deployments are also important. If there had been staggered deployments, the. The damage would have happened, but it would have been localized. I think security people would argue against a staged deployment though, as if it were discovered what the new definitions protected against, an exploit could be developed quickly to put those servers that aren’t in the “canary” group at risk. (At least in theory — I can’t see how staggering deployment over a 6-12 hour window would have been that risky).
There may not be out of the box fuzzers that test device drivers so you hoist all the parser code, build it into a stand-alone application, and fuzz that.
Likely this is a form of technical debt since I can understand not doing all of this day #1 when you have 5 customers but at some point as you scale up you need to change the way you look at risk.
That goes equally if it was a Windows Update rolled out in one motion that broke the falcon agent/driver, or if it was Crowdstrike. There is almost no excuse for a global rollout without telemetry checks, whether it's security agent updates or os patches.
And even testing can't be trusted 100%, because writing code that does the right thing and code that tests things correctly are about equally hard, they just aren't always hard simultaneously.
Policy changes seem more reliable and would catch other, as of yet unknown classes of bugs.
What looks especially bad for Crowdstrike is how many things (relatively simple things) had to fail in order for this to slip through. It's like walking into Fort Knox, grabbing a gold bar, and walking out unimpeded. A complete systemic failure.
But is there any responsibility for the clients consuming the data to have verified these updates prior to taking them in production? I haven't worn the sysadmin hat in a while now, but back when I was responsible for the upkeep of many thousands of machines, we'd never have blindly consumed updates without at least a basic smoke test in a production-adjacent UAT type environment. Core OS updates, firmware updates, third party software, whatever -- all of it would get at least some cursory smoke testing before allowing it to hit production.
On the other hand, given EDR's real-world purpose and the speed at which novel attacks propagate, there's probably a compelling argument for always taking the latest definition/signature updates as soon as they're available, even in your production environments.
I'm certainly not saying that CrowdStrike did nothing wrong here, that's clearly not the case. But if conventional wisdom says that you should kick the tires on the latest batch of OS updates from Microsoft in a test environment, maybe that same rationale should apply to EDR agents?
In the boolean sense, yes. United Airlines (for example) is ultimately responsible for their own production uptime, so any change they apply without validation is a risk vector.
In pragmatic terms, it's a bit fuzzier. Does CrowdStrike provide any practical way for customers to validate, canary-deploy, etc. changes before applying them to production? And not just changes with type=important, but all changes? From what I understand, the answer to that question is no, at least for the type=channel-update change that triggered this outage. In which case I think the blame ultimately falls almost entirely on CrowdStrike.
I would say on the client for buying into CrowdStrike.
And also the client for having no contingencies and just accepting a vendor pinky-swear as meaningful.
CrowdStrike failed at their responsibilities too, I just mean that so did everyone else.
When you cede your own responsibilities to someone else and don't have that backed up with contractually enforced liability to make you whole when they fuck up, and also don't provide your own contingency so it doesn't really matter what some vendor does, that's on you. That's 100% entirely on you and it doesn't matter if a million other people also did the same utterly thoughtless and lazy thing.
I understand this perspective but I think it misses the forest for the trees. You have to evaluate this kind of stuff in context. Purity tests really smack on tech message boards where nobody has any accountability to any kind of business requirements, but basically no real-world organization operates in that way, so it's all a bit irrelevant.
> When you cede your own responsibilities to someone else ...
This framing is a bit naive, I think. It isn't a boolean. Everything is about risk management, cost/benefit analysis.
Honestly, it hadn't even occurred to me that software like this marketed at enterprise customers wouldn't have this kind of control already available. It seems like an obvious thing that any big organization would insist on that I just took it for granted that it existed.
Whoops.
I used to work with regional parks and recreation departments, and they would not approve any updates that did not go through UAW environments that we had set up. All updates had to be deployed to their UAW, thoroughly tested, before going to their production environment.
I get this this is slightly different, but I'd imagine Airlines, Banks, and Hospitals would have far more strict UAW policies to avoid a single vendor from kneecapping operations.
They do, but this update bypassed all of those rules.
No one at my company apparently.
I don't know how you could assert that this is impossible, hence channel files should be treated as code.
My company had a lot of Azure vms impacted by this and I'm not sure who the admin was who should have tested it. Microsoft? I don't think we have anything to do with crowdstrike software on our vms. ( I think - I'm sure I'll find out this week.)
Edit: I just learned the Azure central region failure wasn't related to the larger event - and we weren't impacted by the crowd strike issue - I didn't know it was two different things. So my second part of the comment is irrelevant.
Exactly which team owns the testing is probably left up to each individual company to determine. But ultimately, if you have a team of admins supporting the production deployment of the machines that enable your business, then someone's responsible for ensuring the availability of those machines. Given how impactful this CrowdStrike incident was, maybe these kinds of third-party auto-update postures need to be reviewed and potentially brought back into the fold of admin-reviewed updates.
Which is also why you see every single customer affected - what you are suggesting is simply not an available thing to do at present for them.
At least for now - I imagine that some kind of staggered/slowed/ringed option will have to be implemented in the future if they want to retain customers.
what incentivized the bad decisions that led to this catastrophic failure?
It makes sense in a way - given their fast growth strategy (from nowhere to top 3) and desire to “do things differently” - the iconoclast upstarts that redefine the industry.
Or to summarise - hubris.
The "how" here is AV definition or a way to identify the attack. In CS-speak: content.
Catching 0day quickly results in good reputation that your EDR works well.
If people turn off their AV definition auto-update, they are at-risk. Why use EDR if folks don't want to stop attack quickly?
This is one bsod on windows 10. I saw another kernel panic on specific Linux distro.
What else?
One thing that is funny is that quite a few of their competitors are taking this opportunity to shit on them via Twitter and by marketing themselves as better than CrowdStrike.
Twitter, with all its issues, apparently has a feature to prevent fake news and that feature will show crowd source sentiment to debunk fake news, in this case, Twitter users showed how many times Crowdstrike competitor BSOD windows
There's best practice and then there's customer
The solution to customer reluctance to being the guinea pig is not to force every customer to be a guinea pig.
The red hat one. But they also did it to debian with a different issue, and I think another distro as well.
I'm sorry but this is the customer's fault.
If I'm using your services you work for me and you don't get to bully me into doing whatever you think needs to be done.
People that chose this solution need to be penalized, but they won't.
Compliance also has to share some of the blame here, if best practices (local testing) aren’t allowed to be followed in the name of “security”.
Many don’t have a choice, a lot of compliance is doing x to satisfy a checkbox and you don’t have a lot of flexibility in that or you may not be able to things like process credit cards which is kinda unacceptable depending on your company. (Note: I didn’t say all)
CrowdStrike automatic update happens to satisfy some of those checkboxes.
I never thought such stories were real until I encountered them…
No, for exactly the reason we just saw, and the same reason why vaccines are tested before widespread rollout.
As an example, reaction time is paramount to counter many kinds of attacks - thats why blocklists are so popular, and AS blackholing is a viable option.
Agreed. It's crazy that the top tech companies enforce this in a biblical fashion, despite all sorts of pressure to ship and all that. Crowdstrike went YOLO at a global scale.
Is there anything we can take from other professions/tradecraft/unions/legislation to ensure shops can’t skip the basic best practices we are aware of in the industry like staged rollouts? How do we set incentives to prevent this? Seriously the App Store was raking in $$ from us for years with no support for staged rollouts and no other options.
I'd assume that sort of thing would be covered in the EULA and contract -- but even if it weren't, it seems like allowing customers to define their own definition update strategy would give them a pretty compelling avenue to claim non-liability. If CrowdStrike can credibly claim "hey, we made the definitions available, you chose to wait for 2 weeks to apply them, that's on you", then it becomes much less of a concern.