Pipeline 1 --
Code updates to their software are treated as material changes that require non-production and canary testing before global roll-out of a new "Version".
Pipeline 2 --
Content / channel updates are handled differently -- via a separate pipeline -- because only new malware signatures and the like are distrubuted via this route. The new files are just data files -- they are supposed to be in a standard format and only read, not "executed".
This pipeline itself must have been tested originally and found tobe working satisfactorily -- but inside the pipeline there is no "test" stagethat verifies the integrity of the data fine so generated, nor - more importantly - checking if this new data file works without errors when deployed to the latest versions of the software in use.
The agent software that reads these daily channel files must have been "thoroughly" tested (as part of pipeline 1) for all conceivable data file sizes and simulated contents before deployment. (any invalid data files should simply be rejected with an error ... "obviously")
But the exact scenario here -- possibly caused by a broken pipeline in the second path (pipeline 2) -- created invalid data files with some quirks. And THAT specific scenario was not imagined or tested in the software version dev-test-deploy pipeine (pipeline 1).
If this is true --
The lesson obviously is that even for "data" only distributions and roll-outs, however standardized and stable their pipelines may be, testing is still an essential part before large scale roll-outs. It will increase cost and add latency sure, but we have to live with it. (similar to how people pay for "security" software in the first place)
Same lesson for enterprise customers as well -- test new distributions on non-production within your IT setup, or have a canary deployment in place before allowing full roll-outs into production fleets.
It was mentioned in one of the HN threads, that the update was pushed overriding the settings customer had [1]. What recourse any customer can have in in such a case ?
Sue them and use something else.
We got this update pushed right through.
But the problem here is that the code runs in kernel mode. As such any data that it may consume should have been tested with the same care as the code itself which has never been the case in this industry.
And of of course that cost would be absolutely insignificant relative to the potential risk...
I’m not sure why you find that hard to believe - based on the (admittedly fairly limited) evidence we have right now, it’s highly unlikely that this deployment was tested much, if at all. It seems much more likely to me that they were playing fast and loose with definition updates to meet some arbitrary SLAs[1] on zero-day prevention, and it finally caught up with them. Much more likely than somehow every single real-world pc running their software being affected but their test machines somehow all impervious.
[1] When my company was considering getting into endpoint security and network anomaly detection, we were required on multiple occasions by multiple potential clients to provide a 4-hour SLA on a wide number of CVE types and severities. That would mean 24/7 on-call security engineers and a sub-4-hour definition creation and deployment. Yes, that 4 hours was for the deployment being available on 100% of the targets. Good luck writing and deploying a high-quality definition for a zero day in 4 hours, let alone running it through a test pipeline, let alone writing new tests to actually cover it. We very quickly noped out of the space, because that was considered “normal” (at least to the potential clients we were discussing). It wouldn’t shock me if CS was working in roughly the same way here.
You do not deploy anything, ever on your entire production fleet at the same time and you do not buy software that does that. It's madness and we're not talking about small companies with tiny IT departments here.
Source: https://x.com/patrickwardle/status/1814367918425079934
Note how the incident disproportionally affected highly regulated industries, where businesses don't have a choice to screw "best practice".
The problem here is that this type of update (a content update) should never be able to cause this however badly it goes. In case the software receives a bad content update, it should fail back to the last known good content update (potentially with a warning fired off to CS, the user, or someone else about the failed update).
In principle, updates that could go wrong and cause this kind of issue should absolutely be deployed slowly, but per my understanding, that’s already the practice for non-content updates at CrowdStrike.
You do not know if a content update will screw you over and mark all the files of your company as malware. The "It should never happen" situations are the thing you need to prepare for, the reason we talk about security as an onion, the reason we still do staggered production releases with baking times even after tests and QA have passed...
"But it's cybersecurity" is not a justification. I know that security departments and IT departments and companies in general love dropping the "responsibility" part on someone else, but in the end of the day the thing getting screwed over is the company fleet. You should retain control and make sure things work properly, the fact those billion dollar revenue companies are unable to do so is a joke. A terrible one, since IT underpins everything nowadays.
The CS customer has decided to update whenever 24/7 CS says. The alternative is to arrive on Monday morning to an infected fleet.
The CS customer has decided to offload the responsibility of its fleet to CS. In my opinion that's bullshit and negligence (it doesn't mean I don't understand why they did it), particularly at the scale of some of the customers :)
Incorrect, I believe, given they did not and could not get advance sight of the offending forced update.
Incorrect, I believe, given they could and did not get advance sight of the offending forced update.
Companies choose to work with Crowdstrike. One of the reasons they do that is ‘hands-off’ administration-let a trusted partner do it for you. There are absolutely risks of doing it this way. But there are also risks of doing it the other way.
The difference is, if you hand over to Crowdstrike, you’re not on your own if something goes wrong. If you manage it yourself, you’ve only got yourself working on the problem if something goes wrong.
Or worse, something goes wrong and your vendor says “yes, we knew about this issue and released the fix in the patch last Tuesday. Only 5% of your fleet took the patch? Oh. Sounds like your IT guys have got a lot of work on their hands to fix the remaining 95% then!”.
I am sympathetic to that, but its only possible if both policy and staffing allow.
for policy, there are lots of places that demand CVEs be patched within x hours depending on severity. A lot of times, that policy comes from the payment integration systems provider/third party.
However you are also dependent on programs you install not autoupdating. Now, most have an option to flip that off, but its not always 100% effective.
We are not talking about small companies here. We're talking about massive billion revenue enterprises with enormous IT teams and in some cases multiple NOCs and SOCs and probably thousands consultants all around at minimum.
I find it hard to be sympathetic to this complete disregard of ownership just to ship responsibility somewhere else (because this is the need at the of the day let's not joke around). I can understand it, sure, and I can believe - to a point - someone did a risk calculation (possibility of crowdstrike upgrade killing all systems vs hack if we don't patch a CVE in <4h), but it's still madness from a reliability standpoint.
> for policy, there are lots of places that demand CVEs be patched within x hours depending on severity.
I'm pretty sure leadership when they need to choose between production being down for an unspecified amount of time and taking the risk of delaying (of hours in this case) the patching will choose the delay. Partners and payment integration providers can be reasoned with, contracts are not code. A BSOD you cannot talk away.
Sure, leadership is also now saying "but we were doing the same thing as everyone else, the consultants told us to and how could have we have known this random software with root on every machine we own could kill us?!" to cover their asses. The problem is solved already, since it impacted everyone, and they're not the ones spending their weekend hammering systems back to life.
> However you are also dependent on programs you install not autoupdating. Now, most have an option to flip that off, but its not always 100% effective.
You choose what to install on your systems, and you have the option to refuse to engage with companies that don't provide such options. If you don't, you accept the risk.
And if an attacker does??
- Lack of testing of a deployment - Lack of required procedures to validate a deployment - Engineering management prioritizing release pace over stability/testing - Management prioritizing tech debt/pentests/etc far too low - Sales/etc promising fast turnarounds that can’t be feasibly met while following proper standards - Lack of top-down company culture of security and stability first, which should be a must for any security company
This outage wasn’t caused only by “the intern pushing release.” It was caused by a poor company culture (read: incorrect direction from the top) resulting in a lack of testing of the program code, lack of testing environment for deployments, lack of formal deployment process, and someone messing up a definition file that was caught by 0 other employees or automated systems.
https://news.ycombinator.com/item?id=%2041002864
Very much not surprised to see this now.
Now, if these sorts of things were battle tested before release, and had a (ideally decade+-long) history of stability with well-documented processes to ensure that stability, you can more easily make the argument that it’s worth it. None of those things are close to true though (and more than likely will never be for any AV/endpoint solution), so it is very hard to justify this sort of configuration.
Store something like an `attemptingUpdate` flag before updating, and remove it if the update was successful. Upon system startup, if the flag is present, revert to the previous config and mark the new config bad.
This is why we can’t have nice things, but maybe we just don’t want them anyway? “Mistakes will be made” is way less true if you actually put the effort in to prevent them, but I am beginning to think this has become code for quiet-quitters to telegraph a “I want to get paid for no effort and sympathize with others who feel the same” sentiment and appear compassionate and grimly realistic all at the same time.
yes, billion dollar companies are going to make mistakes, but almost always because of cost cutting, willful ignorance, or negligence. If average people are apologizing for them and excusing that, there has to be some reason that it’s good for them.
I don’t mind meetings but being in a 4 hour emergency meeting because some due diligence wasn’t done is a waste of my time.
Life is easier when you do good work.
This would have been poor, but to have released it with no testing would have been the most staggering negligence.