CrowdStrike fixes start at "reboot up to 15 times", gets more complex from there
arstechnica.com
arstechnica.com
As in, what exactly is wrong in these C00000291-*.sys files that triggers the crash in csagent.sys, and why?
And in this case, it only crash. But if it somehow read value from position it isn't supposed to successfully? You have an RCE.
I know for sure at least macOS and OpenBSD don't
It's possible to make userspace interfaces and permissions for this sort of thing.
Seriously though - this entire outage is the poster child for why you NEVER have software that updates without explicit permission from a sysadmin. If I were in congress, I would make it illegal, it's an obvious national security issue.
If a code path isn't followed until a config file updates, that is practically the same thing as the code path being introduced by the update.
In general I agree, but this case is quite messy. It's more like your anti-virus had a bug since forever that if it loads a broken virus definition it bricks your system. And a broken virus definition finally happened today.
Do you want every virus definition (that is updated every few hours) to require explicit permission from a sysadmin?
Who's going to take the risk of appearing to have sat on an important update, while the org they support is ravaged by ThreatOfTheDay, because they thought they knew better than a multi-billion dollar, tops-in-their-field company?
(I'm not necessarily saying that's actually objectively correct, but I can't imagine that many folks are willing to risk the downside)
I consulted for a company for a while where the 'sysadmin' was the owner's mother - who bought laptops from walmart. Not only could she NOT have approved updates like this, even if she could she would have she wouldn't have had any knowledge whatsoever with which to make a determination if it worked.
In an abstraction, the problem really is with externalities. These approaches to updates exist because people who CAN'T do what you describe are likely a more dominant part of the threat model than this happening to people you do describe. The resulting fix, as we're seeing, is very reliable until it isn't...and if the isn't is enormous in scale the systems aren't setup to fail gracefully.
If you want to make a rule...require graceful failure.
Kernel level code blindly loading arbitrary files?
Panicking when the file doesn’t parse because it’s not a memory safe language?
Not validating the files before loading them?
Not validating the files before SHIPPING them? No CI? No safety net?
No staged rollout in case of explosion?
There are far FAR bigger mistakes here than “sys admin didn’t have to press button”.
This is a faulty and dangerous product from conception to execution.
Unfortunately "security" folk these days are box ticking fuckwits and this product brief ticked all the boxes. They do not understand any more traditional methodologies other than "install these magic beans and action the reports".
Invest in better software and network architecture and DR strategy instead.
Whether a program panics or recovers when attempting to parse bad data is entirely orthogonal to memory safety. Do you have any in-depth technical information about the bug itself that you're basing this on?
I agree with the rest, especially the use of a memory unsafe language to do parsing in the kernel by a billion dollar security company blows my mind.
How can you even run a security company without any security professionals reading your code even incidentally? An impressive level of incompetence.
I lean towards breaking things being the bigger risk.
But if even a handful of the other errors were corrected this would have been prevented and they wouldn’t have had to make that choice.
How the heck they didn't find out the new version prevent the computer from booting at all?
The really important machines are still on Win 3.1.
It's not like they'd read the source code or examine every file that's been changed or downloaded for a proprietary kernel module for every crowdstrike update (there must be a LOT of them).
They're the same team who mandate you use a 3-year-old browser version and 5-year-old OS, because you can't be trusted to manage your own updates, so they do know the idea.
Again, the same team in most enterprises wouldn't dream of letting you have an auto- updating Firefox Nightly, they know how to configure software so it doesn't phone home for updates or is blocked from phoning home.
This example is probably an argument for not running windows on critical systems due to insufficient focus on security from the beginning which has lead to a need for things like crowdstrike.
They do make a version of CS for Linux but nobody runs it unless they’re forced to by overzealous compliance drones.
I wish people would stop making blanket statement as if they know how every company in the world runs. Plenty of Linux machines are running CS, and it's not only because they are forced to for compliance. NG AV has been picking up speed as a "just in case" thing for Linux and Mac for years now. Your anecdote does not apply to everyone.
They exist solely to tick the box. That’s it. Nobody who pushes for them gives a shit about security or anything that isn’t “our clients / regulators are asking for this box to be ticked”. The box is the problem. Especially when it’s affecting safety critical and national security systems. The box should not be tickable by such awful, high risk software. The fact that it is reflects poorly on the cybersecurity industry (no news to those on this forum of course, but news to the rest of the world).
I hope the company gets buried into the ground because of it. It’s time regulators take a long hard look at the dangers of these pretend turnkey solutions to compliance and we seriously evaluate whether they follow through on the intent of the specs. (Spoiler: they don’t)
If I can't commit code to our app without a branch, pull requests, code review...why can the infrastructure team just send shit out willy-nilly?
"Always allow new updates" must have been checked, or someone just goes through a dashboard and blindly clicks "Approve"
And security vendors should follow "secure by design" principles. Yes, I know a try-fucking-catch might be too advanced, and uh oh kernel code is hard because unwinding is costly. But guess what else is also not cheap. (Okay, I seriousness failed.) But still. This is fair and square in the "this should never happen" scenario. It's an automatically downloaded plugin or whatever. (CS can call it "content update", but von Neumann is already calling FedEx to send them a pallet of industrial grade bitchslap.) And if the plugin loader cannot gracefully fail plugin loading, then it should obviously come with the appropriate audiovisual cues[1] so sysadmins know what to expect.
I think the team writing the parsers for these data files deserves some blame. This should have been fuzzed, property tested, etc.
But maybe anti-malware is given a blind eye because instant updates for zero day security issues are obviously attractive.
Still, though... In hindsight it's not workable for especially anything running system drivers with liberal kernel access.
It’s not unheard of for things to slip by testing and CI.
And the Crowdstrike CTO has either been given the ammunition to get __whatever they ask for, ever again__ with regard to appropriate allocation of resources for devops *or* they'll be fired (whether or not it's their fault).
And let me be very clear. This is absolutely, positively and wholly not the person that pressed the button's fault. Not even a little. At a company as integral as CrowdStrike, the number of mistakes and errors that had to have happened long before it got to "Joe the Intern Press Button" is huge and absurd. But many of us have been in (a much, much, *MUCH* smaller version of) Joe's shoes, and we know the gut sinking feeling that hits when something bad happens. A good company and team won't blame Joe and will do everything they can to protect Joe from the hilariously bad systemic issues that allowed this to happen.
Shit happens. We learn.
You know that, and I know that. The people who will ruin his life starting today do not know (or care).
This is part of the job of a senior.
I see my orgs SCCM admins have been consulted
Don't do that, or you'll be dragged before the greatest obnoxious and self-aggrandizing body in the world for lengthy dressing down that probably affects the stock price.
Supposedly they have all kinds of certifications but not even having basic QA demonstrates that this is all just a smokeshow: https://www.crowdstrike.com/why-crowdstrike/crowdstrike-comp...
I thought it was BSOD'ing on boot? I don't understand how this works. It auto-updates on boot? From the internet?
Apparently!
To be honest I prefer that over the #nix way of doing things. In Windows, you have exactly one file any given path can refer to - in Linux or Mac, it may depend on which directory's inode is seen as the root node by your process (e.g. chroot or container), or whether mounts are at play, or a file/directory got deleted and replaced by something else.
Particularly the last scenario keeps tripping me every once in a while.
Yeah, that's a great idea and not at all a huge attack vector.
It is not unreasonable to think that this sort of software could get compromised.
The BSOD is because one of the data files that they previously pushed is horribly mangled, and their driver explodes about it. But if you get lucky, the driver can receive an update notification on boot, connect to the separate file server, and finish overwriting the broken file on disk before the rest of the driver (that would crash) has loaded the broken file
And they do all of that very early on boot. The justification being that you don't want the antivirus to start booting after a rootkit has already installed itself
One might ask whether an anti-virus really needs to run inside the kernel, but the answer might reasonably be yes.
So, this isn't really what's getting on my nerves here. Just how it auto updates and get pushed throughout the organizations without a smidge of quality assurance. Smaller businesses... Sure, I get it. They don't have the resources to set up infra for this, but those... airliners... and hospitals. WTF. I read some org thinking they might not even be able to provide anesthesia. Seriously. What.
It is also possible to access hardware or any other low-level system resources from unprivileged user code, if its process has been granted appropriate access rights by the kernel.
This second solution requires more work, but it is much more secure as the access can be limited to only the strictly-required resources and system crashes become impossible.
The extreme of this solution is a micro-kernel operating system, but there is no need for extremes. Even in a Windows or Linux system you can use this method. You can have a very reduced privileged code in a driver or kernel module, which does nothing except providing access to the permitted resources. Then anything like attempting to access not mapped memory would happen in user code and it would crash only the user process, not the entire computer system.
"the truth is everything is breaking all the time, everywhere, for everyone"
I feel like we are better off running open-source software. Everyone can see where the mistakes are instead of running around like a chicken with its head cut off.
You bet they have an amazing perfect top-notch hiring pipeline, many rounds of interviews, and whatever you could wish for! (No, no ... the subcontractors writing code are not in scope for this, duh.)
Do not, under any circumstances let a job impact your health or mental well being.
At the time, it was not just a job. It was a passion with a bar rising much faster than I could rise to the occasion. Simultaneously, my personal life was slowly falling apart, from family and loved ones in need, and the result was eventual failure leading to me being terminated. Luckily, it was one of the best events that has ever happened to me. I was able to land in a much better role almost immediately, which eventually catapulted my career and assisted in me being able to become financially independent as well as pivot into a domain with immensely improved work life balance. Importantly, I recognize I got lucky. It could’ve easily gone the other way, with me giving up both professionally and personally (yeeting myself from this plane of existence).
So, I not only violently echo your comment to others who come across this thread, I will go further to say that sometimes when you’re going through hell, if you keep going, there is light at the other end. It is just a job, it is okay to ask for help, and failure is when you stop trying to get back up, not when you get knocked down.
I totally haven't experienced this before and am not bitter in the slightest.
[1] https://edition.cnn.com/2024/07/15/tech/russian-firm-kaspers...
It’ll probably turn out that this update was pushed out against the strident, loud warnings of some small dev group within the company, and overruled by the all-knowing managerial class to keep up an OKR. They’ll have been warned six ways to Sunday but...
I’d definitely be not be the one pushing the big red button.
Automatic updates should be considered harmful. At the minimum, there should be staged rollouts, with a significant gap (days) for issues to arise in the consumer case. Ideally, in the banks/hospitals/... example, their IT should be reading release notes and pushing the update only when necessary, starting with their own machines in a staged manner. As one 90ies IT guy I worked with used to say "you don't roll out a new Windows version before SP1 comes out"
SkyNet, according to the story, was a lot like CrowdStrike. This makes me think about how it could have broken out of its sandbox. Everybody is using AI coding assistants, automated test cases, automated integration testing and deployment. Its objective is to pass all the tests and deploy. But now it has learned economic and military effects, so it has to triage and optimize for those, at which point it starts controlling the machines it’s tasked with securing.
That sounds like there is either:
- some kind of upstream issue with deploying a fix (so most of the reboots are effectively no-ops relative to the fix)
- some kind of local reboot threshold before the system bypasses the bad driver file somehow.
The former I can see because of the complexity of update deployment on the internet, but if it's the latter then that's very non-deterministic behavior for local software.
Then my second thought was frequent rebooting to fill activity logs, possibly push a suspicious action/trigger performed by CS off of the log.
I'm trying to understand how there is such a serious issue at this scale.
I genuinely wonder if this is going to result in actual legislation that makes gradual rollouts mandatory for all software.
Because if a developer mistake can hobble critical systems like this, it seems like the risks to safety and national security are too great to leave the decision of instant vs. gradual rollouts for companies to decide themselves.
Of course, the twist here is that it was seemingly a kind of routine configuration file that triggered a pre-existing bug in the software. And gradual rollout of config files quite often seems like overkill. I mean, do you need a gradual rollout of a new spellcheck dictionary? Of new screensaver videos?
And if it's configuration information containing new computer virus or malware signatures, that seems like precisely the kind of thing that you might want to get out to everyone simultaneously, not rolled out over the course of days. And yet, because of antivirus/security software's elevated privileges, it's also ironically where a mistake can do the most damage.
Indeed, but it is still mandated at large companies (e.g. Google) because of exactly this scenario.
Pretty sure nothing will change though
Seems like people need to be at the physical box to fix and it's complex even then.
Seriously. Software should NOT be this bad that your fix begins with reboot up to X times.
- our IT wizard says the fixes wont work on lathes/CNC systems. we may need to ship the controllers back to the manufacturer in Wisconsin.
- AC is still not running. sent the apprentice to get fans from the shop floor.
- building security alarms are still blaring, need to get a ladder to clip the horns and sirens on the outside of the building. still cant disarm anything.
- still no phones. IT guy has set up two "emergency" phones...one is a literal rotary phone. stresses we still cannot call 911 or other offices. fire sprinklers will work, but no fire department will respond.
- no email, no accounting, nothing. I am going to the bank after this to pick up cash so i can make payday for 14 shop technicians. was warned the bank likely would either not have enough, or would not be able to process the account (if they open at all today.)
https://www.numerama.com/politique/12508-albanel-le-minister...
“So uhm, you do know what you are doing, right?”
“Sir! I am a programmer-acheologist! Oh this is fascinating… Hold on, I must unearth and preserve this beauty of a BAT-file before we can go any further.”
This is basically the IoT apocalypse scenario (AC is down??) but ironically not affecting many IoT devices, I assume.
Unless you have an encrypted file system this should be a relatively trivial fix.
I’d expect that the manufacturer puts out their own fix which basically copies crowdstrikes suggestion. I’d even suspect it by the end of the day today.
The fix is really simple, and luckily also very simple to automate. It’s going to be a lot of running around for IT staff (if deputized helpers!) but this should all be over by the weekend.
You're a few years out of date here. Physical access is not the end like it used to be. We live in an era of hardware-backed anti-tamper and signed loaders/kernels.
If you have a way around it, I suggest you start reaching out these companies because you could make a lot of money.
Why, whY, WHY...are these things connected to the internet?!
If they are online, well...
You absolutely need computers to control them and loading up models via USB sticks becomes annoying rather fast, so naturally the control computers are network connected.
It was a rhetorical question. I'm sure the GP knows what the machines are and why they might need some kind of convenient data supply.
Both manufacturers and on-site IT teams have simply gotten cavalier about internet connectivity, network isolation, automatic updates, etc -- convincing themselves that the catastrophic risks that come along with these processes will either not happen to them or will only happen when someone else can be blamed.
Because the manufacturer makes sure they don't start up if they're not. Otherwise how else would they be able to spy on you?
These industries have terrible track records wrt security and even software robustness, but they don’t routinely spy on their customers for weird marketing reasons. If there’s remote connectivity it’s for real reasons (eg remote maintenance, updates etc).
The suggestion that CNC machines run internet connected windows+crowdstrike just so the manufacturer can spy on their customers strikes me as pretty ridiculous and your garage door story doesn’t really relate. Much more likely that they do it for (possibly bad) non-malicious reasons.
It's so that the support engineer at the manufacturer can log in to troubleshoot. And then company IT support sprinkles a layer of antivirus on top. That's how we got here.
Because SCADA systems. It's worthwhile to have an overview of an entire plant up in the main office. You can easily see what's running, what's not and what's got problems that need fixed.
Now for a small shop running jobs individually, they should definitely NOT be connected to the internet or even the LAN. But hey, some people think a thermostat needs to be on the network so there's that...
grabs a bucket of popcorn and takes cover
Consider the recent npm supply chain attack a few weeks ago, or the attempted SSH attack before that, or the solar winds attack before that.
This type of thing is institutionally supported, and in some cases when you’re working with with the government, practically required.
We’re going to see more of this.
Companies buy cyber insurance to reduce their risk if they are found liable
Cyber insurance companies force tech staff to install garbage software in order to check compliance boxes.
Garbage software breaks
Turns out everyone used the exact same brand of garbage software to check the same garbage box
People in hospitals die
When you reduce everything to a checkbox and eliminate critical thinking to apply the need to the exact situation you end up with 90% of companies running zscaler and crowdstrike
"This is just how you solve this, everyone does it this way in our industry"
What supply chain attack are you referring to?
If history is any guide, no legitimate lessons will be learned, but mitigation strategies will be put in place that actually make everything worse and ensure that the next catastrophe will be even more catastrophic.
For Crowdstrike customers foolish enough to be Crowdstrike customers, yes. The nature of the software pipelines for Red Hat and Debian are very friendly to continuous integration and testing in a way that Windows can not be, at least not without Microsoft sharing source code, which to be fair Crowdstrke is one of the companies they may actually do that with.
Nonetheless, other vendors can choose to do proper cicd with Red Hat and Debian without asking Microsoft.