CrowdStrike releases root cause analysis of the global Microsoft breakdown
abc.net.au
abc.net.au
The way I read it was that they coded up this template type that could take 21 parameters. They expected it to take up to 21 parameters, but an intermediate representation expected it to have only 20. The code path which read the 21st parameter was not covered by tests nor the first deployed rapid response content using this template type. When they ran the untested code for the first time, kaboom.
First and foremost: bust mono-cultures in critical IT environments. Meaningful vendor diversity should be mandatory-by-law, and just like in the early days of the Internet, when nobody in their right mind would move into a data center that didn't have at least separate 'Cisco' and 'Juniper' uplink connectivity (so a vendor-specific bug wouldn't wipe you off the net: "ah, that unexpected BGP behavior took down all our peers for several hours" was, like, a weekly occurrence then), we shouldn't accept check-in-counter-workstations running exclusively on Windows today. It's already inexcusable that 'boot into the last OS snapshot that still worked' (which is commercial-off-the-shelf stuff used in most schools, just because 'little Johnny broke the PC again' is just a fact of life there) is apparently beyond oh-so-important mega-corps like Delta, but we need to go further.
So: mandatory hardware switches that select a boot into either Windows (running off one brand of SSD), or Linux or another completely-independent alternative (running on another brand of storage device, so that 'well, today is the day that all Seagate drives refuse to respond, due to an uptime integer overflow' is recoverable). One environment runs Chrome/Edge, Office365-connected-using-that-lovely-propietary-Exchange-protocol, the Microsoft TN3270 emulator and what-have-you. The other Firefox, OpenOffice with Thunderbird-over-IMAPS, Mocha TN3270 -- you get the idea. Run one environment for 3-to-4 days, then a mandatory switch-over, then back again, all year round. The training and making-everything-look-and-work-the-same development efforts will be fun, but nothing insurmountable.
Next up: Clownstrike and cohorts. Much has been made of the fact that custom kernel-mode drivers were used to track system activity, leading to this particular incident, but I can assure you, as a repeat victim of Enterprise Security Solutions, that even in plain old user mode, it's entirely feasible to mess things up beyond recovery using just the programming equivalent of a butter knife. And getting rid of layers like these will just never happen: "simply make the underlying OS secure enough" is not compatible with an open market for general computing, nor with the ambitions of Security Bureaucrats, and there will always be a need for 'zero day mitigation' which, when done reasonably well, can actually have dramatically good outcomes. I see this on a regular basis when figthing email phishing and spam: if you happen to catch a novel attack strategy and block it right away (which often requires an entirely new class of checks), you prevent a significant number of incidents. Snooze for an hour or two, and the (often painful) damage is done. So, "external vendors pushing cutting-edge code to all your devices on a real-time basis" is here to stay, and while certain players of course need to step up their basic validation game, regulation can only do so much.
So, what can be done: force Microsoft to ship https://learn.microsoft.com/en-us/sysinternals/downloads/sys... as a default and supported part of the OS, open-source it, have an industry-leading bug-bounty program, plus provide a rock-solid user-mode API that not only satisfies MS requirements, but also those of accredited third-party vendors. On Linux, where most interesting full-system observability tooling is still proprietary as well, find someone to open-source theirs (hey, Elastic has https://github.com/elastic/ebpf and some related stuff like https://github.com/open-telemetry/opentelemetry-ebpf-profile..., so perhaps that's a good basis?), with pretty much with the same requirements: rock-solid, open to reasonable criticism/improvements and a generally safe interface to whatever observability is deemed to be required.
TL;DR: the time has come for legally-mandated dual-ecosystem critical-use workstations with vendor-agnostic and safe observability APIs. First to deliver wins!
Some companies will divest of crowd strike and smaller players in that market will grow. Some companies will digest of windows, improving platform diversity. Hopefully release engineering at CS improves, strengthening their product, and for those that won’t do anything, insurance rates will rise making them think harder about this problem in the future.
Based on comments here, regulation requiring a checkbox security suite is what caused this to begin with. The answer isn’t more regulation.
The days-long airport chaos that most of the world just went through recently would suggest otherwise?
This was caused by over regulation in the first place. Adding more regulation only entrenches existing players making the market worse when things like this happen, not better. Those that chose monoculture should be punished in the market.
I won't, but you be you...
There are already laws to make people while in these scenarios. Can they replace missed days playing at the beach with children? No, of course not. Will I eve choose to fly an airline that bumped four kids and two below average weight adults because the plane weighed to much. They certainly won’t be my first choice.
I also had a pending flight home on an airline that was hard down for part of the Friday of doom. My flight wasn’t that day, but were it Delta I probably would have been affected. They can’t fulfill their contract, so minimally I would need a refund, and optimally hotels and rebooked flights on another airline. The underlying reason isn’t important and I don’t need more regulation for an already overly regulated industry.
Airline flights, while reliable, are not guaranteed. It’s massively inconvenient when they fail. Plan accordingly.[0]
0 - https://www.transportation.gov/airconsumer/fly-rights#Delaye...
...This is not a natural consequence of monoculture. You trade resilience/fault-tolerance for exposure to high-breadth, high-severity existential breakdowns.
The only people who like monocultures are those tasked with optimizing costs. There, monoculture looks great. In terms of actually building things that fail gracefully though, fuggedaboutit.