CrowdStrike's outage should not have happened
ebellani.github.io
ebellani.github.io
Critical software engineering is a race to the bottom across many domains. Healthcare, banking, flight systems, etc.
Edit: I'd add this goes double when working on safety-critical code, or anything touching protected health data, or payment/financial data. It's just too toxic and valuable to leave to a chance change.
I agree, although in reality it's not chiefly developers themselves who are responsible for quick, lazy approaches, is it? Developers are typically the parties most pained by technical debt. If the discipline of software development is to become software engineering in earnest, there will have to be some pressure all the way up the management chain— pressure strong enough to outweigh software's low cost of iteration. I imagine this is really rare outside of highly regulated industries and very specific applications, and even with a formalized software engineering discipline, many companies will prefer sloppy software development and many competitive markets will 'select for' such companies.
I'd agree with you, except... ooh, a new, shiny, untested language / framework / platform to rewrite the codebase in!
I think the temptation to rewrite also reflects how messy and unworkable we let codebasee get— sometimes that impulse is more about the pain of working with the existing codebase than anything else.
Occasionally, usually because initial requirements were sorely lacking or changed, you can simplify the system via rewrite.
More often, everyone ends up realizing they didn't actually understand that last 5% of edge cases.
And then you've either replaced the working system with a 95% complete solution (so common in modern software) or you produce a system equally ugly once you handle that last 5%.
Lol HN.
Outside of civil/structural disciplines, PE is not required for engineering.
Mechanical, Chemical, Electrical, Nuclear don't require it.
I've literally never met an Aerospace engineer with a PE, and they build planes 'n sht.
---
It's a pure resume padder, like Cisco or AWS certifications.
You're just proving my point in that it's a CTO that dismisses the argument in a rather childish way. They would be the one to be told 'no' by the now-professional software engineers when their license and livelihood is on the line while being pressured to do something that goes against their recommendation. Funny how that power dynamic changes when there's something real on the line and not just an inflated title, huh?
Perhaps if those aerospace and software engineers that attempted to blow the whistle at Boeing were successful and were empowered via their license to say enough and stop development on MCAS, there wouldn’t be 300+ dead people because of a software change rammed through by management. The licensure ain’t just window dressing. It has real, actual impact on real human lives. Don’t be so dismissive.
A certificate would not change the status quo at this point.
Software developers/engineers are - for the most part - seen as essentially blue collar workers. Replaceable gears that MBAs can just "scale up" or "down" to fit their currently desired velocity. Let's ignore the fact that this fundamentally isn't true, but it's what they believe.
The work they do is decided by MBAs, and the time they have to implement these changes is heavily influenced by other MBAs.
Adding a certificate to this mix will change literally nothing
> I bristle at my title containing "engineer" as I don't have a PE
because it is aggressively ignorant of facts.
I have no comment on the potential utility of a PE requirement, software or otherwise.
If there's anything that the software industry needs most is a code of ethics. Companies are built on software that exploits, tricks and deceives their users. They release borderline malware and get rich doing it, either by having complicit investors or fooling them with false valuations. They cover their asses with dishonest PR, and lobby governments to keep the party going. This happens in the largest tech giants and tiny startups alike. And don't get me started on the gaming industry and their predatory practices.
We often exculpate engineers as being cogs in the machine, but they're ultimately choosing to work in these places, and enable this behavior.
The world would be a much better place if software engineers were required to take and uphold the equivalent of the Hippocratic Oath. We don't expect less from health professionals. Why should we from IT ones when the world is run by software?
Maybe there's improvement to be had, but this is not a difference between disciplines.
If one exists, I would like to see your argument that SWEs are adhering to it, and that the software industry is behaving ethically.
Some programmers (evident in the replies) even think it should exist for their profession, a very worrying idea.
And then a lot of "real" engineering now require software anyway: self-driving cars (written in part by people who are hacking together webapps by pulling in thousands of NPM dependencies) comes to mind.
The future is honestly a bit scary looking.
On the bright side things are going to get "interesting". At some point in the past we had many "Uber but for ...". Soon we'll have "Clownstrike but for fridges", "Clownstrike but for cars", etc.
Should be fun.
That absolutely can work, and does for plenty of industries, but it also creates the potential for a false sense of security until planes start falling out of the sky.
My frustration, and disappointment, in the software industry has generally been the complete unwillingness at scale for us to take on the responsibility to ensure safety and reliability without regulations enforcing it. Plenty of this responsibility (blame?) falls on companies led by individuals who are solely focused on profit and self-interest, but we have to own some of the responsibility as we're the ones agreeing to write and ship bad code.
I prefer the idealistic view that each individual can make a change through choice, but the reality is that choice is a privilege that isn't evenly distributed across the population. For example some can afford to not shop at Walmart, others can not - paradoxical as it may be from a local economics perspective.
Regulation is the typical blunt instrument to move the incentives to the business leaders rather than the individual. Other commenters don't think regulation is the answer, but I think most agree doing nothing won't change the status quo soon enough.
While I personally agree with the sentiment of your comment in general, this piece really is part of the blind spot in my opinion.
The assumption here is that everyone has to get all of their for from a grocery store, and the only question is what quality of products you can afford. It doesn't have to be that way, and wasn't until very recently in human history.
We almost always have alternatives. They just often seem so extreme as to not be feasible. People can grow their own food though. And at least in the US, we could go without a huge portion of the crap we spend money on every year. We just choose not to. There's absolutely nothing wrong with that choice, but its important to realize it is a choice.
A good example of this is an urban single parent of multiple kids, time and space are likely very scarce and choices are limited.
Understanding why this happens would be an interesting research project. Part of it might be information asymmetry with customers (shiny new features are very visible at sale time and reliability is totally unknown, so customers tend to weight known features over unknown reliability), and part might be principle agent issues (the decision maker who bought the software will have collected their bonus and retired long before the data breach can be attributed to them), and part might be that the market simply hasn't caught up to the negative consequences of all this change and careless companies will be purged by the market in the future.
I'm not terribly fond of regulation as a solution either. It tends to overconstrain industries, prevent innovation, and leave a hole at the lower end of the market that eventually makes products unaffordable. But there should be some quality mechanism that incentivizes decision makers to do the right thing and invest in quality even when there's a cost in features.
Cute fantasy about pinning everything on management but people do remember the old adage that "shit rolls downhill" don't they? What that will result in is very onerous processes and certifications mandated by "those in charge" on the people at the bottom to generate ironclad proof of no wrongdoing, at least for themselves. Maybe that is ultimately what this industry needs but it is also going to result in a work environment which really sucks a lot.
When management actively makes decisions to prioritize profit over security, for example, they should be held personally liable when a security issue occurs. I'm not really sure what a reasonable argument for that not being the case would look like.
If such a setup did result in a shitty work environment, people would ultimately have the option to not work at certain companies or to work for themselves. We can't assume that people must work for big tech and limit ourselves to what works in that sandbox.
Leaders of an electric company shouldn't be held liable for a lightning strike that starts a fire, for example. But they should be held liable if they purposely decide not to spend the money it takes to maintain power lines and a tree branch that should have been trimmed falls and starts a fire.
There would be consequences of such a system that change what we have today, but I wouldn't expect that to mean we couldn't possibly have things like electric companies.
Management is just making decisions based on what the companies value and companies are just valuing what their shareholders value which is more money for the shareholders.
The best way to fix it would be to reform the stock market system so that companies aren't beholden to uninvolved third parties looking to make a quick buck. Only active employees should own stock in companies and sit on company boards.
This would also require reforming the retirement system so retirement money isn't just dumped into the stock market. It needs to instead go somewhere safe and just sit. Retirement funds being in the stock market creates a huge inflationary feedback loop by demanding constant increases in profits which cause companies to raise prices which causes retirement funds to need to be bigger which causes them to demand more profit increases.
Take CrowdStrike for example. If the company and its leadership wasn't so well shielded from financial and legal liability they likely wouldn't have had a process that allowed rolling out an untested update to the entire world at once. Instead, they have a CEO that did effectively the same thing at McAfee before allowing it at CrowdStrike and the company will likely get little more than a financial slap of the wrist.
Would it solve everything? Absolutely not, and other actions like changes to the stock market could help. But it surely would make a difference if leadership and companies knew they could actually be ruined if they are provably negligent or culpable in damages caused.
This is used as a configuration and data exchange format despite having no formal definition, resulting in different results based on the parser used, and a weak typing system that has caused many bugs in many applications that use it. This despite the fact that many better, more reliable configuration and data interchange systems existed even before YAML got popular.
The Crowdstrike incident is worth billions and people may have died. If you look to the engineer you won't be able to recover billions. Hospitals must absolutely be on the hook as they are the direct interface to their customers; hospitals in turn can sue Crowdstrike.
An event worth billions must have billions in liability in order to prevent perverse incentives. Otherwise hospitals will just say "well McKinsey said it was a good bet, so what gives?"
I bought Cloudflare stock after this because it's obvious to me their customers aren't really thinking.
For those who weren't around, Sophos antivirus was the leading "enterprise-strength" security software around. When I started at Google in 2009, all corp-issued Windows laptops had it. Sometime in 2011, we got an internal email to immediately install a critical system update and reboot our machines. When they rebooted, they didn't have Sophos anymore. Then came another e-mail that said that Windows laptops would only be renewed for critical business exceptions, and everybody else was supposed to switch over to Mac or Linux laptops instead. Note that the first Chromebooks started shipping 3 months later.
A couple days later the full report came out, and a partial version was released to the public. The tl;dr was that Tavis Ormandy (then Google's lead security researcher) had done some cursory pentesting on Sophos, and found that it had so many security holes that it was architecturally impossible to make secure. Rather than attempt to bandage the problem, the company decided it was better to give up on Windows entirely.
[1] https://www.forbes.com/sites/andygreenberg/2011/08/04/google...
That's how _bad_ things are.
That's the thing that amazed me.
How do you regularly YOLO patches worldwide to something that runs with enough permissions to crash a system?
I don't care if this was a configuration update vs a new sensor capability -- universal rollout should never have been allowed by CrowdStrike's release team.
And I believe it because for administrators there is no configuration to delay the rollout of these "content updates", you can only delay the sensor updates.
I'm still awestruck that any engineer would be willing to ship code in that setup, though I guess its also possible that they were being misled about how much testing was going on at QA.
Given it's a security product, you'd want everyone protected on Day 0 that you have a new release, no?
Except on Day -1, no one was protected.
So what's magical about the day CrowdStrike decides to ship an update?
I can imagine there are possibly scenarios where mass-release would make sense (aggressive vulnerability spreading rapidly), but that can't be every day, can it?
I would think all the end-to-end tests of the full system would have been instantly detected the problem and prevented it, because it would have failed all the end-to-end tests.
Did I miss something? Did they never test the complete system as deployed? Looks like it, but maybe I misunderstood something.
If you don't test each config update, by definition you didn't test it.
Rust has many nice properties, but it can panic too. You should still test the whole system before each release if it's critical.
While such techniques are available, would they be really applicable in a very dynamic environment such as with millions of PCs running various windows versions, needing continuous / real-time updates.
And yes, we of course know that QA and testing magically removes all possible failure modes/bugs.
I don't see much difference in complexity between the affected software and the several existing formally verified software. At the very least the parser/interpreter could very much be formally verified.
But my point is, have they tried? They don't seem to be even aware of such.
One strategy Google SRE uses is that the team ensuring reliability has a different reporting path than the product team - so there is always a check and balance when things like rollout policies get worked around by clever product teams.
It's a shame because I hear it's actually a pretty good product.
* Saint Nedelya (church in the blog)
* Satya Nadella (microsoft ceo)
The thing not mentioned in CrowdStrike’s report is anything about people— especially management. Bad management and understaffed teams will defeat any technical solution, any day.
Maybe people with inside knowledge of recent events were trying to make an exit so they had to smash the glass and hit the red button to stop air travel so they could snag them in time?
Making it a perfect update failure is clever enough, but the name of the product is the best part. Imagine a system that can stop breaches even after they occur ;)
Either way, you've got a kernel-mode exception that isn't being caught, and that's a BSOD.
Many of the ideas of Tanenbaum would seem applicable towards preventing such a thing.
This is what happens. Stop skimping on QA.
Anything that is doing something sensitive or critical that can crash the system should be written in rust.
If not, insurance companies would be mandated by law to run static analysis on such C++ code.
We're going to keep seeing these horror stories until C/C++ go away.
Can you provide any references for this?
[0] https://www.whitehouse.gov/oncd/briefing-room/2024/02/26/mem...
[1] https://www.nist.gov/itl/ssd/software-quality-group/safer-la...
Also all GC languages would pretty easily solve the problem including Java and C#.
Static analysis tools (broadly) are not powerful enough to ensure memory safety for nontrivial programs without a high developer burden in the form of annotations or restricted language subsets.
GC languages are generally not appropriate for kernel level systems programming like crowdstrike was doing. Java has some history in this space, but generally without niceties like garbage collection that typical developers are used to.
That's not the reason. The DoD really is focused on Rust [1, 2].
(On a side note, pleasantly surprised to see a Linux kernel android bug today)
You need a license to operate a backhoe, but not to write software.
While Rust is a fine language, there are others which may be better suited for systems which have a global impact such as CrowdStrike does.
Ada: https://en.wikipedia.org/wiki/Ada_(programming_language)
SPARK: https://en.wikipedia.org/wiki/SPARK_(programming_language)
I am an advocate for neither and only reference their applicability for this problem domain based on the languages' design intent.
While Rust's behavior of panicing is saner than C++'s behavior of undefined behavior for a typical index out of bounds bug (`slice[index]`), I would imagine it'd still result in BSODs and outages in a kernel context.
Admittedly, safe bounds checking (`slice.get(index) → Option<T>`) might be easier in Rust than in C++, but it's not a given that a RIIR would've avoided this bug.
There are SO MANY things that went wrong. They don't even seem to care about half of them.
Ok, it wasn't tested. It wasn't staged. It wasn't a slow roll-out over hours. It wasn't verified from the deployment servers.
How about why is it parsed in kernel space at all? I didn't see that listed.
Why didn't you have a mechanism to catch faulty updates at runtime and flag them so that you'd be able to survive the next boot? Mark somewhere on the FS "I'm loading #19583". If you boot up and find that, it means you failed, so skip it next boot. Succeed? Remove the mark. Tada. Resiliency.
Why are you even doing this stuff yourself? There are Microsoft APIs for a bunch of this stuff that would have helped reduce risk. Oh, because your code is deployed on XP and up and you don't want to have 2 or 3 versions? That went well didn't it. I bet if you were using those you wouldn't have crashed the kernel space.
Do you fuzz? We know you don't. You broke Windows. You broke Macs. You broke Linux. You keep doing this. You're not good at being resilient. You had to face-plant in front of the entire world before you decided to do something. I didn't see fuzzing mentioned.
Those mitigations you propose are not enough. They'll help, but you're still vulnerable. You're supposed to be a _security company_, I'd expect a better list of how to secure software from a security company.
By all means, use Rust or Swift or (insert other memory safe language) if you want. That's smart. But if their existing kernel module had no UB and buffer overruns, it would still be doing things a kernel module shouldn't. The language involved is just a very tiny part.
Not even, it only runs on 7 / 2008 Server and up.
Better yet, it's high time that PCs booted like (some) mobile phones do: there's two copies of the OS. The running copy is read-only. The other copy can only be replaced with a monolithic "image" that has is validated with a HS256-based digitally signed Merkle hash. Either the entire image is bit-for-bit valid, or it is rejected.
If booting to the new image fails twice, automatically roll back to the previous known-good image.
This isn't "hard", it isn't even that much work. Heck, Windows can already boot from WIM images and VHDX disk images! The code is already in there, it just needs to be turned on.
Also I believe Apple Silicon Macs do what you suggest, at least in part. I don’t know if they can fall back after an upgrade failure but everything is validated and signed and tamper proof.