Let's blame the dev who pressed "Deploy"
yieldcode.blog
yieldcode.blog
> Neither the offerings nor crowdstrike tools are for use in the operation of aircraft navigation, nuclear facilities, communication systems, weapons systems, direct or indirect life-support systems, air traffic control, or any application or installation where failure could result in death, severe physical injury, or property damage.
(Originally in all-caps, presumably to make it sound more legally binding)
As far as I know, those applications weren't affected either - airports were affected because the ticketing systems were offline. Companies not using Crowdstrike were able to fly just fine.
Edit: 911 was affected in many places, and that is definitely concerning as it is very much life critical.
But that's not what this "small-print" says, i.e., the text you quoted and the terms and conditions page linked to in your comment.
Generally, it is not possible to disclaim liability for death, physical injury or property damage. The Crowdstrike disclaimer does not attempt to do so.
Nor do the terms ask the software user to assume any risk of death, physical injury or property damage. (Except for a warning about using "Malware Samples".)
I don't know where the delusion comes from, but some of them to go see the folly of their mistakes in the last few days. They will ride this off in the excuse that 'it affected so many other people, see, its not my fault !'.
From my own limited experience, those air gapped systems are often no more well managed than anything else. Perhaps having one more hop between the update channel and the secure network is enough to catch crowdstrike, but don’t be surprised if it isn’t.
Why though? Is it just "because we do it on every other machine", scared to fail audit, or what? Obviously the regulatory environment is a problem but IT incompetence is also another.
In any reasonably rational system, there should be some sort of balance between the two. That is hard to reach, as the various stakeholders are fundamentally at odds with one another, but it is easily believable that the dev who presses 'deploy' rarely has much of either.
People were running Crowdstrike on AWS (which in most (all?) cases involves a hypervisor), and recovery was still painful for them.
" Disclaimer. Neither Microsoft, nor the device manufacturer or installer, gives any other express warranties, guarantees, or conditions. Microsoft and the devicemanufacturer and installerexclude all implied warranties and conditions, including those of merchantability, fitness for a particular purpose, and non-infringement..."
https://www.microsoft.com/en-us/Useterms/Retail/Windows/10/U...
I thought this blog might have some substance about proper postmortem investigations and how to evaluate and address the circumstances that led to a failure like this, but it has none of that. It’s just a very angry rant about CEOs and middle management. The premise is that engineers can’t bear any responsibility for their actions because they don’t get “respect”
This has to be the 10th time I’ve seen arguments that “blame” is the right action in this case, but with the key exception that we’re only allowed to blame people other than the engineers. The last article was a lengthy rant about how it’s actually QA’s fault and engineers shouldn’t be expected to ensure their own code is correct, therefore engineers are blameless.
This is empty calories for people who like ragebait, but nothing more.
Go back to middle management.
Nah, that's not what the blog post says. It says engineers aren't given sufficient responsibility, so that they can assume the respective blame, when something goes wrong.
If an engineer says it would take 1 month to build a feature so it's sufficiently reliable, and the manager says "nah that's ridiculous, you have 1 week", then when the cobbled together feature breaks in production the manager should take the blame, since they effectively took the responsibility away from the engineer, upon themselves.
Developers deep in the trenches tend to have a bad feeling for business requirements or constraints; coupled with a knack for perfectionism and premature optimisation, that really often results in ridiculous time frames that are just plain unrealistic and would ruin the organisation long term.
I don’t have any profound insights, though: The only sane mantra can be keeping things in balance. Too much management, you drive your devs insane; too much engineer control, and the architecture astronauts reinvent the wheel every other day.
In my experience, this is because the developers are removed from interacting with the business. How are they supposed to make good decisions if they don't talk to their customers and understand what they are aiming for?
And I don't mean have the business folks show a roadmap once a year.
I have a similar opinion of being micromanaged. The micromanager is like a chaos demon that keeps pointing me in random directions. I lose all internal vision / intuition and turn into an unhappy task robot.
This is just standard corporate accountability avoidance on the side of the engineer though. Most people don’t want to be accountable for any risk so they advise against it, or give impractical advice, so that somebody else has to make the decision and hold the accountability.
That’s not how it works in any industry, ever. A single person can’t launch nukes, blow up a reactor, collapse a bridge, or otherwise cause billions of dollars in damages by accidents pressing one button.
The systems are designed to prevent that.
The premise of the post was a response to the ridiculous claim that when something goes bad, we need to blame the engineer(s) who pressed the button.
I tried, through rant, demonstrate that there are other people to blame, starting from politicians who are incompetent in what they do, to CEOs who get compensated for taking the risk, to managers who cut corners, etc.
The culmination of the post is that if you want o blame someone, you might as well blame any of the involved parties. But instead, if we want to prevent such issues in the future, we need to understand that the entire process or broken, rather than throwing individuals under the bus.
I hope this clarifies it a bit
CEO will of course blame some low level employees who did not follow procedures...
"CrowdStrike Significantly Invests in India Operations to Continue Protecting Businesses from Modern Cyber Attacks" - https://www.crowdstrike.com/press-releases/crowdstrike-inves...
And from now on every engineer now has the autonomy to refuse to use certain libraries and software stacks they are unfamiliar with, can refuse to submit a change, will have control over whether to push out software they worked on, etc.
And a huge pay bonus as well since they have all this CEO-like risk/responsibility now.
What the author of the article doesn't know is that the structural engineer can also get fired if management deems the engineer too troublesome. The only difference is that the PE has a code of ethics and professional responsibility to uphold, and failing to do so would mean risking their license to practice engineering.
This in a way covers the issue. Never in all my decades developing have had an estimate accepted. All the developers get is "It must be done by ...., no exceptions".
So you end up pulling all nighters for weeks, and because you are tired, errors or bad decisions creep in. So yes, the issue is fully with upper management.
I left a company because a big project was about to start and I could see it would be a big cluster**. People who stayed told me that is what happened.
Major logic error. See how much chaos ensued when that display went down? That's why it needs protection to keep the display up. Dumb or not is irrelevent. It is mission-critical regardless.
I know why, because there is, probably, a regulation that says that if you run an airline company, you need to have malware protection on all machines. I bet, some IT guy even tried to question the need to run EDR on a non-mission-critical machine, but he was stopped by a wall of "it is what it is".
Instead of assuming a regulation and writing a blog about it, do the research and find out. To quote the irreplaceable Benny Hill, "You mustn't assume, because it will make an ass out of you and me."
Also, and more important, why default to regulation and not airline directors pushing ill-advised modernization strategies pushed by M$?
We kept our other controls, we just added edr as well, because just having it appeased auditors. If you try to explain to an auditor your other controls, it could change a part of the audit from five minutes to multiple days.
We don’t use crowdstrike, but this was years ago.
So it's not necessarily government regulation per se, but a combination of things:
1. It's much safer (in terms of personal liability) for the decision makers at large companies to follow "standard industry practices" (however ridiculous they are). For example, no-one will get fired outside of Crowd Strike for this incident precisely because everyone was affected. "How could we have foreseen this when noone else did?"
2. The Cyber Security Insurance provider may not cover this kind of incident given there was no breach and so as far as they are concerned installing something like Crowd Strike is always profitable.
3. The insurance provider has no way to effectively evaluate the security posture of the enterprise they are insuring, so rely on basic indicators such as this checkbox, which completely eliminates any nuance and leads to worse outcomes (but not to the insurance provider!)
4. "Bad checkboxes" propagate down the supply chain the same way that "good checkboxes" do (eg. there are generally sections on these due diligence questionnaires about modern slavery regulation, and that's something you really want to propagate down the supply chain!)
Overall I would say the main cause of this issue is simply "big organisation problems". At a certain scale it seems to become impossible for everyone within the organization to commicate effectively and to make correct, nuanced decisions. This leads to the people at the top seeing these huge (and potentially real) risks to the business because of their lack of information. The person ultimately in charge of security can't scale to understand every piece of software, and so ends up having to make organisation-wide decisions with next to no information. The entire thing is a house of cards that noone can let fall down because it's simply too big to fail.
Making these large organisations work effectively is a very hard problem, but I think any solution must involve mechanisms to allow parts of the business to fail withing taking everything down. Allowing more decisions to be taken locally, but also the responsibilities and repercussions of those decisions to be felt locally.
Also I doubt any slimmed version of Windows is sufficiently malware proof without added EPS.
In the case of a display at a check-in counter:
- The display needs to be on a network, because it needs to collect information from elsewhere to display it.
- It's on a network, so it needs to be kept updated, because a compromised host elsewhere on the same network will be able to compromise it, and anyway the display vendor won't support you if your product is nine versions behind current.
- Since it needs updates for various components, it almost certainly needs some amount of outbound internet access, and it's also vulnerable to supply-chain attacks from those updates.
- Since it is on a network, and has internet access, it needs to be running some kind of EDR or feed for a SIEM, because it is compromisable and the last thing you want is an unmonitored compromised host on your internal network talking back to C2.
Anything that can be used for lateral movement will be used for lateral movement, and if we can get logs from it we want logs from it. A cross-platform EDR solution is perfect for these scenarios.
"- It's on a network, so it needs to be kept updated, because a compromised host elsewhere on the same network will be able to compromise it"
the suggested solution was "an isolated network".0
The problem there is the operator would have to use SD cards to update the adverts... :)
Because the thermostat on a fish tank has been used as a critical entry point into a casino network[1], and the point of EDR is not just to prevent that sort of thing if possible but also provide the telemetry into a SIEM for incident responders to know that it has happened after the fact and get the adversary out. So there is value in running it anywhere it can run.
I've seen a lot of contempt on HN threads today for compliance regulations and insurance demands that require things like EDR be installed where possible. As a Red Teamer I used to share that contempt for the non-technical types, but I don't now. It's true compliance is not security, but also true that Chesterton's Fence should apply here: just because you shouldn't be checking the box blindly doesn't mean you shouldn't be either checking it or documenting why not. The people who created the box were (probably) not actually idiots. It's there because somebody else had a very bad day.
1. https://www.washingtonpost.com/news/innovations/wp/2017/07/2...
I am pretty sure that removing EDR from an internetted Windows airport display would ensure it goes down. And probably then come up with a ransomware demand.
… arguing for 100% test coverage or SOLID principles are more like philosophies and anecdotes without a lot of hard data supporting them.
Software engineering looks less like engineering and more like a lot of bike shedding conversations.
It’s just nobody wants to pay for these specialist skills and would rather slap something on top of windows and call it a day
Now try to pushback on your managers request to “cut this long deploy process just once because this big client wants it fast”.
Blaming the developers (or any specific individual/group for that matter) is a cop out, it’s easy and lazy and doesn’t get to the root of the problem, which is more often than not a lack of processes, tools, information and lack of time/desire from leadership to address “technical debt” (for lack of a better term), no matter how many times the devs bring that up.
When you blame an individual or a group you can close the case shut on the post mortem and not get to any substantive improvements, meaning this can and will happen again, just to somebody else.
That’s why blameless postmortem and a blameless culture is so important. This is a good article about that philosophy:
> My summary of blameless culture is: when there is an outage, incident, or escaped bug in your service, assume the individuals involved had the best of intentions, and either they did not have the correct information to make a better decision, or the tools allowed them to make a mistake.
There is incompetence and complacency showing at every level after the CrowdStrike outage. Both the people selling CrowdStrike and the ones implementing it.
> I remember times when leaders had dignity and self-respect. They would go on stage and apologize
Etiquette rules change over time, but a constant throughout history is that people in power don't take responsibility voluntarily. The "good old days" where leaders had dignity and self-respect never existed.
> [...] delusional claim how software engineers should bear the responsibility for bugs and outages
The /opinion/ that people /should/ be held responsible can't be delusional. Apparently the author believes that software engineers should bear zero responsibility -- even when their software kills people -- because we don't get enough respect. I don't agree with that opinion but it's a bit rich to call other people delusional when making one unfounded claim after another.
The post as a whole is way too angry and too cynical for me.
I mean, it is a bit of a cluster fuck.
Most of the issues I've seen are due to speed or cost. We don't have time for tests.we don't have time to record the business requirements in a collective place (just look through multiple JIRA stories and piece it together). We don't have money for dedicated QA roles. So of course problems happen. Luckily my team only works on a lower criticality site.
None of this is really a dev's fault. It's leadership and the culture they incentivise.
So which mangers bear the blame here? The ones who likely met their SLAs, or the ones who knew their systems relied on software that wasn't 100% reliable and didn't have a backup plan?
But the same argument about lack of testing can be made for the companies using the system. Do they have tests for what happens when pieces of their systems or infrastructure go down? Did they ever do a disaster recovery drill? Honestly, there's blame all around because almost nobody is doing it right. Even the ones who are have failures.
Yes, but this forced them to run the worst DR routine short of Microsoft going rogue. The scale of the testing problem is orders of magnitude larger on one side: people trusted them to be minimally competent and they just weren’t.
I mention that last because that’s what happened to a lot of people here. They had DR plans assuming that they had their management infrastructure or could quickly bring it back online, but then they had things like CrowdStrike taking out the servers holding BitLocker recovery keys and other critical infrastructure. One of the under-appreciated outcomes from the general push to secure things has been that a lot of systems are now less robust because they depend on a few security critical components with no easy path to recovery if those fail. Full disk encryption is great from the perspective of data loss but it also means key management is mission critical in a way senior management probably wasn’t fully appreciative of when setting funding plans.
The article says:
> You want software engineers to be accountable for their code, then give them the respect they deserve.
The problem is, respect is something that's taken as much as its something given. We can't even decide for ourselves if software is worthy of respect. Can anyone learn to code from a coding bootcamp, where getting a job is all that matters? Or is it a discipline that takes years to master, where mistakes can and will bring down the global economy? Are we glorified plumbers, or are we mathematicians and civil engineers combined?
If you see yourself as a code monkey, of course you can't be "held responsible" for the results of your work. Coding bootcamps don't teach infosec. Its up to your company to set good practices and your job is just to follow them.
Its only if you personally want to take your role in society seriously that it makes sense to consider not just your job, but the effect your job has on the wider world. I'm personally of the opinion that this mindset is almost always long term positive for your career. Its less "blame the dev for hitting deploy" and more "I'm the dev. No, I won't hit deploy on that code in the state its in."
I didn't go to a coding bootcamp. I went to university. There they forced all of us engineering & CS students to do an ethics course - which was actually fantastic. There they taught us about the Therac-25: a computer controlled radiation machine which killed a bunch of people. The engineers on the ground knew that it needed more testing, but the company insisted it was fine and pushed it out the door.
Here's the question: If you were one of those engineers, what would you do? If you knew, or suspected, that a bug in your code could bring down 911 services and hospitals, ground planes or give people a lethal dose of radiation, do you really trust your manager to make the engineering call? Do you think the CEO understands the risk that is being taken by hitting deploy?
Of course, if we're playing the blame game, the blame ultimately the blame falls on the CEO of the company or something.
But forget the blame game. You won't be fired. You won't lose your cushy job. The question is: Who do you want to be in situations like this? Think about this question now. You won't have time in the moment when your boss tells you to hit deploy, and you have second thoughts.
Me? I want to be someone who would say no.
But in order to be the "person who would say no", the industry needs to understand that your opinion and expertise--matter. It could be a cultural shift, or a gate-keeping style shift where we protect the title "Engineer", like they do in some professions in some countries.
But given the current state, you can't on one hand blame the developers, and on the other hand treat them like spoiled kids who make too much money and in any way AI can replace most of them. It doesn't work this way. A Structural Engineer bears the responsibility because he has the authority, and respect to his knowledge, to refuse to sign off a broken design. This is not the case in software engineering.
To what extent is this because we act like spoiled kids? I really do mean "we" here; I probably have acted like that sometimes. I wonder if we, the post-microcomputer generations, are messed up to some extent because we started programming as a fun distraction from the work we were supposed to be doing, rather than learning programming as a serious job from the beginning like our predecessors who learned on mainframes or minicomputers in college.
But we don’t need to boil the ocean for things to improve. You personally can still decide you don’t want to make software that harms society. You personally can push back against your company if they want you to sell your users data. Nobody really knows how much they should respect your opinions and skills, so they’ll try things on. If you don’t respect yourself, nobody else will either.
Therefore the current identity crisis: the avg SWE is expected to be skilled and autonomous, but also report to daily standups and use mythical "story points" to estimate Very Serious Projects.
I would say devs are neither plumbers nor engineers but managers (of automation), but to recognize this is not Agile so instead we'll treat them as cheap cogs and wonder why the best escape to run their own shop where they can recapture the responsibilities and compensation eroded by bloat and grifters.
This is why the corp narrative is so wishful toward AI replacing engineers: an abstract magic box that prints money with no concept of accountability, copyright, liability is a dream for CEOs.
This is so naive that it is impossible to take the author seriously.
I haven’t heard of a surgeon who said “this operation will take as much time as needed”, and the hospital manager pressured him to “finish it in 8 hours, and not use too many syringes”.
If we simply accepted the fact that the wizards of the arcane machines knew what they do and the rest of us is at their mercy, then we would also loose any organisational control. I’m sure that sounds tempting to some engineers, but business people actually do serve a purpose. Acting like the all-knowable developers will surely do right just isn’t a good business strategy.