Initial details about why CrowdStrike's CSAgent.sys crashed
twitter.com
twitter.com
In the past ten years or so of having done somewhat serious computing and zero cybersecurity whatsoever, I have my mind concluded, feel free to disagree.
Approximately 100% of CVEs, crashes, bugs, slowdowns, and pain points of computing have to do with various forms of deserialising binary data back into machine-readable data structures. All because a) human programmers forget to account for edge cases, and b) imperative programming languages allow us to do so.
This includes everything from: decompression algorithms; font outline readers; image, video, and audio parsers; video game data parsers; XML and HTML parsers; the various certificate/signature/key parsers in OpenSSL (and derivatives); and now, this CrowdStrike content parser in its EDR program.
That wager stands, by the way, and I'm happy to up the ante by £50 to account for my second theory.
I'd make that 98%. Outside of rounding errors in the margins, the remaining two percent is made up of logic bugs, configuration errors, bad defaults, and outright insecure design choices.
Disclosure: infosec for more than three decades.
OWASP top-10 was dominated by those for a very long time. They have only recently been overtaken by authorization failures.
> Note "channel updates ...bypassed client's staging controls and was rolled out to everyone regardless"
> A few IT folks who had set the CS policy to ignore latest version confirmed this was, ya, bypassed, as this was "content" update (vs. a version update)
If your content updates can break clients, they should not be able to bypass staging controls or policies.
This is going to be what most customers did not realize. I'm sure Crowdstrike assured them that content updates were completely safe "it's not a change to the software" etc.
Well they know differently now.
Also go read Parse Don’t Validate https://lexi-lambda.github.io/blog/2019/11/05/parse-don-t-va...
Looking at how this whole thing is pasted together, there's probably a regex engine in one of those sys files somewhere that was doing the "parsing"...
I agree to the second but disagree on the first. Parser generator frameworks produce a lot of code that is hard to read and understand and they don't necessarily do a better job of error handling than you would. A hand-written recursive descent parser will usually be more legible, will clearly line up with the grammar that you're supposed to be parsing, and will be easier to add better error handling to.
Once you're aware of the risks of a bad parser you're halfway there. Write a parser with proper parsing theory in mind and in a language that forces you to handle all cases. Then fuzz the program, turn bad inputs that turn up into permanent regression tests, and write your own tests with your knowledge of the inner workings of your parser in mind.
This isn't like rolling your own crypto because the alternative isn't a battle-tested open source library, it's a framework that generates a brand new library that only you will use and maintain. If you're going to end up with a bespoke library anyway, you ought to understand it well.
1. The battle testing of the generator (e.g. lex yacc is widely used)
2. The readability of the grammar file, which enables good review
You don’t (often) read the generated code, only test it. Much like you wouldn’t read the binary generated by a compiler.
1. Poorly written code in the kernel module crashed the whole OS, and kept trying to parse the corrupted files, causing a boot loop. Instead of handling the error gracefully and deleting/marking the files as corrupt.
2. Either the corrupted files slipped through internal testing, or there is no internal testing.
3. Individual settings for when to apply such updates were apparently ignored. It's unclear whether this was a glitch or standard practice. Either way I consider it a bug(it's just a matter of whether it's a software bug or a bug in their procedures).
4. This was pushed out everywhere simultaneously instead of staggered to limit any potential damage.
5. Whatever caused the corruption in the first place, which is anyone's guess.
It is not possible that only this pattern caused the crash, and fuzzing omitted to try this unfuzzy pattern?
A fun example is that if you point AFL at a JPEG parser, it will eventually "learn" to produce valid JPEG files as test cases, without ever having been told what JPEG file is supposed to look like. https://lcamtuf.blogspot.com/2014/11/pulling-jpegs-out-of-th...
I think most AV companies now have a helper process to do that.
If you successfully exploit the helper process, the worst damage you ought to be able to do is falsely find files to be clean.
Ought. But it depends on the way the communication with the main process is done. I wouldn't be surprised if the main process trusts the output from the parser just a tiny bit too much.
I'll bet you this company has some arbitrary unit test coverage requirements for PRs which developers game be mocking the heck out of dependencies. I am sure they have some vanity sonarqube integration to ensure great "code quality". This likely also went through manual QA.
However I am sure the topic of fuzz testing would not have come up once. These companies sell checkbox compliance, and they themselves develop their software the same way. Checking all the "quality engineering" boxes with very little regards for long term engineering initiatives that would provide real value.
And I am not trying to kick Crowdstrike when they are down. It's the state of any software company run by suits with myopic vision. Their engineering blogs and their codebases are poles apart.
6) some form of sandboxing/error handling/api changes to make it possible to write safer kernel modules (not sure if it already exists and was just not used). It seems like the design could be better if a bad kernel module can cause a boot loop in the OS…
One of the problems with this outage was that you couldn’t even boot into safe mode without having the bit locker recovery key.
I don’t know the exact reasoning why safe mode requires the BitLocker recovery key, but presumably not doing so would open up an attack vector defeating the BitLocker protection.
Not staggering the updates is what blew my mind.
They probably considered it low risk, had done similar things of times hundreds of times before, etc.
Wild that anyone would consider anything in the “critical path” low risk. I would bet that they just don’t do rolling releases normally since it never caused issues before.
Not everyone is working on the same timezone.
I see remote, Israel, Canada.
https://crowdstrike.wd5.myworkdayjobs.com/en-US/crowdstrikec...
This one specifically Spain and Romania
I know they bought companies all over the globe from Denmark to other locations.
All the other engineering locations seem even less likely.
Bugs in code I can almost always understand and forgive, even the ones that seem like they’d be obvious with hindsight. But this is just an egregious lack of the most basic rollout standards.
what incentivized the bad decisions that led to this catastrophic failure?
It makes sense in a way - given their fast growth strategy (from nowhere to top 3) and desire to “do things differently” - the iconoclast upstarts that redefine the industry.
Or to summarise - hubris.
The "how" here is AV definition or a way to identify the attack. In CS-speak: content.
Catching 0day quickly results in good reputation that your EDR works well.
If people turn off their AV definition auto-update, they are at-risk. Why use EDR if folks don't want to stop attack quickly?
This is one bsod on windows 10. I saw another kernel panic on specific Linux distro.
What else?
One thing that is funny is that quite a few of their competitors are taking this opportunity to shit on them via Twitter and by marketing themselves as better than CrowdStrike.
Twitter, with all its issues, apparently has a feature to prevent fake news and that feature will show crowd source sentiment to debunk fake news, in this case, Twitter users showed how many times Crowdstrike competitor BSOD windows
There's best practice and then there's customer
The solution to customer reluctance to being the guinea pig is not to force every customer to be a guinea pig.
The red hat one. But they also did it to debian with a different issue, and I think another distro as well.
I'm sorry but this is the customer's fault.
If I'm using your services you work for me and you don't get to bully me into doing whatever you think needs to be done.
People that chose this solution need to be penalized, but they won't.
Compliance also has to share some of the blame here, if best practices (local testing) aren’t allowed to be followed in the name of “security”.
Many don’t have a choice, a lot of compliance is doing x to satisfy a checkbox and you don’t have a lot of flexibility in that or you may not be able to things like process credit cards which is kinda unacceptable depending on your company. (Note: I didn’t say all)
CrowdStrike automatic update happens to satisfy some of those checkboxes.
I never thought such stories were real until I encountered them…
But is there any responsibility for the clients consuming the data to have verified these updates prior to taking them in production? I haven't worn the sysadmin hat in a while now, but back when I was responsible for the upkeep of many thousands of machines, we'd never have blindly consumed updates without at least a basic smoke test in a production-adjacent UAT type environment. Core OS updates, firmware updates, third party software, whatever -- all of it would get at least some cursory smoke testing before allowing it to hit production.
On the other hand, given EDR's real-world purpose and the speed at which novel attacks propagate, there's probably a compelling argument for always taking the latest definition/signature updates as soon as they're available, even in your production environments.
I'm certainly not saying that CrowdStrike did nothing wrong here, that's clearly not the case. But if conventional wisdom says that you should kick the tires on the latest batch of OS updates from Microsoft in a test environment, maybe that same rationale should apply to EDR agents?
My company had a lot of Azure vms impacted by this and I'm not sure who the admin was who should have tested it. Microsoft? I don't think we have anything to do with crowdstrike software on our vms. ( I think - I'm sure I'll find out this week.)
Edit: I just learned the Azure central region failure wasn't related to the larger event - and we weren't impacted by the crowd strike issue - I didn't know it was two different things. So my second part of the comment is irrelevant.
Exactly which team owns the testing is probably left up to each individual company to determine. But ultimately, if you have a team of admins supporting the production deployment of the machines that enable your business, then someone's responsible for ensuring the availability of those machines. Given how impactful this CrowdStrike incident was, maybe these kinds of third-party auto-update postures need to be reviewed and potentially brought back into the fold of admin-reviewed updates.
In the boolean sense, yes. United Airlines (for example) is ultimately responsible for their own production uptime, so any change they apply without validation is a risk vector.
In pragmatic terms, it's a bit fuzzier. Does CrowdStrike provide any practical way for customers to validate, canary-deploy, etc. changes before applying them to production? And not just changes with type=important, but all changes? From what I understand, the answer to that question is no, at least for the type=channel-update change that triggered this outage. In which case I think the blame ultimately falls almost entirely on CrowdStrike.
I used to work with regional parks and recreation departments, and they would not approve any updates that did not go through UAW environments that we had set up. All updates had to be deployed to their UAW, thoroughly tested, before going to their production environment.
I get this this is slightly different, but I'd imagine Airlines, Banks, and Hospitals would have far more strict UAW policies to avoid a single vendor from kneecapping operations.
I don't know how you could assert that this is impossible, hence channel files should be treated as code.
Honestly, it hadn't even occurred to me that software like this marketed at enterprise customers wouldn't have this kind of control already available. It seems like an obvious thing that any big organization would insist on that I just took it for granted that it existed.
Whoops.
They do, but this update bypassed all of those rules.
No one at my company apparently.
I would say on the client for buying into CrowdStrike.
And also the client for having no contingencies and just accepting a vendor pinky-swear as meaningful.
CrowdStrike failed at their responsibilities too, I just mean that so did everyone else.
When you cede your own responsibilities to someone else and don't have that backed up with contractually enforced liability to make you whole when they fuck up, and also don't provide your own contingency so it doesn't really matter what some vendor does, that's on you. That's 100% entirely on you and it doesn't matter if a million other people also did the same utterly thoughtless and lazy thing.
I understand this perspective but I think it misses the forest for the trees. You have to evaluate this kind of stuff in context. Purity tests really smack on tech message boards where nobody has any accountability to any kind of business requirements, but basically no real-world organization operates in that way, so it's all a bit irrelevant.
> When you cede your own responsibilities to someone else ...
This framing is a bit naive, I think. It isn't a boolean. Everything is about risk management, cost/benefit analysis.
Which is also why you see every single customer affected - what you are suggesting is simply not an available thing to do at present for them.
At least for now - I imagine that some kind of staggered/slowed/ringed option will have to be implemented in the future if they want to retain customers.
If any one of the five points above hadn’t happened, this event would have been avoided. However, if number 1 had been addressed - any of the others could have happened (or all at the same time) and it would have been fine.
I understand that we should assume that bugs will be present anywhere, which is why staggered deployments are also important. If there had been staggered deployments, the. The damage would have happened, but it would have been localized. I think security people would argue against a staged deployment though, as if it were discovered what the new definitions protected against, an exploit could be developed quickly to put those servers that aren’t in the “canary” group at risk. (At least in theory — I can’t see how staggering deployment over a 6-12 hour window would have been that risky).
There may not be out of the box fuzzers that test device drivers so you hoist all the parser code, build it into a stand-alone application, and fuzz that.
Likely this is a form of technical debt since I can understand not doing all of this day #1 when you have 5 customers but at some point as you scale up you need to change the way you look at risk.
Policy changes seem more reliable and would catch other, as of yet unknown classes of bugs.
What looks especially bad for Crowdstrike is how many things (relatively simple things) had to fail in order for this to slip through. It's like walking into Fort Knox, grabbing a gold bar, and walking out unimpeded. A complete systemic failure.
That goes equally if it was a Windows Update rolled out in one motion that broke the falcon agent/driver, or if it was Crowdstrike. There is almost no excuse for a global rollout without telemetry checks, whether it's security agent updates or os patches.
And even testing can't be trusted 100%, because writing code that does the right thing and code that tests things correctly are about equally hard, they just aren't always hard simultaneously.
Is there anything we can take from other professions/tradecraft/unions/legislation to ensure shops can’t skip the basic best practices we are aware of in the industry like staged rollouts? How do we set incentives to prevent this? Seriously the App Store was raking in $$ from us for years with no support for staged rollouts and no other options.
No, for exactly the reason we just saw, and the same reason why vaccines are tested before widespread rollout.
As an example, reaction time is paramount to counter many kinds of attacks - thats why blocklists are so popular, and AS blackholing is a viable option.
Agreed. It's crazy that the top tech companies enforce this in a biblical fashion, despite all sorts of pressure to ship and all that. Crowdstrike went YOLO at a global scale.
I'd assume that sort of thing would be covered in the EULA and contract -- but even if it weren't, it seems like allowing customers to define their own definition update strategy would give them a pretty compelling avenue to claim non-liability. If CrowdStrike can credibly claim "hey, we made the definitions available, you chose to wait for 2 weeks to apply them, that's on you", then it becomes much less of a concern.
This is the most interesting question to me because it doesn't seem like there is an obviously guessable answer. It seem very unlikely to me that a company like CrowdStrike pushes out updates of any kind without doing some sort of testing, but the widespread nature of the outage would also seem to suggest any sort of testing setup should have caught the issue. Unless it's somehow possible for CrowdStrike to test an update that was different than what was deployed, it's not obvious what went wrong here.
The only reason any of these things don't cause issues in practice is checksums and error correcting codes.
Errors can show up any time, and usually show up between the parts that checksums correct. On the wire, TCP protects you with a (weak) checksum. Off the wire, your computer and filesystem can still screw things up. Even CPU bugs can do this.
I've built a couple of kernel drivers over the years and what I know is that ".sys" files are to the kernel as ".dll" files are to user-space programs in that the ones with code in them run only after they are loaded and a desired function is run (assuming boilerplate initialization code is good).
I've never made a data-only .sys file, but I don't see why someone couldn't. In that case, I'd guess that no one ever checked it was correct, and the service/program that loads it didn't do any verification either -- why would it, the developers of said service/program would tend to trust their own data .sys file would be valid, never thinking they'd release a broken file or consider that files sometimes get corrupted -- another failure mode waiting to happen on some unfortunate soul's computer.
Most importantly it was never tested at all :D
Any SWE job I've worked over my entire career, nothing is deployed with new versions of dependencies without testing them against a staging environment first.
If the crowdstrike selloff continues, I'm betting this will be why.
(There's a chance I'll make trading decisions based on this rationale in the next 72 hours, though I'm not certain yet)
I've heard that said elsewhere, but I haven't found a source for it at all. Are you able to point to one for me?
Thing is, as far as I can see, deploying this database update to a Windows machine will result promptly and unconditionally in a BSOD. That implies that this update was tried on exactly zero machines before it was shipped.
The bug can't have "slipped through internal testing"; it would have failed immediately on any machine it was loaded on.
Both of these are undergraduate-level techniques. Heck, they are covered in most first-semester programming courses. Either of these failures is inexcusable in a professional product, much less one that is running with kernel-level privileges.
Bet: CrowdStrike has outsourced much of its development work.
You think everyone writing code in the US would give two shits about the quality of their output if they see the CEO pocketing another private jet while they can barley make big-city rent?
Hell, even well paid devs at top companies in the US can be careless and lazy if their company doesn't care about quality. Have you seen some of the vulnerabilities and bugs that make it into the Android source code and on Pixel devices? And guess what, that code was written by well paid developers in the US, hired at Google leetcode standards, yet would give far-east sweatshops a run for their money in terms of carelessness. It's what you get when you have a high barrier of entry but a low barrier of output quality where devs just care about "rest and vest".
That said, I have had some experience with classic offshoring. Cultural differences make a huge difference!
My experience with "typical" programmers from India, China, et al is that they do exactly what they are told. Their boss makes the design decisions down to the last detail, and the "programmers" are little more than typists. I specifically remember one sweatshop where the boss looped continually among the desks, giving each person very specific instructions of what they were to do next. The individual programmers implemented his instructions literally, with zero thought and zero knowledge of the big picture.
Even if the boss was good enough to actually keep the big picture of a dozen simultaneous activities in his head, his non-thinking minions certainly made mistakes. I have no idea how this all got integrated and tested, and I probably don't want to know.
Sure but there's no proof yet that was the case here. That's just masive speculations based on anecdotes on your side. There's plenty of offshore devs that can run rings around western devs.
It doesn’t mean ‘Murica better, just that the origin story of staff matter, especially if you don’t have good processes around things like rca.
Every stereotype exists for a reason.
Alternatively, there might be a general assumption that lower development costs equate to inferior quality, which is a flawed yet prevalent human bias.
Don’t we have those kind of failures in almost every professional product? I’ve been working in the industry for over a decade and in every single company we had those bugs. The only difference was that none of those companies were developing kernel modules or whatever. Simple saas. And no, none of the bugs were outsourced (the companies I worked for hired only locals and people in the range of +- 2h time zone)
This problem has a promising solution, WUFFS, "a memory-safe programming language (and a standard library written in that language) for Wrangling Untrusted File Formats Safely."
HN discussion: https://news.ycombinator.com/item?id=40378433
HN discussion of Wuffs implementation of PNG parser: https://news.ycombinator.com/item?id=26714831
Ignore this failure which was catastrophic, this was a bad design asking to be exploited by criminals.
Of course it doesn't have to be either/or; you can have fast + secure, but it costs a lot more to design, develop, maintain and validate. What you can't have is a "why don't they just" simple and obvious solution that makes it cheap without making it either less secure, less performant, or both.
Given all the other mishaps in this story, it is very well possible that the software is insecure (we know that), slow and also still very expensive. There's a limit to how high you can push the triangle, but there's not bottom to how bad it can get.
The problem wasn't storing raw memory offsets, it was not having some way to validate the data at runtime.
Yes you still need to validate the keys to avoid logic errors, but you can avoid the memory errors.
The addresses aren't being generated internal to the program, so there are no "handles". They are referencing external data by design.
That's like saying "you shouldn't use a hard-coded volatile pointer to reference a hardware device". No, you literally need to do that sometimes; especially in embedded software.
What's that, three pints in a pub inside the M25? :P
Completely agree with this sentiment though, we've known that handling of binary data in memory unsafe languages has been risky for yonks. At the very least, fuzzing should've been employed here to try and detect these sorts of issues. More fundamentally though, where was their QA? These "channel files" just went out of the door without any idea as to their validity? Was there no continuous integration check to just .. ensure they parsed with the same parser as was deployed to the endpoints? And why were the channel files not deployed gradually?
For the record, the top 25 common weaknesses for 2023 are listed at:
* https://cwe.mitre.org/top25/archive/2023/2023_top25_list.htm...
Deserialization of Untrusted Data (CWE-502) was number fifteen. Number one was Out-of-bounds Write (CWE-787), Use After Free (CWE-416) was number four.
CWEs that have been in every list since they started doing this (2019):
* https://cwe.mitre.org/top25/archive/2023/2023_stubborn_weakn...
Out-of-bounds Write
Improper Neutralization of Input During Web Page Generation (‘Cross-site Scripting’)
Improper Neutralization of Special Elements used in an SQL Command (‘SQL Injection’)
Use After Free
Improper Neutralization of Special Elements used in an OS Command ('OS Command Injection')
Improper Input Validation
Out-of-bounds Read
Improper Limitation of a Pathname to a Restricted Directory (‘Path Traversal’)
Cross-Site Request Forgery (CSRF)
NULL Pointer Dereference
Improper Authentication
Integer Overflow or Wraparound
Deserialization of Untrusted Data
Improper Restriction of Operations within Bounds of a Memory Buffer
Use of Hard-coded Credentials
Parsing is the reverse—taking an untrusted string (or binary string) that is meant to be code and converting it into a data structure.
Both are the result of taking untrusted data and assuming it'll look like what you expect, but both are not parsing issues.
Which is precisely why parsing should've been used here instead. The correct way to do this is to work at the level after parsing, not before it. "SELECT * FROM foo WHERE bar LIKE ${untrusted input}" is dumb. Parsing the query with a placeholder in it, replacing it as an abstract node in the parsed form with data, and then serializing to string if needed to be sent elsewhere, is the correct way to do it, and is immune to injection attacks.
(Probably one other factor is that SQL was designed in a peculiar way, for "readability to non-programmers", which tends to result with languages that don't map well to simple data structures. Still, there are tools that let you construct a tree, and will generate a valid SQL from that.)
HTML is a better example, because it's inherently tree-structured, and trees tend to be convenient to work with in code. There it's more obvious when you're crossing from dumb string to parsed representation, and then back.
The only situation where a parser makes sense over simple escaping routines is if you actually intended to accept a subset of the language that you're injecting into rather than plain text, in which case you'll need more than just a parser to ensure you don't have anything dangerous—you'd need to do a lot of error-prone analysis of the AST afterward as well.
You shouldn't do "escaping" and string concatenation. That's just parsing and unparsing while cutting corners, which is how you get injection bugs.
> The only situation where a parser makes sense over simple escaping routines is if you actually intended to accept a subset of the language that you're injecting into rather than plain text
And that's exactly what you're doing. With escaping, you're taking a serialized form of some data, and splice into it some other data, massaged in a way you hope will make it always parse to string when something parses this later. It's going to eventually bite you; not necessarily with XSS - web template breakage is another common occurrence.
Working in string space is tricky, dangerous, and dumb - parsing, working on the parsed representation, and unparsing at the end, is how you do it correctly and safely.
(Another way to put it: plaintext is a wire format; you don't work in it if the data is structured.)
Note that the API may look like you're doing text - see JSX - but it internally goes through a parsing stage, and makes it impossible for you to do stupid things that break or transform the program, like working in string space lets you.
If you instead escape the user-provided unstructured text by replacing the very well-known set of special characters that could create tags, you know your users cannot produce active code, only text nodes.
It's the principle of least power: if you don't need users to access anything other than unstructured text then why feed their input into a parser that produces a data structure that represents code? Make illegal states unrepresentable by just escaping the text nodes as they're saved!
Working on the data structures after parsing makes it impossible to accidentally break the structure itself. Like, maybe your string escaping is perfect, but if you do:
$content = $templatePrefix + $sanitizedString + $templateSuffix;
Then you're still vulnerable to trivial errors in your template breaking the structure and creating an exploitable vulnerability, despite the $sanitizedString being correct. If you instead work at parsed level and do: $result = $template.findNode("#foo").setText($unsanitizedInput)
Then there's just no way this can break (except bugs in the HTML parser and DOM API in general, which are much less likely to exist, and much easier to find and fix). $result = $template.findNode("#foo").setText($unsanitizedInput)
This is not parsing the user input, this is letting the native API escape the input for you, which is exactly what I'm advocating for. See my note above:> escape the HTML using the language's standard library or the web framework's tooling.
This is what parsing the user input would look like with the DOM API:
const newDiv = document.createElement('div');
newDiv.innerHTML = untrustedUserInput;
// Do some work to attempt to sanitize the new HTML elements
document.body.appendChild(newDiv);
To me this is definitively a Bad Idea™, and I thought this was what you were advocating for.What you actually proposed is just escaping the HTML, not parsing user input, with the only twist being that you prefer to inject user input into your templating system imperatively with something resembling the DOM API instead of declaratively with something resembling JSX. That's fine, but not relevant to the question of what method we use to sanitize the untrusted input that we're injecting. On that front it sounds like we're in agreement that parsing user input is a terrible idea.
Surely many of these originate from deserialization of untrusted data (e.g., trusting a supplied length). It’s probably documented but I’m passively curious how they disambiguate these cases.
> Surely many of these originate from deserialization of untrusted data (e.g., trusting a supplied length).
Then they would presumably be classified under "Deserialization of Untrusted Data", number fifteen.
“Deserialization of untrusted data” isn’t even a security bug like an out of bounds write is. Every meaningful program deserializes external input. It’s a common area where bugs occur, but it’s not a type of bug in and of itself. Every bug in that category “belongs” in a more proximate category.
I wouldn't blame imperative programming.
Eg Rust is imperative, and pretty good at telling you off when you forgot a case in your switch.
By contrast the variant of Scheme I used twenty years ago was functional, but didn't have checks for covering all cases. (And Haskell's ghc didn't have that checked turned on by default a few years ago. Not sure if they changed that.)
This. One year ago UK air traffic control collapsed due to inability to properly parse "faulty" flight plan: https://news.ycombinator.com/item?id=37461695
The whole concept of "we fuck with the system in kernel based on data downloaded from the internet" is just not very sound and safe.
next time you'd be adding /s to your posts
Which is precisely the rationale which led to Standard Operating Procedures and Best Practices (much like any other Sector of business has developed).
I submit to you, respectfully, that a corporation shall never rise to a $75 Billion Market Cap without a bullet-proof adherence to such, and thus, this "event" should be properly characterized and viewed as a very suspicious anomaly, at the least
https://news.ycombinator.com/item?id=41023539 fleshes out the proper context.
28c3: The Science of Insecurity (2011)
By now, if you write any parser that deals with any outside data and don't fuzz the heck out of it, you are willfully negligent. Fuzzers are pretty easy to use, automatic and would likely catch any such problem pretty soon. So, did they fuzz and got very very unlucky or do they just like to live dangerously?
The problem is the entire category of surveillance software. It should not exist. Companies that use it don't understand security, and don't trust their employees. They're not good places to work at.
Complaints voiced by others included false positives (flagging something as a threat when it wasn't, or alerting that a system wasn't in place when it was), being too intrusive and affecting their workflow, and privacy concerns (reading and reporting all files, web browsing history, etc.). There were others I'm not remembering, as I mostly tried to stay away from the discussion, but it was generally disliked by the (mostly technical) workforce. Everyone just accepted it as the company deemed it necessary to secure some enterprise customers.
Also, Kolide's whole spiel about "honest security"[1] reeks of PR mumbo jumbo whose only purpose is to distance themselves from other "bad" solutions in the same space, when in reality they're not much different. It's built by Facebook alumni, after all, and relies on FB software (osquery).
> being too intrusive and affecting their workflow
Kolide is a reporting tool, it doesn't for example remove files or put them in quarantine. You also cannot execute commands remotely like in Crowdstrike. As you mentioned, it's based on osquery which makes it possible to query machine information using SQL. Usually, Kolide is configured to send a Slack message or email if there is a finding, which I guess can be seen as intrusive but IMO not very.
> reading and reporting all files
It does not read and report all files as far as I know, but I think it's possible to make SQL queries to read specific files. But all files or file names aren't stored in Kolide or anything like that. And that live query feature is audited (ens users can see all queries run against their machines) and can be disabled by administrators.
> web browsing history
This is not directly possible as far as I know, but maybe via a file read query but it's not something built-in out of the box/default. And again, custom queries are transparent to users and can be disabled.
> Kolide's whole spiel about "honest security"[1] reeks of PR mumbo jumbo whose only purpose is to distance themselves from other "bad" solutions in the same space
While it's definitely a PR thing, they might still believe in it and practice what they preach. To me it sounds like a good thing to differentiate oneself from bad actors.
Kolide gives users full transparency of what data is collected via their Privacy Center, and they allow end users to make decisions about what to do about findings (if anything) rather than enforcing them.
> It's built by Facebook alumni, after all, and relies on FB software (osquery).
For example React and Semgrep is also built by Facebook/Facebook alumni, but I don't really see the relevance other than some ad-hominem.
Full disclosure: No association with Kolide, just a happy user.
That said, since Kolide/osquery is a very flexible product, the complaints might not have been directed at the product itself, but at how it was configured by the security department as well. There are definitely some growing pains until the company finds the right balance of features that everyone finds acceptable.
Re: intrusiveness, it doesn't matter that Kolide is a report-only tool. Although, it's also possible to install extensions[1,2] that give it a deeper control over the system.
The problem is that the policies it enforces can negatively affect people's workflow. For example, forcing screen locking after a short period of inactivity has dubious security benefits if I'm working from a trusted environment like my home, yet it's highly disruptive. (No, the solution is not to track my location, or give me a setting I have to manage...) Forcing automatic system updates is also disruptive, since I want to update and reboot at my own schedule. Things like this add up, and the combination of all of them is equivalent to working in a babyproofed environment where I'm constantly monitored and nagged about issues that don't take any nuance into account, and at the end of the day do not improve security in the slightest.
Re: web browsing history, I do remember one engineer looking into this and noticing that Kolide read their browser's profile files, and coming up with a way to read the contents of the history data in SQLite files. But I am very vague on the details, so I won't claim that this is something that Kolide enables by default. osquery developers are clearly against this kind of use case[3]. It is concerning that the product can, in theory, be exploited to do this. It's also technically possible to pull any file from endpoints[4], so even if this is not directly possible, it could easily be done outside of Kolide/osquery itself.
> Kolide gives users full transparency of what data is collected via their Privacy Center
Honestly, why should I trust what that says? Facebook and Google also have privacy policies, yet have been caught violating their users' privacy numerous times. Trust is earned, not assumed based on "trust me, bro" statements.
> For example React and Semgrep is also built by Facebook/Facebook alumni, but I don't really see the relevance other than some ad-hominem.
Facebook has historically abused their users' privacy, and even has a Wikipedia article about it.[5] In the context of an EDR system, ensuring trust from users and handling their data with the utmost care w.r.t. their privacy are two of the most paramount features. Actually, it's a bit silly that Kolide/osquery is so vocal in favor of preserving user privacy, when this goes against working with employer-owned devices where employee privacy is definitely not expected. In any case, the fact this product is made by people who worked at a company built by exploiting its users is very relevant considering the type of software it is. React and Semgrep have an entirely different purpose.
[1]: https://github.com/trailofbits/osquery-extensions
[2]: https://github.com/hippwn/osquery-exec
[3]: https://github.com/osquery/osquery/issues/7177
[4]: https://osquery.readthedocs.io/en/stable/deployment/file-car...
[5]: https://en.wikipedia.org/wiki/Privacy_concerns_with_Facebook
There is a better alternative too. Make it a fair game for coworkers to send an invitation to a beer from the forgetful worker's machine to the whole company / department. It works wonders.
I would imagine an open source version of crowdstrike would not have had such a bad outcome.
The only reason this kind of software is used is so that companies can tick a certification checkbox that gives the appearance of running a tight ship.
I realize it's the easy way out, and possibly the only practical solution for a large corporation, but then this type of issues is unavoidable. Whether the product is free or proprietary makes no difference.
You highlight training as a control. Training is expensive - to reduce cost and enhanced effectiveness, how do you focus training on those that need it without any method to identify those that do things in insecure ways?
Additionally, I would say a major function of these systems is not surveillance at all - it is preventive controls to prevent compromise of your systems.
Overall, your comment strikes me a naive and not based on operational experience.
This is to say, there are costs and threats caused by deploying these systems too, and they should be considered when making security decisions.
The years I spent doing IT at that level, every time, every single time I got a request for admin privileges to be granted to a user or for software to be installed on an endpoint we already had a solution in place for exactly what the user wanted, installed and tested on their workstation that was taught in onboarding and they simply "forgot".
Just like the users I had to reset their passwords for every monday because they forgot their passwords. It's an irritation but that doesn't mean they didn't do their job well. They met all performance expectations, they just needed to be handheld with technology .
The real world isn't black and white and this isn't Reddit.
For example by doing continuous scans that consume so much CPU the machine stays thermally throttled at all times.
(Yes, really. I've seen a colleague raising a ticket about AV making it near-impossible to do dev work, to which IT replied the company will reimburse them for a cooling pad for the laptop, and closed the issue as solved.)
The problem is so bad that Microsoft, despite Defender being by far the lightest and least bullshit AV solution, created "dev drive", a designated drive that's excluded by design from Defender scanning, as a blatant workaround for corporate policies preventing users and admins from setting custom Defender exclusions. Before that, your only alternative was to run WSL2 or a regular VM, which are opaque to AVs, but that tends to be restricted by corporate too, because "sekhurity".
And yes, people in these situations invent workarounds, such as VMs, unauthorized third-party SaaS, or using personal devices, because at the end of the day, the work still needs to be done. So all those security measures do is reduce actual security.
People making decisions about purchasing, deploying and configuring those systems are separated by many layers from rank-and-file employees. The impact on business downstream is diffuse and doesn't affect them directly, while the direct incentives they have are not aligned with the overall business operations. The top doesn't feel the damage this is doing, and the bottom has no way of communicating it in a way that will be heard.
It does build distrust, but not necessarily in the sense that "company thinks I'm a potential criminal" - rather, just the mundane expectation that work will continue to get more difficult to perform with every new announcement from the security team.
Also I'm unsure I've ever seen an AV even come close to stressing a machine I would spec for dev work. Likely misconfigured for the use case but I've been there and definitely understand the other side of the coin, sometimes a beer or pizza with someone high up at IT gets you much further than barking. We all live in a society with other people.
I would also hazard a guess that the defender drive is more a matter of just making it easier for IT to do the right thing, requested by IT departments more than likely. I personally have my entire dev tree excluded from AV purely because of false positives on binaries and just unnecessary scans because the fines change content so regularly. That can be annoying to do with group policy if where that data is stored isn't mandated and then you have engineers who would be babies about "I really want my data in %USERPROFILE%/documents instead oF %USERPROFILE%/source" now IT can much easier just say that the Microsoft blessed solution is X and you need to use it.
Regarding WSL, if it's needed for you job then go for it and have you manager out in a request. However if you are only doing it to circumvent IT restrictions, well don't expect anyone to play nice.
On the person devices note. If there's company data on your device it and all it's content can be subpoenad in a court case. You really want that? Keep work and personal seperate, it really is better for all parties involved.
That's true, but it gets tricky in a large multinational, when the rules are set by some team in a different country, whose responsibilities are to the corporate HQ, and the IT department of the merged-in company I worked for has zero authority on the issue. I tried, I've also sent tickets up the chain, they all got politely ignored.
From the POV of all the regular employees, it looks like this: there are some annoying restrictions here and there, and you learn how to navigate the CPU-eating AV scans; you adapt and learn how to do your work. Then one day, some sneaky group policy update kills one of your workarounds and you notice this by observing that compilation takes 5x as long as it used to, and git operations take 20x as long as they should. You find a way to deal (goodbye small commits). Then one day, you get an e-mail from corporate IT saying that they just partnered with ESET or CrowdStrike or ZScaler or not, and they'll be deploying the new software to everyone. Then they do, and everything goes to shit, and you need to start to triple every estimate from now on, as the new software noticeably slows down everything across the board. You think to yourself, at least corporate gave you top-of-the-line laptops with powerful CPUs and absurd amount of RAM; too bad for sales and managers who are likely using much weaker machines. And then you realize that sales and management were doing half their work in random third-party SaaS, and there is an ongoing process to reluctantly in-house some of the shadow IT that's been going on.
Fortunately for me, in my various corporate jobs, I've always managed to cope by using Ubuntu VMs or (later) WSL2, and that this always managed to stay "in the clear" with company security rules. Even if it meant I had to figure out some nasty hacks to operate Windows compilers from inside Linux, or to stop the newest and bestest corporate VPN from blackholing all network traffic to/from WSL2 (was worth it, at least my work wasn't disrupted by the Docker Desktop licensing fiasco...). I never had to use personal devices, and I learned long ago to keep firm separation between private and work hardware, but for many people, this is a fuzzy boundary.
There was one job where corporate installed a blatant keylogger on everyones' machines, and for a while, with our office IT's and our manager's blessing, our team managed to stave it off - and keep local admin rights - by conveniently forgetting to sign relevant consent forms. The bad taste this left was a major factor in me quitting that job few months later, though.
Anyway, the point to these stories is, I've experienced first-hand how security in medium and large enterprises impacts day-to-day work. I fought both alongside and against IT departments over these. I know that most of the time, from the corporate HQ's perspective, it's difficult to quantify the impact of various security practices on everyone's day-to-day work (and I briefly worked in cybersecurity, so I also know this isn't even obvious to people this should be considered!). I also know that large organizations can eat a lot of inefficiency without noticing it, because at that size, they have huge inertia. The corporate may not notice the work slowing down 2x across the board, when it's still completing million-dollar contracts on time (negotiated accordingly). It just really sucks to work in this environment; the inefficiency has a way of touching your soul.
EDIT:
The worst is the learned helplessness. One day, you get fed up with Git taking 2+ minutes to make a goddamn commit, and you whine a bit on the team channel. You hope someone will point out you're just stupid and holding it wrong, but no - you get couple people saying "yeah, that's how it is", and one saying "yeah, I tried to get IT to fix that; they told me a cooling stand for the laptop should speed things a bit". You eventually learn that security people just don't care, or can't care, and you can only try to survive it.
(And then you go through several mandatory cybersecurity trainings, and then you discover a dumb SQL injection bug in a new flagship project after 2 hours of playing with it, and start questioning your own sanity.)
And let's see if we can agree that likely corporate multinationals are probably a bad thing, or at least micromanaging from the stratosphere when you cannot see how youe decision effects things. That however is likely a management antipattern and if it is really negatively effecting your mental health but you are still meeting performance expectations I'm not against you making a decision to walk.
Sometimes the only way to solve those problems is to cause turnover and make management look twice, and a lot of time one key person leaving can cause an exodus that will force change.
Not being negative here, sometimes you are just in a toxic relationship and need to get out.
Imagine you are a bank. Imagine you have no way to ensure no employee is a crook.
It does happen.
Wait, are you saying we have gotten rid of all the crooks in a bank/or those that handle money?
What should these companies understand about security exactly?
And aren’t they kinda right to not trust their employees if they employ 50,000 people with different skills and intentions?
These endpoint security companies latch onto people making decisions, those people want security and these software vendors promise to make the process as easy as possible. No need to change the way a company operates, just buy our stuff and you're good. That's the scam.
Truthfully, it must be practically infeasible to transform security practices of a large company overnight. Most of the time they buy into these products because they're chasing a security certification (ISO 27001, SOC2, etc.), and by just deploying this to their entire fleet they get to sidestep the actually difficult part.
The irony is that at the end of this they're not anymore "secure" than they were before, but since they have the certification, their customers trust that they are. It's security theater 101.
Yes, in a 50k employee company, the CEO won't know every single employee and be able to vouch for their skills and intentions.
But in a non-dysfunctional company, you have a hierarchy of trust, where each management level knows and trusts the people above and below them. You also have siloed data, where people have access to the specific things they need to do their jobs. And you have disaster mitigation mechanisms for when things go wrong.
Having worked in companies of different sizes and with different trust cultures, I do think that problems start to arise when you add things like individual monitoring and control. You're basically telling people that you don't trust them, which makes them see their employer in an adversarial role, which actually makes them start to behave less trustworthy, which further diminishes trust across the company, harms collaboration, and eventually harms productivity and security.
A Marxist reading would suggest alienation, but a more modern one would realize that it is a bit more than that: to enable modern business practices (both good and bad!) we designed systems of management to remove or reduce trust and accountability in the org, yet maintain as similar results to a world that is more in line with the one you believe is possible.
A security professional though would tell you that even in such a world, you can not expect even the most diligent folks to be able to identify all risks (e.g. phishing became so good, even professionals can’t always discern the real from fake), or practice perfect opsec (which probably requires one to be a psychopath).
Even in a company of two sometimes a husband or a wife betrays the trust. Now multiply that probability by 50000.
A user doesn’t have to do anything wrong for the computer to become compromised, or even if they do, being able to limit the blast radius and lock down the computer or at least after the fact have collected the data to be able to identify what went wrong seems important.
How would you secure a network of computers without an agent that can do anti-virus, detect anomalies, and remediate them? That is to say, how would you manage to secure it without doing something that has monitoring and lockdown capabilities? In your words, signaling that you do not trust the users?
There is no straightforward answer to this question. Assuming that your infrastructure is "secure" because you deployed an EDR solution is wrong. It only gives you a false sense of security.
The reality is that security takes a lot of effort from everyone involved, and it starts by educating people. There is no quick bandaid solution to these problems, and, as with anything in IT, any approach has tradeoffs. In this case, and particularly after the recent events, it's evident that an EDR system is as much of a liability as it is an asset—perhaps even more so. You give away control of your systems to a 3rd party, and expect them to work flawlessly 100% of the time. The alarming thing is how much this particular vendor was trusted with critical parts of our civil infrastructure. It not only exposes us to operational failures due to negligence, but to attacks from actors who will seek to exploit that 3rd party.
Any security certification has a section on regularly educating employees on the topic.
To your point, I agree that companies are attempting to bypass the hard work by deploying a tool and thinking they are done.
Personally, I don't know how to solve that problem.
It is not considered a silver bullet by the security team, rather a last-resort detection mechanism for suspicious behavior (for example if the network segmentation or access control fails, or someone managed to get foothold by other means). It also helps them identify which employees need more training as they keep downloading random executables from the web.
But this will be the work of at least two human generations; our tools and work practices are woefully inadequate, so even if the pointy haired bosses (fearing imprisonment for gratuitous failure) and grasping, greedy investors fear (for the destruction of “hard earned” capital), it’s not going to be done in the snap of our fingers, not least because the people occupying technology industry - and this is an overgeneralisation, but I’m pretty angry so I’m going to let it stand - Just Don’t Care Enough.
If we cared, it would be nigh on impossible for my granny to get tricked to pop her Windows desktop by opening an attachment in her email client.
It wouldn’t be possible to sell (or buy!) cloud services for which we don’t get security data in real time and signal about what our vendor advises to do if worst comes to worst.
And on and on.
As the other reply to your comment said: the world is not 'fair' or 'honest', that's just a lie told to children. Apart from geuinely evil people, there are unlimited variables that dictate people's behavior. Culture, personality, nutrition, financial situation, mood, stress, bully coworkers, intrinsic values, etc etc. To think people are all fair and honest "unless" is a really harmful worldview to have and in my opinion the reason for a lot of bad things being allowed to happen and continue (troughout all society, not just work).
Zero-trust in IT is just the digitized version of "trust is earned". In computers you can be more crude and direct about it, but it should be the same for social connections and interactions.
You have to start with that premise otherwise organizations and society fail. Every hour of every day, even people in high security organizations have opportunities to betray the trust bestowed on them. Software and processes are about keeping honest people honest. The dishonest ones you cannot do too much about but hope you limit the damage they can cause.
If everyone is treated as dishonest then there will eventually be an organizational breakdown. Creativity, high productivity, etc... do not work in a low/zero trust environment.
It's like you let one company build your office building and then bring in another contractor to randomly add walls and have others removed while having never looked at the blueprints and then one day "whoopsie, that was a supporting wall I guess".
Why is it not just completely normal but even expected that an OS vendor can't build an OS properly, or that the admins can't properly configure it, but instead you need to install a bunch of crap that fucks around with OS internals in batshit crazy ways? I guess because it has a nice dashboard somewhere that says "you're protected". Checkbox software.
If you manage a fleet of tens of thousands of systems and you need to protect against well funded organized crime? Employees running malicious code under their user is a given and can't be prevented. Buying crowdstrike sensor doesn't seem like such a bad idea to me. What would you do instead?
As said, limit the user's abilities as much as possible with features of the OS and software in use. Maybe if you want those other metrics, use a firewall, but not a Tls-breaking virus scanning abomination that has all the same problems, but a simple one that can warn you on unusual traffic patterns. If soneone from accounting starts uploading a lot of data, connects to Google cloud when you don't use any of their products, that should be odd.
If we're talking about organized crime, I'm not convinced crowdstrike in particular doesn't actually enlarge the attack surface. So we had what now as the cause, a malformed binary ruleset that the parser, running with kernel privileges, choked on and crashed the system. Because of course the parsing needs to happen in kernel space and not a sandboxed process. That's enough for me to make assumptions about the quality of the rest of the software, and answer the question regarding attack surface.
Before this incident nobody ever really looked at this product at all from a security standpoint, maybe because it is (supposed to be) a security product and thus cannot have any flaws. But it seems now security researchers all over the planet start looking at this thing and are having a field day.
Bill gates sent that infamous email in the early 2000s, I think after sasser hit the world, that security should be made the no1 priority for Windows. As much as I dislike windows for various reasons, I think overall Microsoft does a rather good job about this. Maybe it's time those companies behind these security products start taking security serious too?
If you only knew how absurd of a statement that is. But in any case, there are just too many threats network IDS/IPS solutions won't help you with, any decent C2 will make it trivial to circumvent them. You can't limit the permissions of your employees to the point of being effective against such attacks while still being able to do their job.
You don't seem to know either since you don't elaborate on this. As said, people are picking this apart on Twitter and mastodon right now. Give it a week or two and I bet we'll see a couple CVEs from this.
For the rest of your post you seem to ignore the argument regarding attack surface, as well as the fact that there are companies not using this kind of software and apparently doing fine. But I guess we can just claim they are fully infiltrated and just don't know because they don't use crowdstrike. Are you working for crowdstrike by any chance?
But sure, at the end of the day you're just gonna weigh the damage this outage did to your bottom line and the frequency you expect this to happen with, against a potential hack - however you even come up with the numbers here, maybe crowdstrike salespeople will help you out - and maybe tell yourself it's still worth it.
The trouble is that people still need local file access, and use network file shares. You have hundreds of apps used by a handful of users that need to run locally. And a few intranet apps that are mission critical and have dubious security. That creates the necessity for wrapping users in firewalls, vpns, tls interception, end point security etc. And the less well it all works the more you need to fill the gaps.
Fun fact an attacker only needs to steal credentials from the home directory to jump into a companies AWS account where all the juicy customer data lives, so there are reasons we want this control.
Frankly I'd like to see the smart people complaining help write better solutions rather than hinder.
Doing it right requires very capable individuals and a significant effort. Less than it used to take, more than most companies are ready to invest.
The alternative is to replace you with AI yes?
Globally networked personal computers were kind of cultural revolution against the setting you describe. Everyone had their own private compute and compute time and everyone could share their own opinion. Computers became our personal extensions. This is what IBM, Atari, Commodore, Be, Microsoft and Apple (and later desktop Linux) sold. Now given this ideology, can a company own my limbs? If not, they can't own my computers.
Well, presuming that:
1. the employee is issued a computer, that they have possession of even if not ownership (i.e. they bring the computer home with them, etc.)
2. and the employee is required to perform creative/intellectual labor activities on this computer — implying that they do things like connecting their online accounts to this computer; installing software on this computer (whether themselves or by asking IT to do it); doing general web-browsing on this computer; etc.
3. and where the extent of their job duties, blurs the line between "work" and "not work" (most salaried intellectual-labor jobs are like this) such that the employee basically "lives in" this computer, even when not at work...
4. ...to the point that the employee could reasonably conclude that it'd be silly for them to maintain a separate "personal" computer — and so would potentially sell any such devices (if they owned any), leaving them dependent on this employer-issued computer for all their computing needs...
...then I would argue that, by the same chain of reasoning as in the GP post, employers should not be legally permitted to “issue” employees such devices.
Instead, the employer should either purchase such equipment for the employee, giving it to them permanently as a taxable benefit; or they should require that the employee purchase it themselves, and recompense them for doing so.
Cyberpunk analogy: imagine you are a brain in a vat. Should your employer be able to purchase an arbitrary android body for you; make you use it while at work; and stuff it full of monitoring and DRM? No, that'd be awful.
Same analogy, but with the veil stripped off: imagine you are paraplegic. Should your employer be allowed to issue you an arbitrary specific wheelchair, and require you to use it at work, and then monitor everything you do with it / limit what you can do with it because it’s “theirs”? No, that’d be ridiculous. And humanity already knows that — employers already can't do that, in any country with even a shred of awareness about accessibility devices. The employer — or very much more likely, the employer's insurance provider — just buys the person the chair. And then it's the employee's chair.
And yes, by exactly the same logic, this also means that issuing an employee a company car should be illegal — at least in cases where the employee lives in a non-walkable area, and doesn't already have another car (that they could afford to keep + maintain + insure); and/or where their commute is long enough that they'd do most non-employment-related car-requiring things around work and thus using their company car. Just buy them a car. (Or, if you're worried they might run away with it, then lease-to-own them a car — i.e. where their "equity in the car" is in the form of options that vest over time, right along-side any equity they have in the company itself.)
> Does that also apply to the operator at a switchboard…
Actually, no! Because an operator of a switchboard is not a “user” of the computer that powers the switchboard, in the same sense that a regular person sitting at a workstation is a "user" of the workstation.
The system in this case is a “kiosk computer”, and the operator is performing a prescribed domain-specific function through a limited UX they’re locked into by said system. The operator of a nuclear power plant is akin to a customer ordering food from a fast-food kiosk — just providing slightly more mission-critical inputs. (Or, for a maybe better analogy: they're akin to a transit security officer using one of those scanner kiosk-handhelds to check people's tickets.)
If the "computer" the nuclear-plant operator was operating, exposed a purely electromechanical UX rather than a digital one — switches and knobs and LEDs rather than screens and keyboards[1] — then nothing about the operator's workflow would change. Which means that the operator isn't truly computing with the computer; they're just interacting with an interface that happens to be a computer.
[1] ...which, in fact, "modern" nuclear plants are. The UX for a nuclear power plant control-center has not changed much since the 1960s; the sort of "just make it a touchscreen"-ification that has infected e.g. automotive has thankfully not made its way into these more mission-critical systems yet. (I believe it's all computers under the hood now, but those computers are GPIO-relayed up to panels with lots and lots of analogue controls. Or maybe those panels are USB HID devices these days; I dunno, I'm not a nuclear control-systems engineer.)
Anyway, in the general case, you can recognize these "the operator is just interacting with an interface, not computing on a computer" cases because:
• The machine has separate system administrators who log onto it frequently — less like a workstation, more like a server.
• The machine is never allowed to run anything other than the kiosk app (which might be some kind of custom launcher providing several kiosk apps, but where these are all business-domain specific apps, with none of them being general-purpose "use this device as a computer" apps.)
• The machine is set up to use domain login rather than local login, and keeps no local per-user state; or, more often, the machine is configured to auto-login to an "app user" account (in modern Windows, this would be a Mandatory User Profile) — and then the actual user authentication mechanism is built into the kiosk app itself.
• Hopefully, the machine is using an embedded version of the OS, which has had all general-purpose software stripped out of it to remove vulnerability surface.
> If employers allowed employees to "bring their own devices", and then didn't force said employees to run MDM software on those devices, then how in the world could the employer guarantee the integrity of any line-of-business software the employee must run on the device; impose controls to stop PII + customer-shared data + trade secrets from being leaked outside the domain; and so forth?
My answer to that question: it's safe to say that most people in the modern day are fine with the compromise that your device might be 100% yours most of the time; but, when necessary — when you decide it to be so — 99% yours, 1% someone else's.
For example, anti-cheat software in online games.
The anti-cheat logic in online games, is this little nugget of code that runs on a little sub-computer within your computer (Intel SGX or equivalent.) This sub-computer acts as a "black box" — it's something the root user of the PC can't introspect or tamper with. However:
• Whenever you're not playing a game, the anti-cheat software isn't loaded. So most of the time, your computer is entirely yours.
• You get to decide when to play an online game, and you are explicitly aware of doing so.
• When you are playing an online game, most of your computer — the CPU's "application cores", and 99% of the RAM — is still 100% under your control. The anti-cheat software isn't actually a rootkit (despite what some people say); it can't affect any app that doesn't explicitly hook into it.
• In a brute-force sense, you still "control" the little sub-computer as well — in that you can force it to stop running whatever it's running whenever you want. SGX and the like aren't like Intel's Management Engine (which really could be used by a state actor to plant a non-removable "ring -3" rootkit on your PC); instead, SGX is more like a TPM, or an FPGA: it's something that's ultimately controlled by the CPU from ring 0, just with a very circumscribed API that doesn't give the CPU the ability to "get in the way" of a workload once the CPU has deployed that workload to it, other than by shutting that workload off.
As much as people like Richard Stallman might freak out at the above design, it really isn't the same thing as your employer having root on your wheelchair. It's more like how someone in a wheelchair knows that if they get on a plane, then they're not allowed to wheel their own wheelchair around on the plane, and a flight attendant will instead be doing that for them.
How does that translate to employer MDM software?
Well, there's no clear translation currently, because we're currently in a paradigm that favors employer-issued devices.
But here's what we could do:
• Modern PCs are powerful enough that anything a corporation wants you to do, can be done in a corporation-issued VM that runs on the computer.
• The employer could then require the installation of an integrity-verification extension (essentially "anti-cheat for VMs") that ensures that the VM itself, and the hypervisor software that runs it, and the host kernel the hypervisor is running on top of, all haven't been tampered with. (If any of them were, then the extension wouldn't be able to sign a remote-attestation packet, and the employer's server in turn wouldn't return a decryption key for the VM, so the VM wouldn't start.)
• The employer could feel free to MDM the VM guest kernel — but they likely wouldn't need to, as they could instead just lock it down in much-more-severe ways (the sorts of approaches you use to lock down a server! or a kiosk computer!) that would make a general-purpose PC next-to-useless, but which would be fine in the context of a VM running only line-of-business software. (Remember, all your general-purpose "personal computer" software would be running outside the VM. Web browsing? Outside the VM. The VM is just for interacting with Intranet apps, reading secure email, etc.)
(Why yes, I am describing https://en.wikipedia.org/wiki/Multilevel_security.)
> The anti-cheat software isn't actually a rootkit (despite what some people say); it can't affect any app that doesn't explicitly hook into it.
Out of all examples you could have cited, you chose this one.
https://www.theregister.com/2016/09/23/capcom_street_fighter...
https://twitter.com/TheWack0lian/status/779397840762245124
There you go. An anti-cheat rootkit so ineptly coded it serves as literal privilege escalation as a service. Can we stop normalizing this stuff already?
My computer is my computer, and your computer is your computer.
The game company owns their servers, not my computer. If their game runs on my machine, then cheating is my prerrogative. It is quite literally an exercise of my computer freedom if I decide to change the game's state to give myself infinite health or see through walls or whatever. It's not their business what software I run on my computer. I can do whatever I want.
It's my machine. I am the god of this domain. The game doesn't get to protect itself from me. It will bend to my will if I so decide. It doesn't have a choice in the matter. Anything that strips me of this divine power should be straight up illegal. I don't care what the consequences are for corporations, they should not get to usurp me. They don't get to create little extraterritorial islands in our domains where they have higher power and control than we do.
I don't try to own their servers and mess with the code running on them. They owe me the exact same respect in return.
Sure.
However, due to the nature of how these games work, cheating cannot be prevented serverside only.
So, if you want to play the game, you have to agree to install the anti-cheat because it's the only way to actually stop cheating.
The *only other alternative is to sell a separate category of gaming machines where users wouldn't have access to install cheats, using something like the TPM to enforce.
Sure, you don't have to agree the earth isn't flat either. But then, as with here, you'd be entirely wrong.
> We're not about to sacrifice our power and freedom for the sake of preventing cheating in video games
Sure we are, gladly. Maybe not you or me, but most people absolutely.
If you want to play AAA games, that's the compromise. Until they release limited gaming PCs that are basically consoles.
> we're going to impose some of our terms and conditions on these things.
I doubt that but wish all the best. It won't change that the consumer end needs to be locked down to prevent cheating though.
What a bizarre leap of logic. Can Fedex employees reasonably sell their non-uniform clothes? Just because the employer in this scenario didn't 100% lock down the computer (which is a good thing because the alternative would be incredibly annoying for day-to-day work), doesn't mean the the employee can treat it as their own. Even from the privacy perspective, it would be pretty silly. Are you going to use the employer provided computer to apply to your next job?
Also, many people don't own a separate "personal" computer in the first place. Especially, again, poor people. (I know many people who, if needing to use "a PC" for something, would go to a public library to use the computers there.)
Not every job is a software dev position in the Bay Area, where everyone has enough disposable income to have a pile of old technology laying around. Many jobs for which you might be issued a work laptop still might not pay enough to get you above the poverty line. McDonald's managers are issued work laptops, for instance.
(Also, disregarding economic class for a moment: in the modern day, most people who aren't in tech solve most of their computing problems by owning a smartphone, and so are unlikely to have a full PC at home. But their phone can't do everything, so if they have a work computer they happen to be sat in front of for hours each day — whether one issued to them, or a fixed workstation at work — then they'll default to doing their rare personal "productivity" tasks on that work computer. And yes, this does include updating their CV!)
---
Maybe you can see it more clearly with the case of company cars.
People sometimes don't own any other car (that actually works) until they get issued a company car; so they end up using their company car for everything. (Think especially: tradespeople using their company-logo-branded work box-truck for everything. Where I live, every third vehicle in any parking lot is one of those.)
And people — especially poorer people — also often sell their personal vehicle when they are issued a company car, because this 1. releases them from the need to pay a lease + insurance on that vehicle, and 2. gets them possibly tens of thousands of dollars in a lump sum (that they don't need to immediately reinvest into another car, because they can now rely on the company car.)
There are also fairly obvious differences between work-issued computers and all of your other analogies:
1. A car (and presumably the cyberpunk android body) is much more expensive than a computer, so the downside of owning both a personal and a work one is much higher.
2. A chair or a wheel chair doesn't need security monitoring because it's a chair (I guess you could come up with an incredibly convoluted scenario where it would make sense to put GPS tracking in a wheelchair, but come on).
> just buys the person the chair. And then it's the employee's chair.
It's not because there's a law against loaning chairs, it's because the chair is likely customized for a specific person and can't be reused. Or if you're talking about WFH scenarios, they just don't want to bother with return shipping.
Which is, again, a situation so shitty that we've outlawed it entirely! And then also imposed further regulations on regular, non-employer landlords, about what kinds of conditions they can impose on tenants. (E.g. in most jurisdictions, your landlord can't restrict you from having guests stay the night in your room.)
Tenants' rights are actually a great analogy for what I'm talking about here. A company-issued laptop is very much like an apartment, in that you're "living in it" (literally and figuratively, respectively), and that you therefore should deserve certain rights to autonomous possession/use, privacy, freedom from restriction/compromise in use, etc.
While you don't literally own an apartment you're renting, the law tries to, as much as possible, give tenants the rights of someone who does own that property; and to restrict the set of legal justifications that a landlord can use to punish someone for exercising those (temporary) rights over their property.
IMHO having the equivalent of "tenants' rights" for something like a laptop is silly, because that'd be a lot of additional legal edifice for not-much gain. But, unlike with real-estate rental, it'd actually be quite practical to just make the "tenancy" case of company IT equipment use impossible/illegal — forcing employers to do something else instead — something that doesn't force employees into the sort of legal area that would make "tenants' rights" considerations applicable in the first place.
Agreed. I'm fine with this, as long as the employer also accepts that I will never use a personal device for work, that I will never use a minute of personal time for work, and that my productivity is significantly affected by working on devices and systems provided and configured by the employer. This knife cuts both ways.
That's a strange analogy, since the vault is meant to safeguard customer assets from the public, not from bank employees. Besides, the vault doesn't make the teller's job more difficult.
> What is the mechanism that would prevent your computer from getting infected by a 0-day if only your employer trusted you?
There isn't one. What my employer does is trust that I take care of their assets and follow good security practices to the best of my abilities. Making me install monitoring software is an explicit admission that they don't trust me to do this, and with that they also break my trust in them.
Every Google client device has it.
The other issue is that detection engineering is really expensive, so the detections that are included with CrowdStrike out of the box are your problem if you're using a free product. From a cost perspective you're not getting off a lot cheaper and trying to sell open source and a detection engineer's salary to a CISO who can just buy CrowdStrike instead is understandably a pretty tough sell. Or it was until this weekend, anyway.
As a red teamer developing malware for my team to evade EDR solutions we come across, I can tell you that EDR systems are essential. The phrase "root kit powered endpoint surveillance" is a mischaracterization, often fueled by misconceptions from the gaming community. These tools provide essential protection against sophisticated threats, and they catch them. Without them, my job would be 90% easier when doing a test where Windows boxes are included.
> So the main tool would be open source and it would be transparent what it does exactly and that it is free of backdoors or really bad bugs.
Open-source EDR solutions, like OpenEDR [1], exist but are outdated and offer poor telemetry. Assembling various GitHub POCs that exist for production EDR is impractical and insecure.
The EDR sensor itself becomes the targeted thing. As a threat actor, the EDR is the only thing in your way most of the time. Open sourcing them increases the risk of attackers contributing malicious code to slow down development or introduce vulnerabilities. It becomes a nightmare for development, as you can't be sure who is on the other side of the pull request. TAs will do everything to slow down the development of a security sensor. It is a very adversarial atmosphere.
> On the other hand it could still be a business model to supply malware signatures as a security team feeding this system.
It is actually the other way around. Open-source malware heuristic rules do exist, such as Elastic Security's detection rules [2]. Elastic also provides EDR solutions that include kernel drivers and is, in my experience, the harder one to bypass. Again, please make an EDR without drivers for Windows, it makes my job easier.
> *It could be audited by the public."
The EDR sensors already do get "audited" by security researchers and the threat actors themselves. Reverse engineering and debugging the EDR sensors to spot weaknesses that can be "abused." If I spot things like the EDR just plainly accepting kernel mode shellcode and executing it, I will, of course, publicly disclose that. EDR sensors are under a lot of scrutiny.
[1] https://github.com/ComodoSecurity/openedr [2] https://github.com/elastic/detection-rules
This is a such tired non-sequitur argument with no evidence whatsoever to back it up that the risk is actually higher for open source versus closed source.
I can just easily argue that a state or non-state actor could buy[1], bribe or simply threaten to get weak code in a proprietary system, without users having any means to ever find out. On the other hand, it is always easier(easier not easy) to discover compromise in open-source like it happened with xz[2] and verify such reports independently.
If there is no proof that compromise is less likely with closed source and it is far easier to discover them in open-source, the logical conclusion is simply open source is better for security libraries.
Funding defensive security infrastructure which is open source and freely available for everyone to use even with 1/100th of the NSA budget that is effectively only offensive, would improve info-security enormously for everyone not just from nation state actors, but also from scammers etc. Instead we get companies like CS that have enormous vested interest in seeing that never happens and trying to scare the rest of us that open-source is bad for security.
I feel having the solution open sourced isn't bad from a code security standpoint, but rathee that it is simply not economically viable. To my knowledge most of the major open source technologies are currently funded by FAANG and purely because it's needed by them to conduct business and the moment it becomes inconvenient for them to support it they fork it or develop their own, see Terraform/Redis...
I also cannot get behind a government funding model purely because it will simply become a design by committee nightmare because this isn't flashy tech. Just see how many private companies have beaten NASA to market in a pretty well funded and very flashy industry. The very government you want to fund these solutions are currently running on private companies infrastructure for all their IT needs.
Yes opensouring is definitely amazing and if executed well will be better, just like communism.
Government has to fund not run it like any other grant works today. The existing foundations and non profits like Apache or even mixed ones like Mozilla are fairly capable of handling the grants.
Expecting private companies or dedicated volunteers to maintain mission critical libraries like xz is not a viable option as we are doing it now.
Eventually that's something that gets exposed anyways, but I think the crucial part is timing and being a few steps ahead in the cat and mouse game. Otherwise I'm not sure what kind of proof would even be meaningful here.
That is not what am saying, I am saying open sourcing doesn’t cause more problems than proprietary systems which is the argument OP was making .
Open source is not a panacea, it is just not objectively worse as OP implies.
On the other hand, anyone serious about malware development already has "the actual source code", either for defensive operations and offensive operations.
Bazzar works absolutely fine for security, Linux kernel is one project which does this , all security infrastructure uses it one way or another. The tens of thousands of patches and forks has not once been discovered to have the subtle bug/vulnerability scenario intentionally submitted yet in 30 years .
There seems to be a lot of misconceptions in this thread what open source is or can do. Most of my points have been made by people much better than me for decades now.
How exactly is this is mischaracterization? Technically these EDR tools are identical to kernel level anticheat and they are identical to rootkits, because fundamentally they're all the same thing just with a different owner. If you disagree it would be nice if you explained why.
As for open source EDRs becoming the target, this is just as true of closed source EDR. Cortex for example was hilariously easy to exploit for years and years until someone was nice enough to tell them as much. This event from CrowdStrike means that it's probably just as true here.
The fact that the EDR is 90% of the work of attacking a Windows network isn't a sign that we should continue using EDRs. It means that nothing privileged should be in a Windows network. This isn't that complicated, I've administered such a network where everything important was on Linux while end users could run Windows clients, and if anything it's easier than doing a modern Windows/AD deployment. Good luck pivoting from one computer to another when they're completely isolated through a Linux server you have no credentials for. No endpoint should have any credentials that are valid anywhere except on the endpoint itself and no two endpoints should be talking to each other directly: this is in fact not very restrictive to end users and completely shuts down lateral movement - it's a far better solution than convoluted and insecure EDR schemes that claim to provide zero-trust but fundamentally can't, while following this simple rule actually provides you zero-trust.
Look at it this way - if you (and other redteamers) can economically get past EDR systems for the cost of a pentest, what do you think competent hackers with economies of scale and million dollar payouts can do? For now there's enough systems without EDRs that many just won't bother, but as it spread more they will just be exploited more. This is true as well of the technical analogue in kernel anticheat, which you and I can bypass in a couple days of work.
Where we are is that we're using EDRs as a patch over a fundamentally insecure security model in a misguided attempt to keep the convenience that insecurity brings.
People don't go around complaining that Microsoft Defender is "rootkit powered endpoint surveillance". It's intent is to protect the system.
There is a lot more suspicion around kernel level anti-cheat software developed by the likes of Epic games due to their ownership than they Crowdstrike or Microsoft.
People have been complaining about rootkit powered antimalware for a long time. It didn't start with CrowdStrike: there was a whole debacle about it in the Windows XP days when Microsoft stopped antiviruses from patching the kernel.
You can very easily shoot your own foot off instead of slaying the monster, use the wrong ammunition to be effective, or in this case a poorly crafted gun can explode in your hand when you are holding it.
DAT-style content updates and signature-based prevention are very archaic. Directly loading content into memory and a hard-coded list of threats? I was honestly shocked that CS was still doing DAT-style updates in an age of ML and real-time threat feeds. There are a number of vendors who've offered it for almost a decade. We use one. We have to run updates a couple of times a year.
SMH. The 90's want their endpoint tech back.
I’m not sure why you find that hard to believe - based on the (admittedly fairly limited) evidence we have right now, it’s highly unlikely that this deployment was tested much, if at all. It seems much more likely to me that they were playing fast and loose with definition updates to meet some arbitrary SLAs[1] on zero-day prevention, and it finally caught up with them. Much more likely than somehow every single real-world pc running their software being affected but their test machines somehow all impervious.
[1] When my company was considering getting into endpoint security and network anomaly detection, we were required on multiple occasions by multiple potential clients to provide a 4-hour SLA on a wide number of CVE types and severities. That would mean 24/7 on-call security engineers and a sub-4-hour definition creation and deployment. Yes, that 4 hours was for the deployment being available on 100% of the targets. Good luck writing and deploying a high-quality definition for a zero day in 4 hours, let alone running it through a test pipeline, let alone writing new tests to actually cover it. We very quickly noped out of the space, because that was considered “normal” (at least to the potential clients we were discussing). It wouldn’t shock me if CS was working in roughly the same way here.
You do not deploy anything, ever on your entire production fleet at the same time and you do not buy software that does that. It's madness and we're not talking about small companies with tiny IT departments here.
Source: https://x.com/patrickwardle/status/1814367918425079934
Note how the incident disproportionally affected highly regulated industries, where businesses don't have a choice to screw "best practice".
The problem here is that this type of update (a content update) should never be able to cause this however badly it goes. In case the software receives a bad content update, it should fail back to the last known good content update (potentially with a warning fired off to CS, the user, or someone else about the failed update).
In principle, updates that could go wrong and cause this kind of issue should absolutely be deployed slowly, but per my understanding, that’s already the practice for non-content updates at CrowdStrike.
You do not know if a content update will screw you over and mark all the files of your company as malware. The "It should never happen" situations are the thing you need to prepare for, the reason we talk about security as an onion, the reason we still do staggered production releases with baking times even after tests and QA have passed...
"But it's cybersecurity" is not a justification. I know that security departments and IT departments and companies in general love dropping the "responsibility" part on someone else, but in the end of the day the thing getting screwed over is the company fleet. You should retain control and make sure things work properly, the fact those billion dollar revenue companies are unable to do so is a joke. A terrible one, since IT underpins everything nowadays.
The CS customer has decided to update whenever 24/7 CS says. The alternative is to arrive on Monday morning to an infected fleet.
The CS customer has decided to offload the responsibility of its fleet to CS. In my opinion that's bullshit and negligence (it doesn't mean I don't understand why they did it), particularly at the scale of some of the customers :)
Incorrect, I believe, given they did not and could not get advance sight of the offending forced update.
Incorrect, I believe, given they could and did not get advance sight of the offending forced update.
Companies choose to work with Crowdstrike. One of the reasons they do that is ‘hands-off’ administration-let a trusted partner do it for you. There are absolutely risks of doing it this way. But there are also risks of doing it the other way.
The difference is, if you hand over to Crowdstrike, you’re not on your own if something goes wrong. If you manage it yourself, you’ve only got yourself working on the problem if something goes wrong.
Or worse, something goes wrong and your vendor says “yes, we knew about this issue and released the fix in the patch last Tuesday. Only 5% of your fleet took the patch? Oh. Sounds like your IT guys have got a lot of work on their hands to fix the remaining 95% then!”.
I am sympathetic to that, but its only possible if both policy and staffing allow.
for policy, there are lots of places that demand CVEs be patched within x hours depending on severity. A lot of times, that policy comes from the payment integration systems provider/third party.
However you are also dependent on programs you install not autoupdating. Now, most have an option to flip that off, but its not always 100% effective.
We are not talking about small companies here. We're talking about massive billion revenue enterprises with enormous IT teams and in some cases multiple NOCs and SOCs and probably thousands consultants all around at minimum.
I find it hard to be sympathetic to this complete disregard of ownership just to ship responsibility somewhere else (because this is the need at the of the day let's not joke around). I can understand it, sure, and I can believe - to a point - someone did a risk calculation (possibility of crowdstrike upgrade killing all systems vs hack if we don't patch a CVE in <4h), but it's still madness from a reliability standpoint.
> for policy, there are lots of places that demand CVEs be patched within x hours depending on severity.
I'm pretty sure leadership when they need to choose between production being down for an unspecified amount of time and taking the risk of delaying (of hours in this case) the patching will choose the delay. Partners and payment integration providers can be reasoned with, contracts are not code. A BSOD you cannot talk away.
Sure, leadership is also now saying "but we were doing the same thing as everyone else, the consultants told us to and how could have we have known this random software with root on every machine we own could kill us?!" to cover their asses. The problem is solved already, since it impacted everyone, and they're not the ones spending their weekend hammering systems back to life.
> However you are also dependent on programs you install not autoupdating. Now, most have an option to flip that off, but its not always 100% effective.
You choose what to install on your systems, and you have the option to refuse to engage with companies that don't provide such options. If you don't, you accept the risk.
And if an attacker does??
- Lack of testing of a deployment - Lack of required procedures to validate a deployment - Engineering management prioritizing release pace over stability/testing - Management prioritizing tech debt/pentests/etc far too low - Sales/etc promising fast turnarounds that can’t be feasibly met while following proper standards - Lack of top-down company culture of security and stability first, which should be a must for any security company
This outage wasn’t caused only by “the intern pushing release.” It was caused by a poor company culture (read: incorrect direction from the top) resulting in a lack of testing of the program code, lack of testing environment for deployments, lack of formal deployment process, and someone messing up a definition file that was caught by 0 other employees or automated systems.
https://news.ycombinator.com/item?id=%2041002864
Very much not surprised to see this now.
Now, if these sorts of things were battle tested before release, and had a (ideally decade+-long) history of stability with well-documented processes to ensure that stability, you can more easily make the argument that it’s worth it. None of those things are close to true though (and more than likely will never be for any AV/endpoint solution), so it is very hard to justify this sort of configuration.
Store something like an `attemptingUpdate` flag before updating, and remove it if the update was successful. Upon system startup, if the flag is present, revert to the previous config and mark the new config bad.
This is why we can’t have nice things, but maybe we just don’t want them anyway? “Mistakes will be made” is way less true if you actually put the effort in to prevent them, but I am beginning to think this has become code for quiet-quitters to telegraph a “I want to get paid for no effort and sympathize with others who feel the same” sentiment and appear compassionate and grimly realistic all at the same time.
yes, billion dollar companies are going to make mistakes, but almost always because of cost cutting, willful ignorance, or negligence. If average people are apologizing for them and excusing that, there has to be some reason that it’s good for them.
I don’t mind meetings but being in a 4 hour emergency meeting because some due diligence wasn’t done is a waste of my time.
Life is easier when you do good work.
Pipeline 1 --
Code updates to their software are treated as material changes that require non-production and canary testing before global roll-out of a new "Version".
Pipeline 2 --
Content / channel updates are handled differently -- via a separate pipeline -- because only new malware signatures and the like are distrubuted via this route. The new files are just data files -- they are supposed to be in a standard format and only read, not "executed".
This pipeline itself must have been tested originally and found tobe working satisfactorily -- but inside the pipeline there is no "test" stagethat verifies the integrity of the data fine so generated, nor - more importantly - checking if this new data file works without errors when deployed to the latest versions of the software in use.
The agent software that reads these daily channel files must have been "thoroughly" tested (as part of pipeline 1) for all conceivable data file sizes and simulated contents before deployment. (any invalid data files should simply be rejected with an error ... "obviously")
But the exact scenario here -- possibly caused by a broken pipeline in the second path (pipeline 2) -- created invalid data files with some quirks. And THAT specific scenario was not imagined or tested in the software version dev-test-deploy pipeine (pipeline 1).
If this is true --
The lesson obviously is that even for "data" only distributions and roll-outs, however standardized and stable their pipelines may be, testing is still an essential part before large scale roll-outs. It will increase cost and add latency sure, but we have to live with it. (similar to how people pay for "security" software in the first place)
Same lesson for enterprise customers as well -- test new distributions on non-production within your IT setup, or have a canary deployment in place before allowing full roll-outs into production fleets.
It was mentioned in one of the HN threads, that the update was pushed overriding the settings customer had [1]. What recourse any customer can have in in such a case ?
Sue them and use something else.
We got this update pushed right through.
But the problem here is that the code runs in kernel mode. As such any data that it may consume should have been tested with the same care as the code itself which has never been the case in this industry.
And of of course that cost would be absolutely insignificant relative to the potential risk...
This would have been poor, but to have released it with no testing would have been the most staggering negligence.
Why the blas radius was so huge?
I have deployed much less important services much more slowly with automatic monitoring and rollback in place.
You first deploy to beta, where you don't get customers traffic, if everything goes right to a small part of your fleet, and slowly increase the percentage of hosts that receives the updates.
This would have stopped the issue immediately, and I somehow I thought it was common practices...
There are so many ways to avoid this issue, or at least minimize the risk of it happening, but as always profits come before people.
The expectation being that people want up-to-date virus detection rules constantly even if they don't want potentially breaking changes.
The missed edge case being an untested config that breaks existing code.
Source: Pure speculation, don't quote this in news articles.
This situation is akin to the immune system overreacting and melting the patient in response to a papercut. This sometimes happens, but it's considered a serious medical condition, and I believe the treatment is to nuke someone's immune system entirely with hard radiation, and reinstall a less aggressive copy. Take from that analogy what you want.
Yes they do? And it’s more akin to a shared immune system than a single organism.
In this case, it’s not like viruses move fast relative to the total population of machines, but within the population of machines being targeted they do move fast.
Just one.
To the smart people below:
It’s clear to everyone that 70 minutes is not 1 month. The point is that it’s not a fair comparison: it would simply not have been possible to infect that many computers in 70 minutes: the internet infrastructure just wasn’t there.
It’s like saying “the Spanish flu didn’t do that much damage because there where less people on the planet” - it’s a meaningless absolute comparison, whereas the relative comparison is what matters.
Be better.
Computer security as a whole has improved, whilst the complexity of interconnected systems has exponentially increased.
This has made the barrier to entry for malware higher, and so means we no longer have the same historic examples of large scale worms targeting consumer machines that we used to.
At the same time the financial rewards for finding and exploiting a vulnerability within an organisations complex stack have greatly increased. The rewards are coupled to the time it takes to execute on the vulnerability.
This leads to what we have today: localised, and often specialised attacks against valuable targets that are executed as fast as possible in order to minimise the chance a target has to respond or the vulnerability they are exploiting to be burned.
Of course the “smart people belw” must know this, so it’s unclear why they are pretending to be dumb.
Yup, exactly that.
So what I'm saying it, it's beyond idiotic to combat this with a kernel-level backdoor managed by one entity and deployed across half the Internet. If anyone manages to breach that, they have a way to make their attack much simpler and much less localized (though they're unlikely to be prepared to capitalize on that). A fuckup on the defense side, on the other hand, can kill everything everywhere all at once. Which is what just happened.
It's a "cure" for disease that happens to both boost the potency of the disease, and, once in blue moon, randomly kills the patient for no reason.
The fact is that this does help organisations. Definitely not all of the orgs that buy Crowdstrike, but rapid defence against evolving threats is a valuable thing for companies.
So, individually it’s good for a company. But as a whole, and as currently implemented, it’s not good for everyone.
However that doesn’t matter. Because individually it’s a benefit.
Which is why I'm hoping that this incident will make both security professionals and regulators reconsider the idea of endpoint security as it's currently done, and that there will be some cultural and regulatory pushback. Maybe this will incentivize people to come up with other ideas on how to secure systems and companies, that don't look like a police state on steroids.
The underlying fault in this drama is Microsoft - third party code shouldn’t be able to have the impact it did, regardless of how it is loaded or what it does. Their commitment to supporting legacy interfaces has shot them in the foot here.
If HP pushed a dodgy printer driver (and if those still lived in the kernel) that nuked tens of millions of machines, would you be out here saying “regulators and security professionals need to re-consider printers”?
Microsoft will shit bricks, start to do something to isolate kernel modules, Crowdstrike will be the first shining user of this, and life will go on.
> infected millions of Windows computers worldwide within a few hours of its release
See: https://en.wikipedia.org/wiki/Timeline_of_computer_viruses_a...
And operating systems aren't that bad anymore. You don't have services out of the box opening ports on all the interfaces, no firewalls, accepting connections from everywhere, and using well-known default (or no) credentials.
Even stuff like the recent OpenSSH bug that is remotely exploitable and grants root access wasn't anything close to this kind of disaster because (a) most computers are not running SSH servers on the public internet (b) the exploit is rather difficult to actually execute. Eventually it might not be, but that gives people a bit of breathing space to react.
Most cyberattacks use old, unpatched vulnerabilites against unprotected systems combined with social engineering to get the payload past the network boundary. If you are within a pretty broad window of "up to date" on your OS and antivirus updates, you are pretty safe.
It just needs to infect 200k devices to get to the pot: hundred million dollars of ransomware.
(I expect this to tally up to double-digit billions and thousands of lives lost directly to the outages when the dust settles.)
The organization like MGM and London Drugs?
[SQL] Slammer spread incredibly quickly, even though the vulnerability was patched in the prior year.
> As it began spreading throughout the Internet, it doubled in size every 8.5 seconds. It infected more than 90 percent of vulnerable hosts within 10 minutes.
Worms are not technically viruses, but they can have similar impacts/perform similar tasks on an infected host.
Also keep in mind 8.5 million is likely the count of machines fully impacted and are not counting the machines impacted but were able to be automatically recovered.
Can you cite something? This is HN, not reddit.
> Also keep in mind 8.5 million is likely the count of machines fully impacted and are not counting the machines impacted but were able to be automatically recovered.
Do you have evidence of this? Please bring sources with you.
1. Crowdstrike didn’t test adequately
2. Viruses can move pretty fast once a foothold is gained
There is no real "speed limit" on malware spread.
The proper place would be Ring 1, which doesn't exist on Windows.
And being a kernel-level operation, it has the capability to crash the whole system before the actual OS has any chance to intervene.
Their precognitive intelligence suggested that a world wide attack was only moments away. The same precognitive system showed that the virus was so totally incapacitating that the only safe response was to incapacitate the server.
Knowing that the virus was capable of taking down every crowdstrike server, they didn’t waste time trying it on a subset of servers.
When you know you know.
Hmm, maybe you could have companies pay more to be in the first rollout group? That'd go over well too.
It would be rather easier to understand and explain if it were intentional. Likely not able to be discussed though.
Anyone able to do that here?
While they were not directly responsible for the bug that caused the crashes, Microsoft does hold an effective monopoly position over workstation computing space (I'd consider this as infrastructure at this point) and therefore have a duty of care to ensure the security/reliability and capabilities of their product.
Without competition, Microsoft have been asleep at the wheel on innovations to Windows - some of which could have prevented this outage.
For example; Crowdstrike runs in user space on MacOS and Linux - does Windows not provide the capabilities needed to run Crowdstrike in user space?
What about innovations in application sandboxing which could mitigate the need for level of control CrowdStrike requires?
The fact is; Microsoft is largely uncontested in holding the keys to the world's computing infrastructure and they have virtually no oversight.
Windows has fallen from making over 80% of Microsoft's revenue to 10% today - there is nothing wrong with being a private company chasing money - but when your product is critical to the operation of hospitals, airlines, critical infrastructure, you can't be out there tickling your undercarriage on AI assistants and advertisements to increase the product's profitability.
IMO Microsoft have dropped the ball on their duty of care to consumers and CrowdStrike is a symptom of that. Governments need to seriously consider encouraging competition in the desktop workspace market. That, or regulate Microsoft's Windows product
The one upside of the fall of revenue share from Windows is that it means that MS probably won't hold it as an untouchable sacred cow any more.
Microsoft has long desired to kick AV vendors out of kernel space and has even attempted to do so prior, however because of its dominant position in the market, it is unable to do so. I was at MS when an iteration of this effort was underway, and the EU said no.
See, Windows is a highly regulated OS today, and making a change like kicking out AV vendors from the kernel runs afoul of antitrust laws.
Example: https://www.techtarget.com/searchsecurity/news/450420491/Mic...
Microsoft does provide user-space capabilities: https://learn.microsoft.com/en-us/windows/win32/amsi/antimal... but vendors are not required to use it, nor can Microsoft require vendors to use it (for the aforementioned antitrust reasons).
Microsoft also has ELAM: https://learn.microsoft.com/en-us/windows-hardware/drivers/i... which is a rootkit / bootkit defensive mechanism. A defect in the definition files (as noted in the twitter thread) is what caused the crash in an ELAM driver. CrowdStrike obviously was not following the required process for ELAM drivers.
Mind you, the claim about CrowdStrike not impacting Linux is also bogus: https://www.neowin.net/news/crowdstrike-broke-debian-and-roc...
My understanding was that CrowdStrike breaking on Debian was actually the motivation for them moving to user-space on Linux. I'm surprised that, assuming they have the capability to do so, they haven't done the same on Windows.
It seems great pains are made to ensure the CS driver is installed first _and_ cannot be uninstalled (presumably the remote monitor will notice) or tampered with (signed driver).
Then the driver goes and loads unsigned data files that can be arbitrarily deleted by end users? Can these files also be arbitrarily added by end users to get the driver to behave in ways that it shouldn’t? What prevents a malicious actor from writing a malicious data file and starting another cascade of failing machines or worse, getting kernel privileges?
I sure hope the certificate authorities and other crypto folks get to keep that stuff off their systems at least.
Well it's likely we don't see that because we might be one of the millions.
Any input for something that runs at such high privilege should be at least integrity checked. That’s the basics.
And the fact that you can simply delete these channel files suggests there isn’t even an anti-tamper mechanism.
(This is just an early guess from looking at some of the csagent in ida decompiler, haven't validated that all the sanity checks can be bypassed as these channel files appear to have some kind of signature attached to them.)
eBPF, much the same thing, is actually thought about and well designed. If it wasn't it would be easy to crash linux.
This is what they do and they are doing badly. I bet it's just shit on shit under the hood, developed by somewhat competent engineers, all gone or promoted to management.
This was obviously a bug in RHEL kernel because even if the bpf program was bunk it should not cause the kernel to panic. However, it's almost like CrowdStrike does zero testing of their software and looks at their end users as Test/QA.
https://access.redhat.com/solutions/7068083
> 4bb7ea946a37 bpf: fix precision backtracking instruction iteration
I’m not sure how much early warning RH gives to folks when a kernel change comes in via a point release. Looking at https://www.redhat.com/en/blog/upcoming-improvements-red-hat..., it seems like it’s changing for 9.5. I hope CrowdStrike will be able to start testing against those beta kernels.
Some still-open questions in my mind:
- was the broken rule in the config file (C-00000291-...32.sys) human authored and reviewed or machine-generated?
- was the config file syntactically or semantically invalid according to its spec?
- what is the intended failure mode of the kernel driver that encounters an invalid config (presumably it's not "go into a boot loop")?
- what automated testing was done on both the file going out and the kernel driver code? Where would we have expected to catch this bug?
- what release strategy, if any, was in place to limit the blast radius of a bug? Was there a bug in the release gates or were there simply no release gates?
Given what we know so far, it seems much more likely that this was a "disaster waiting to happen" but I still think there's a lot more to know. I look forward to the public post-mortem.
Seems to me 3rd party code, running in the kernel, on parsed inputs, that can be remotely updated is enough to be disaster waiting to happen gestures breezily at Friday
That's, in the Taleb parlance, a Fat Tony argument, but barring it being a cosmic ray causing a uncorrected bit flop during deploy, I don't think there's room to call it anything but "a disaster waiting to happen"
this code is executed only once during the driver initialization, so shouldn't be much overhead, but will greatly improve reliability against broken channel file
the natural cost of these bits we sell is zero, so in the long run, if the bar is "just write a good & tested kernel driver", there will always be one more subsequent market entrant who will go too cheap on engineering. Then, they touch the hot wire and burn down the establishment.
That doesn't mean capitalism bad, but it does mean I expect only Microsoft is capable of writing and maintaining this type of software in the long run.
Ex. The dentist and dental hygienist were asking me who was attacking Microsoft on Friday, and they were not going to get through to the the subtleties of 3rd kernel driver release gating strategy.
MS has a very strong incentive to fix this. I don't know how they will. But I love when incentives align and assume they always will, in the long run.
If they weren't following these practices, this is kind of a boring incident with not much to be learned, despite how dramatic the scale is. Practices like staged rollout of changes exist precisely because we've learned these lessons before.
My Tesla remote updates ... hmph.
It doesn't feel like this is inherently impossible. It feels more like not enough design/process to mitigate the risks.
Given the number of systems infected, if you could push code that rebooted every client into a compromised state you’d still have run of some % of the lot until it was halted. That time window could be invaluable.
Now, imagine if you screw up the code and just boot loop everything.
I’d say business wise it’s better for crowd strike to let people think it’s an own-goal.
The truth may be mundane but a hack is as reasonable a theory as “oops we pushed boot loop code to world+dog”.
No it's not. There are many signs that point to this being a mistake. There are very few that point to it being a hack. You can't just go "oh it being a hack is one of the options therefore it is also something worth considering".
And while I'm 99% for Hanlon's razor here, I don't see a reason to be sure it wasn't even a completely successful DoS attack.
I fail to see why this is so difficult to understand.
Many corporations have pretty strict rules on system update scheduling so as to ensure business continuity in case of situations like this but all of those were completely circumvented and we had fully synchronised global failure. It really does not seem like business as usual situation.
which crowdstrike gets to bypass because they claime themselves as an antivirus and malware detection platform - at least, this is what the executives they've wined and dined into the purchase contracts have been told. The update schedule is independently controlled by crowdstrike, rather than by a system admin i believe.
However, I doubt they need an instantaneous rollout for every deployment.
Only this time, crowdstrike itself has become indistinguishable from malware.
Because the point of these updates is to be rolled out quickly and globally. It wasn't a system/driver update, but a data file update: think antivirus signature file. (Yes, I know it can get complicated, and that AV signatures can be dynamic... not the point here.)
Why those data updates skipped validity testing at the source is another question, and one that CrowdStrike better be prepared to answer; but the tempo of redistribution can't be changed.
Is it realistic that there's a threat actor that will be attacking every computer on the whole planet at once?
I can understand that it's most practical to update everyone when pushing an update to protect a few actively under attack but I can also imagine policies where that isn't how it's done, while still getting urgent updates to those under attack.
Is this what people are paying CS for? Absolutely.
Given the choice I bet there's going to be surprisingly large number of orgs go "we'll take n+24hrs thanks"
Because it worked good for them so far? There are plenty of companies that do the same and we don’t hear about them until something goes wrong.
Quality starts with good design, good people, etc. the process parts come much after that. I'd like to think that if you do this "right" then this sort of stuff simply can't happen.
If we have organization/culture/engineering/process issues then we're likely not going to get an in-depth public most-mortem. I'd love to get one just for all of us to learn from it. Let's see. Given the cost/impact having something like the Challenger investigation with some smart uninvolved people would be good.
They don't say why it was invalid or really what the file is, but it seems like it is some kind of relatively complex set of rules that are evaluated by the kernel module. Presumably they are manually authored and reviewed and it seems possible the bug was missed in review because this was a relatively new type of rule.
So this isn't a case of an incident that slipped through a rigorous testing and release process process following industry best practice, but rather a "disaster waiting to happen". Further, CrowdStrike CEO George Kurtz should have known better, considering an analogous incident happened under his watch as CTO of McAfee in 2010.
https://www.crowdstrike.com/falcon-content-update-remediatio...
Typically there would also be some clauses where CS is the only one that is allowed to determine an SLA breach, SLA breaches only result in future licence credits no cash, and if you disagree it's limited to mandatory arbitration...
The biggest impact is probably only their reputation taking a huge hit. Loosing some customers over this and making it harder to win future business.
Commenter on stack exchange had an interesting counter: In some jurisdictions, any attempt to sidestep consumer law may be interpreted by the courts as conspiracy, which can prove more serious than merely accepting the original penalties.
i would imagine a class action suit instead of individual cases if this were to happen.
To be clear, I feel investors are a bit delusional, I just thought it was an interesting perspective to share.
On the other hand, this may open the veil for a lot of companies to dump them.
one possibility is: clawback or refunds for past payments equal to business damage caused by the flawed product.
Surely CrowdStrike encrypts and signs their channel files, and I'm wondering if a file full of 0's inadvertently signaled to the validating software than a 'null' or 'none' encryption algo was being used.
This could imply the file full of zeros is just fine, as the null encryption passes, because it's not encrypted.
That could explain why it tried to reference the null memory location, because the null encryption file full of zeroes just forced it to run to memory location zero.
The risk is, if this is true, then their channel loading verification system is critically exposed by being able to load malicious channel drivers through disabled encryption on channel files.
Just a hunch.
April 21, 2010
In 2010 McAffe caused a global IT meltdown due to a faulty update. CTO at this time was George Kurtz. Now he is CEO of crowdstrike
I don't think there's anything more to the inclusion of "French" in their comment beyond it being in the original line.
https://www.youtube.com/watch?v=VFevH5vP32s
and the successful version: https://www.youtube.com/watch?v=qb1KndrrXsY
I know there's safe mode, but that's the nuclear option, and safe mode isn't really "usable".
Couldn't a lot of this been avoided if Windows could just retry its boot after BSOD without the faulting module, and then they could push out a new module with a fix shortly after?
where are tweets from sama and amodei on how agi is going to fix these issues ?
Surely an OS doesn't have to go completely kaput due to one service crashing.
https://manifold.markets/ChrisGreene/why-didnt-the-crowdstri...
It's not guaranteed that NULL is 0.
Still, I don't think you'd find a counterexample in the wild these days.
That's true for the literal constant 0. For 0 in a variable it is not necessarily true. Basically when a literal 0 is assigned to a pointer or compared to a pointer the compiler takes that 0 to mean whatever bit pattern represents the null pointer on the target system.
int *array = NULL
int position = 0x9C
int a = *(array[pos]) //equivalent to *(array + 0x9C) - dereferencing NULL+0x9C, which is just 0x9C
This will segfault (or equivalent) due to reading invalid memory at address 0x9C. Most people would call array[pos] a null pointer dereference casually, even though it’s actually a 0x9C pointer dereference, because there’s very little effective difference between them.
Now, whether this case was actually something like this (dereferencing some element of a null array pointer) or something like type confusion (value 0x9C was supposed to be loaded into an int, or char, or some other non-pointer type) isn’t clear to me. But I haven’t dug into it really, someone smarter than me could probably figure out which it is.
struct BigObject {
char stuff[0x9c]; // random fields
int field;
}
BigObject* object = nullptr;
printf("%d", object->field);
That will result in "Attempt to read from address 0x9c". Just because it's not trying to read from literal address 0x0 doesn't mean it's not nullptr error.R8 is 0x9c in that example, which is somewhat typical for null+offset, but in the twitter thread it's 0xffff9c8e0000008a.
So the actual bug is further back. It's not a null pointer dereference, but it somehow results in the mov r8, [rax+r11*8] instruction reading random data (could be anything) into r8, which then gets used as a pointer.
Maybe this is a use-after-free?
So no an unmapped address is a completely different BSOD, usually PAGE_FAULT_IN_UNPAGED_AREA which is a very bad sign
(DRIVER_)IRQL_NOT_LESS_OR_EQUAL[2][3] is not this case, but it's probably one of the most common reasons drivers crash the system generally. Like you said it's basically attempting to access pageable memory at a time that paging isn't allowed (i.e. when at DISPATCH_LEVEL or higher).
[1]: https://learn.microsoft.com/en-us/windows-hardware/drivers/d...
[2]: https://learn.microsoft.com/en-us/windows-hardware/drivers/d...
[3]: https://learn.microsoft.com/en-us/windows-hardware/drivers/d...
They could have named their files "foo.cfg", "foo.dat", "foo.bla" and been equally protected.
The use of ".sys" here is probably related to the fact it is used by their system driver. I don't think anybody was trying to pretend the files there are system drivers themselves, and a quick look at the exports/disassembly would make that apparent anyway.
Of course we didn't have had any third party loading code into our boxes out of our control (and we run linux)
She said, "I installed it before for my cybersecurity course but I think it was just a trial"
Assumptions eh.
User space downloads file.
User space sets up probation dir.
User space requests kernel to load once the new file.
After that, after a successful boot or 36 hours the file is marked as safe and set to autoload.
Or, you know, just load it. It will be cheaper. The ROI on loading it immediately is far greater and that's what counts.
Not in the security industry, but my take is that basically the desktop OS permissions and security model is wrong for a lot of these devices, but there is no alternative that is suitable or that companies are willing to invest in. Probably many of the highest-profile affected machines (airport terminals, signage, medical systems, etc.) should just resemble a phone/iPad/Chromebook in terms of security/trust, but for historical/cost/practical reasons are Windows PCs with Crowdstrike.
Correct me if I'm wrong but isn't kernel-level access essentially God Mode on every computer their software is installed on? Including spying on the entire memory, running any code, deleting data, installing ransomware? This feels like an insane amount of power concentrated into the hands of a single entity, on the level of a nuclear submarine. Wouldn't that make them a prime target for all sorts of nation-state actors?
This time the damage was (likely) unintentional and no data was lost (save for lost BitLocker keys), but were we really all this time one compromised employee away from the largest-ever ransomware attack, or even worse?
But the kernel driver obviously contains some bugs, so it's possible that those definition updates can inject code. There might be a bug inside the driver that allows code execution (it happens all the time that some file parsing code can be tricked into executing parts of the data). I'm not sure, but I guess a lot of kernel memory is not fully protected by NX bits.
I still have the gut feeling, that this incident was connected to some kind of attack. Maybe a distraction from another attack while everyone is busy about fixing all the clients. During this incident security measures were for sure lowered, lists with BitLocker keys printed out for service technicians to fix the systems. Even the fix itself was to remove some parts of the CroudStrike protection. I would really like to know what was inside the C-00000291*.sys file before the update replaced it with all zeros. Maybe it was a cleanup job to remove something concerning that went wrong. But Hanlon's razor tells me not to trust my gut: "Never attribute to malice that which is adequately explained by stupidity."
Could be wrong here so if anyone knows better and can correct me...plz do!
there are some certification requirements to do pentests/red teaming and then those security folk will all tell them to install an EDR so they picked crowdstrike, but the security people have a very valid technical case for that recommendation.
it doesn't shift liability to crowdstrike, thats not how this works. In this specific case they are very likely liable due to gross negligence, but that is different
Data was lost in the knock on effects of this, I assure you.
> largest-ever ransomware attack
A ransomware attack would be a terrible use of this power. A terrorist attack or cover while a country invades another country is a more appropriate scale of potential damage here. Perhaps even worse.
Isn't that every antivirus software and game anticheat?
And if any part of it is necessary, then that's a failure of the operating system. It should be a feature of Active Directory or Windows.
So, great job sales team, you earned your commissions, now get ready to jump ship, 'cause this one is sinking.
It’s hard not to be a little nihilistic about security.
My opinion is that CS is trying to say the null bytes themselves aren't the actual root cause of the issue, but merely a trigger for the actual root cause, which is that CSAgent.sys has a problem where malformed input vectors can cause it to crash. Well designed programs should error out gracefully for foreseeable errors, like corrupted config files.
If we interpret that quoted sentence such that "this" is referring to "the logical error", and that "the logical error" is the error in CSAgent.sys that causes it to crash upon reading a bad channel file, then that statement makes sense.
This is a bit of a stretch, but so far my impression with CS corporate communication regarding this issue has been nothing but abject chaos, so this is totally on-brand for them.
My opinion is they say "unrelated" because they are trying to say unrelated - and hence no, this was not a trigger.
It seems really scary to me, that crowdstrike is able to push updates in real time to most of their customers systems. I don't know of any other system, that would provide a similar method to inject code at kernel level. Not even windows updates, as they always roll out with some delay and not to all computers at the same time
If you want to attack high profile systems, crowdstrike would be one of the best possible targets.
Did any nation states/other groups have 0-days on this?
Did this event reveal something known to the public, or did this screw up accidentally protect us from someone finding + exploiting this in the future?
Most people don't need this stuff. Just keeping shit up to date, no not on the nightly build branch, but like installing windows update atleast a day or two after they come out. Or maby regular antivirus scans.
But let's be honest, your kernal drivers are useless if your employees fall for phishing or social engineering. See then its not malware, its an authorized user on the system....just copying data onto a USB drive or a rouge employee taking your customer list to your competition. That fancy pants kernal driver might be really good at stopping sophisticated threats and I'm sure the marketing majors at any company cram products full of buzz words. But remember, you can't fix incompetent or malicious employees unless your taking steps to prevent it.
What's more likely: some foreign government hacking khols? Or a script kiddie social engineers some poor worker pretending to be the support desk?
Not here to shit on this product, it has its place and it obviously does a good job....(heard its expensive but most xrd/edr is)
Seems like we are learning how vulnerable certain things are once again. As a fellow security fellow, I must say that Jia Tan must be so envious that he couldn't have this level of market impact.
After they cuss the hackers under their breath exclaiming something like: "they should be locked up in jail for the rest of their lives!...", tell them that's exactly what happened, but CS were the hackers, and maybe they should reconsider mandating installing that crap everywhere.
[1]: https://en.wikipedia.org/wiki/2020_United_States_federal_gov...
https://www.techtarget.com/whatis/feature/SolarWinds-hack-ex...
I'm hopeful that the fallout from Crowdstrike will be a larger emphasis on software BOM risk - when your systems regularly phone home for updates, you're at the mercy of the weakest link in that chain, and that applies to CI/CD and end user devices alike.
if you can get control of the crowdstrike deployment machinery
Or combine a lack of certificate pinning with BGP hijacking.To trigger the crash, you need to write a bad file into C:\Windows\System32\drivers\CrowdStrike\
You need Administrator permissions to write a file there, which means you already have code execution permissions, and don't need an exploit.
The only people who can trigger it over network are CrowdStrike themselves... Or a malicious entity inside their system who controls both their update signing keys, and the update endpoint.
Even if the HTTPS channel is compromised with a man-in-the-middle attack, the attacker shouldn't be able to craft a valid update, unless they also compromised CrowdStrke's keys.
However, the fact that this update apparently managed to bypass any internal testing or staging release channels makes me question how good CrowdStrike's procedures are about securing those update keys.
It's wild to me that it's so normal to install software like this on critical infrastructure, but questions about how they do code signing is a closely guarded/obfuscated secret.
Though, I prefer to give people benefit of doubt for this type of thing. IMO, the level of incompetence to parse a binary file before checking the signature is significantly higher (or at least different) than simply pushing out a bad update (even if the latter produces a much more spectacular result).
Besides, we don't need to speculate. We have the driver. We have the signature files [1]. Because of the publicity, I bet thousands of people are throwing it into Binary RE tools right now, and if they are doing something as stupid as parsing a binary file before checking it's signature (or not checking a signature at all), I'm sure we will hear about it.
We can't see how it was signed because that's happening on Cloudstrike's infrastructure, but checking the signature verification code is trivial.
[1] Both in this zip file: https://drive.google.com/file/d/1OVIWLDMN9xzYv8L391V1ob2ghp8...
That is, the data is signed and they don't want to use the real signing key during testing / in the continuous build because then it is too exposed.
So it's added after as something that "could not break". But it of course did.
This wasn't a code update, just a configuration update. Maybe they don't put config update though QA at all, assuming they are safe.
It's possible that QA is different enough from production (for example debug builds, or signature checking disabled) that it didn't detect this bug.
Might be an ordering issue, and that they tested applying update A then update B, but pushed out update B first.
The fact that it instantly went out to all channels is interesting. Maybe they tested it for the beta channel it was meant for (and it worked, because that version of the driver knew how to cope with that config) but then accidentally pushed it out to all channels, and the older versions had no idea what to do wiht it.
Or maybe they though they were only sending it to their QA systems but pushed the wrong button and sent it out everywhere.
Configuration is data, data is code.
Microsoft supposedly has source IP addresses known by their update clients, so that DNS spoofing won't work.
Microsoft has leaked keys that weren't used for code signing. I've been on the receiving end of this actually, when someone from the Microsoft Active Protections Program accidentally sent me the program's email private key.
Microsoft has been tricked into signing bad code themselves, just like Apple, Google, and everyone else who does centralized review and signing.
Microsoft has had certificates forged, basically, through MD5 collisions. Trail of Bits did a good write-up of this years ago.
But I can't think of a case of Microsoft losing control of a code signing key. What are you referring to?
Attackers today may be willing to spend a few million dollars to access those keys.
This would probably be classified as a terrorist attack and frankly it’s just a matter of time until we get one some day. A small dedicated team could pull it off. It’s just so happens that the people with the skills currently either opt for cyber criminality (crypto lockers and such), work for a state actor (think Stuxnet) or play defense in a cyber security firm.
I just don't understand how they still have users.
Because this post is here and not somewhere else. Strong network effects.
I'll concede the confusing part but all the major Mastodon servers I interact with regularly are pretty quick so I'm not sure where that part comes from.
Not since February. But it's for the best that the Eternal September has remained quarantined on Twitter.
No one wants to pay for anything, and that's the true root of every issue around this. People complain YouTube has ads, but wont buy premium. People hate Elon and Twitter but won't take even an ounce of temporary inconvenience to try and solve it.
Threads exists, I'm happy they integrate with Activity Pub, which should give us the best of both worlds. Why don't people use Threads? I'd a little more popular outside the US but personally, I think the "algorithm" pushes a lot of engagement bait nonsense.
Perhaps if buying into a service guaranteed that they would not be sold out then there would be more engagement. When someone signs up it is pretty much a rock-hard guarantee that their personal information will be marketed and sold to any entity with the money and interest to buy it - paying customers, free-loaders, etc.
When someone chooses to buy your app or SaaS then they should be excluded from the list of users that you sell or trade between "business partners".
When paying for a service guarantees that you're selling all details of your engagement with that service to unrelated business entities you have a disincentive to pay.
People are wising up to all this PII harvesting and those clowns who sold everyone out need to find a different model or quit bitching when real people choose to avoid their "services" since most of these things are not necessary for people to enjoy life anyway. They are distractions.
EDIT: This is not intended as a personal attack on you but is instead a general observation from the perspective of someone who does not use or pay for any apps or SaaS services and who actively avoids handing out accurate personal information when the opportunity arises.
In my experience, Mastodon is nice until you want to partake in discussions. To do so, you need an account.
With an account you can engage in civilized discussions. Some people don't agree with you, and you don't agree with some people. That's fine, maybe you'll learn something new. It's a discussion.
And then, suddenly, a secret court convenes and kills your account just like that; no reason will be given, no recourse will be available, admins won't reply, and you can do two things: go away for good, or try again on a different server.
I'm happy with a read-only Mastodon via a web interface.
But read-write? Never again, I probably don't have the correct ideology for it.
You're being too dismissive of Threads. It's fine, there are adults there.
What weirdo doesn't have an insta?
Some people don't jump on every fad out there. Most of the people who miss out on fads quickly realize that they aren't losing out on much simply because fads are so ephemeral. As far as I can tell, this is normal (though different people will come to that realization at different stages of their life).
As for ChatGPT, I'm sure time will prove it is a fad. That doesn't mean that LLMs are a fad (though it is too early to tell).
My wife only uses Facebook, and even then pretty sparingly.
The quality vs crap ratio is stellar on mastodon. Not so much on anywhere else.
The probable spam thing was nuts to me too. My guess was it's maybe trying to detect users with lower engagement. Like people who aren't moving the investigation forward but are trying to follow it and be in the discussion.
The basic problem is, no moderation results in a deluge of spam and algorithmic moderation is hot garbage that can only filter out the bulk of the spam by also filtering out like half of the legitimate comments. Human moderation is prohibitively expensive unless you want to hire Mechanical Turk-level moderators and not give them enough time to do a good job, in which case you're back to hot garbage.
Nobody really knows how to solve it outside of the knob everybody knows about that can improve the false negative rate at the expense of the false positive rate or vice versa. Do you want less ham or more spam?
The problem is also getting significantly worse because it's trivial to generate entire pages of inorganic content with LLMs.
The backstories of inorganic accounts are also much more convincing now that they can be generated by LLMs. Before LLMs, backstories all focused on a small handful of topics (e.g. sports, games) because humans had to generate them from playbooks of best pracitces. Now they can be into almost anything.
Most everything else goes through a filter and pasteurization before public consumption.
I always assumed that the reason legit answers often fall under "Show probable spam" is because of the inevitable reports coming in on controversial topics. It seems like the community notes feature works well most of the time.
Before this wave of insane bot spam, the comments had started to be so much better than what they used to be (low effort, boomer spam). In fact I think they were much better than the absolute cringy mess that comments on dedicated forums like Reddit turned into