CrowdStrike ex-employees: 'Quality control was not part of our process'
semafor.com
semafor.com
This type of article - built upon disgruntled former employees - is worth about as much as the apology GrubHub gift card.
Look, I think just as poorly about CrowdStrike as anyone else out there... but you can find someone to say anything, especially when they have an axe to grind and a chance at some spotlight. Not to mention this guy was a designer and wouldn't be involved in QC anyway.
> Of the 24 former employees who spoke to Semafor, 10 said they were laid off or fired and 14 said they left on their own. One was at the company as recently as this summer. Three former employees disagreed with the accounts of the others. Joey Victorino, who spent a year at the company before leaving in 2023, said CrowdStrike was “meticulous about everything it was doing.”
So basically we have nothing.
Except the biggest IT outage ever. And a postmortem showing their validation checks were insufficient. And a rollout process that did not stage at all, just rawdogged straight to global prod. And no lab where the new code was actually installed and run prior to global rawdogging.
I'd say there's smoke, and numerous accounts of fire, which this can be taken in the context of.
Sometimes people are up in arms "where's my next version" (eg when adaptive headlights was introduced), yet Tesla prioritise a safe, slow roll out. Sometimes the updates fail (and get resolved individually), but never on a global scale. (None experienced myself, as a TM3 owner on the "advanced" update preference).
I understand the premise of Crowdstrike's model is to have up to date protection everywhere but clearly they didn't think this through enough times, if at all.
When I read the same release notes so often I begin to question whether this redesign is really taking more than six months to roll out. And then I read the Sonos app disaster and I thought that was the other extreme.
Google is terrible at release notes. Since several years ago, the release notes for the "Google" app on the Android app store always shows the exact same four unchanging entries, loosely translating from Portuguese: "enhanced search page appearance", "new doodles designed for app experience", "offline voice actions (play music, enable Wi-Fi, enable flashlight) - available only in the USA", "web pages opened directly within the app". I heavily doubt it's taking these many years to roll out these changes; they probably simply don't care anymore, and never update these app store release notes.
There's always a chance of failure even for the most meticulous companies.
Now I'm not defending or excusing the company, but a singular event like this can happen to anyone and nothing is 100%.
If thorough investigation revealed poor quality control investment compared to what would be appropriate for a company like this, then we can say for sure.
Nobody ran this update
The update was pushed globally to all computers
With that alone we know they have failed the simplest of quality control methods for a piece of software as widespread as theirs. This is even excluding that there should have been some kind of error handling to allow the computer to boot if they did push bad code.
It's clear to me that CrowdStrike saw this as a data update vs. a code update, and that they had much more stringent QA procedures for code updates that they did data updates. It's very easy for organizations to lull themselves into this false sense of security when they make these kinds of delineations (sometimes even subconsciously at first), and then over time they lose site of the fact that a bad data update can be just as catastrophic as a bad code update. I've seen shades of this issue elsewhere many times.
So all that said, I think your point is valid. I know Crowdstrike had the posture that they wanted to get vulnerability files deployed globally as fast as possible upon a new threat detection in order to protect their clients, but it wouldn't have been that hard to build in some simple checks in their build process (first deploy to a test bed, then deploy globally) even if they felt a slower staged rollout would have left too many of their clients unprotected for too long.
Hindsight is always 20/20, but I think the most important lesson is that this code vs data dichotomy can be dangerous if the implications are not fully understood.
Obviously the system wasn't very robust, as a simple, within specs change could break it. A company like CrowdStrike, which routinely deals with memory exploits and claims to do "zero trust" should know better.
As often, there is a good chance it is an organization problem. The team in charge of the parsing expected that the team in charge of the data did their tests and made sure the files weren't broken, while on the other side, they expected the parser to be robust and at worst, a quick rollback could fix the problem. This may indeed be the sign of a broken company culture, which would give some credit to the ex-employees.
From my limited understanding, the file was corrupted in some way. Lots of NULL bytes, something like that.
https://www.crowdstrike.com/wp-content/uploads/2024/08/Chann...
It cannot have been a surprise to Crowdstrike that pushing bad data had the potential to bork the target computer. So if they had such an attitude that would indicate striking incompetence. So perhaps you are right.
> Hindsight is always 20/20, but I think the most important lesson is that this code vs data dichotomy can be dangerous if the implications are not fully understood.
But it's not some new condition that the industry hasn't already been dealing with for many many decades (i.e. code vs config vs data vs any other type of change to system, etc.).
There are known strategies to reduce the risk.
When you have the trifecta of regex, *argv packing and uninitialized memory you're reaching levels of incompetence which require being actively malicious and not just stupid.
They had previous bricked linux hosts earlier with a similar type of update.
So we also know that they don't learn from their mistakes.
This is the point I would emphasize. A kernel module that parses configuration files must defend itself against a failed parse.
We don't really need that thorough of an investigation. They had no staged deploys when servicing millions of machines. That alone is enough to say they're not running the company correctly.
I’d consider staggering a rollout to be the absolute basics of due diligence.
Especially when you’re building a critical part of millions of customer machines.
Before the incident, if you asked a customer if they would like to get updates faster even if it means that there is a remote chance of a problem with them... I bet they'd still want to get updates faster.
Specifically because this is about fighting against malicious actors, where time can be of essence to deploy some sort of protection against a novel threat.
If there's deadlines that you can go over, and nothing bad happens, for sure. Always have canary releases, and perfect QA, monitoring everything thoroughly, but I'm just saying, there can be cases where damage that could be done if you don't act fast enough, is just so much worse.
And I don't know that it wasn't the case for them. I just don't know.
This is severely overstating the problem: an extra few minutes is not going to be the difference between their customers being compromised. Most of the devices they run on are never compromised, because anyone remotely serious has defense in depth.
If it was true, or even close to true, that would make the criticism more rather than less strong. If time is of the essence, you invest in things like reviewing test coverage (their most glaring lapse), fuzz testing, and common reliability engineering techniques like having the system roll back to the last known good configuration after it’s failed to load. We think of progressive rollouts as common now but they got to get that mainstream in large part because the Google Chrome team realized rapid updates are important but then asked what they needed to do to make them safe. CrowdStrike’s report suggests that they wanted rapid but weren’t willing to invest in the implementation because that isn’t a customer-visible feature – until it very painfully became one.
Guess which part took down much of the corporate world?
from Preliminary Post Incident Review at https://www.crowdstrike.com/falcon-content-update-remediatio... :
"CrowdStrike delivers security content configuration updates to our sensors in two ways: Sensor Content that is shipped with our sensor directly, and Rapid Response Content that is designed to respond to the changing threat landscape at operational speed.
...
The sensor release process begins with automated testing, both prior to and after merging into our code base. This includes unit testing, integration testing, performance testing and stress testing. This culminates in a staged sensor rollout process that starts with dogfooding internally at CrowdStrike, followed by early adopters. It is then made generally available to customers. Customers then have the option of selecting which parts of their fleet should install the latest sensor release (‘N’), or one version older (‘N-1’) or two versions older (‘N-2’) through Sensor Update Policies.
The event of Friday, July 19, 2024 was not triggered by Sensor Content, which is only delivered with the release of an updated Falcon sensor. Customers have complete control over the deployment of the sensor — which includes Sensor Content and Template Types.
...
Rapid Response Content is used to perform a variety of behavioral pattern-matching operations on the sensor using a highly optimized engine.
Newly released Template Types are stress tested across many aspects, such as resource utilization, system performance impact and event volume. For each Template Type, a specific Template Instance is used to stress test the Template Type by matching against any possible value of the associated data fields to identify adverse system interactions.
Template Instances are created and configured through the use of the Content Configuration System, which includes the Content Validator that performs validation checks on the content before it is published.
On July 19, 2024, two additional IPC Template Instances were deployed. Due to a bug in the Content Validator, one of the two Template Instances passed validation despite containing problematic content data.
Based on the testing performed before the initial deployment of the Template Type (on March 05, 2024), trust in the checks performed in the Content Validator, and previous successful IPC Template Instance deployments, these instances were deployed into production."
When there's 0day, how enterprisey you would like to catch the 0day?
A patch should, at minimum:
1. Let the app run 2a. Block the offending behaviour 2b. Allow normal behaviour
Part 1. can be assumed if Parts 2a and 2b work correctly.
We know CrowdStrike didn't ensure 2a or 2b since the app caused the machine to reboot when the patch caused a fault in the app.
CrowdStrike's Root Cause Analysis, https://www.crowdstrike.com/wp-content/uploads/2024/08/Chann..., lists what they're going to do:
====
Mitigation: Validate the number of input fields in the Template Type at sensor compile time
Mitigation: Add runtime input array bounds checks to the Content Interpreter for Rapid Response Content in Channel File 291 - An additional check that the size of the input array matches the number of inputs expected by the Rapid Response Content was added at the same time. - We have completed fuzz testing of the Channel 291 Template Type and are expanding it to additional Rapid Response Content handlers in the sensor.
Mitigation: Correct the number of inputs provided by the IPC Template Type
Mitigation: Increase test coverage during Template Type development
Mitigation: Create additional checks in the Content Validator
Mitigation: Prevent the creation of problematic Channel 291 files
Mitigation: Update Content Configuration System test procedures
Mitigation: The Content Configuration System has been updated with additional deployment layers and acceptance checks
Mitigation: Provide customer control over the deployment of Rapid Response Content updates
====
/s
I thought the new code was actually installed, the running part depends on the script input...?
However, every company yes-man is paid to be a yes-man and will speak in favor of the company without exception - that literally is the job. Otherwise they will be fired and will join the ranks of the aforementioned people.
So logically it makes more sense for me to believe the former more than the latter. The two-sides are not equivalent (as you may have alluded) in term of trustworthiness.
Mostly disappointed.
That makes it easier to trust one side over another.
If the employees fucked up then you’ll say the company still fucked up because it wasn’t managing the employees well.
And then in that situation you’ll still believe the lying employees who say its the company’s fault while leaving out their culpability.
No, what we have is a publication who is claiming that the people they talked to were credible and had points that were interesting and tended to match one another and/or other evidence.
You can make the claim that Semafor is bad at their jobs, or even that they're malicious. But that's a hard claim to make given that in the paragraph you've quoted they are giving you the contrary evidence that they found.
And this is a process many of us have done informally. When we talk to one ex-employee of a company, well maybe it was just that guy, or just where he was in the company. But when a bunch of people have the same complaint, it's worth taking it much more seriously.
Yes, I'm more likely to leave reviews if I'm unsatisfied. Yes, people are more likely to leave CS if they were unhappy. Biased data, but still useful data.
Software security and quality is the responsibility of everyone on the team. A good UX designer should be thinking of ways a user can escape the typical flow or operate in unintended ways and express that to testers. And in decisions where management is forcing untested patches everyone should chime in.
Granted, a company that isn’t focused on the user experience as much as it is on other things might not prioritise this as much in the first place.
I bet my ass UX is not anywhere close to the low-level OS team.
UX is definitely embedded in the App level team but not in low-level.
Ex-employees said bugs caused the log monitor to drop entries. Crowdstrike responded the project was never designed to alert in real time. But Crowdstrike's website currently advertises it as working in real time.
Ex-employees said people trained to monitor laptops were assigned to monitor AWS accounts with no extra training. Crowdstrike replied that "there were no experienced ‘cloud threat hunters’ to be had" in 2022 and that optional training was available to the employees.
Is anyone really surprised or learned any new information? For us that have worked for tech companies, this is one of those repeating complaints that you hear across orgs that indicates a less than stellar engineering culture.
I've worked with numerous F500 orgs and I would say 3/5 orgs that I worked in, their code was so bad that it made me wonder how they haven't had a major incident yet.
Source: me, a developer who also codes in free time and notices how bad fs perf is especially.
I've had the CrowdStrike sensor, and my current company is using cyberhaven.
So.. while 2 data points don't technically make a pattern, it does begin to raise suspicion.
To you and me, maybe. To the insurers and airlines paying out over the problem, maybe not.
I would expect any contractor that may have worked for CrowdStrike, or done something like a third-party audit, would be under an NDA covering their work.
Who's left to speak out with any meaningful details?
Except the fact that CrowdStrike fucked up the one thing they weren't supposed to fuck up.
So yeah, at this point I'm taking the ex-employees' word, because it confirms the results that we already know -- there is no way that update could have gone out had there been proper "safety first" protocols in place and CrowdStrike was "meticulous".
Employees don't typically have much reputation to ruin. I am perfectly content putting this on their shoulders.
My guess is that people are identifying with sentence said just before: "Speed [of shipping] is everything." Aka "Move fast and break things."
The culture described by this article must mirror many of our lived experiences. The pure pleasure of shipping code, putting out fires, making an impact (positive or negative)... and then leaving it to the next engineers & managers to sort out, ignoring the mess until it explodes. Even when it does, no one gets blamed for the outage and soon everyone goes back to building features that get them promoted, regardless of quality.
Through that ZIRP light, these process failures must look like a feature, not a bug. The emphasis on "quality" must also look like annoying roadblocks in the way of having fun on the customer's dime.
Not very long ago we had this client who ordered a custom high security solution (using a kernel driver). I can't reveal too much but basically they had this offline computer running this critical database and they needed a way to account for every single system call to guarantee that any data could have not been changed without the security system alerting and logging the exact change. No backups etc were allowed to leave the computer ever. We were even required to check ntdll (this was on Windows) for hooks before installing the driver on-site & other safety precautions. Exceptions, freezes or a deadlock? No way. Any system call missed = disaster.
We took this seriously. Whenever we made a change to the driver code we had to re-test the driver on 7 different computers (in-office) running completely different hardware doing a set test procedure. Last test before release entailed an even more extensive test procedure.
This may sound harsh but CrowdStrike are total amateurs, always been. Besides, what have they contributed to the cyber security community? - Nothing! Their research are at a level of a junior cyber security researcher. They are willing to outright lie and jump to wild conclusions which is very frowned upon in the community. Also heard others comment on how CS really doesn't really fit the mold of a standard cyber security company.
Nah, CS should take a close look at true professional companies like Kaspersky and Checkpoint; industry leaders who've created proven top notch security solutions (software/services) but not least actually contributed their valuable research to the community for free, catching zero-days, reporting them before no one even had a chance of exploiting them.
They deserve some criticism.
I once worked with a senior engineer who loved running incidents. He felt it was real engineering. He loved debugging thorny problems on a strict timeline, getting every engineer in a room and ordering them about, while also communicating widely to the company. Then, there's the rush of the all-clear and the kudos from stakeholders.
Specific to his situation, I think he enjoyed the inflated ownership that the sudden urgency demanded. The system we owned was largely taken for granted by the org; a dead-end for a career. Calling incidents was a good way to get visibility at low-cost, i.e., no one would follow-up on our postmortem action items.
It eventually became a problem, though, when the system we owned was essentially put into maintenance mode, aka zero development velocity. Then I estimate (balancing for other variables) the rate the senior engineer called an incident for not-incidents went up by 3x...
To be fair to the senior, it was a bad situation made worse by everyone's self-interest. Myself included.
What I've learned is that fixing things for these people (and even having entire teams fixing things for weeks) just leads to a continued lax attitude to testing, and leaving the fallout for others to deal with. To them, it all worked out in the end, and they get kudos for rapidly getting a solution in place.
I'm done fixing their work. I'd rather work on my own tasks than fix all the problems with theirs. I'm strongly considering moving on, as this has become an entrenched pattern.
My favorite repeated reorg strategy over the years is “that we will train everyone in engineering to be hot swappable in their domains”. Talk about spinning wheels.
For example, you might think "if a big security exploit happens, the stock price might tank". So if they value the stock price, they'll focus on security, right?. In reality what they do is focus on burying the evidence of security exploits. Because if nobody finds out, the stock price won't tank. Much easier than doing the work of actually securing things. And apparently it's often legal.
And when it's not a bizarre incentive, often people just ignore risks, or even low-level failures, until it's too late. Four-way intersections can pile up accidents for years until a school bus full of kids gets T-boned by a dump truck. We can't expect people to do the right thing even if they notice a problem. Something has to force the right thing.
The only thing I have ever seen force an executive to do the right thing is a law that says they will be held liable if they don't. That's still not a guarantee it will actually happen correctly, course. But they will put pressure on their underlings to at least try to make it happen.
On top of that, I would have standards that they are required to follow, the way building codes specify the standard tolerances, sizes, engineering diagrams, etc that need to be followed and inspected before someone is allowed into the building. This would enforce the quality control (and someone impartial to check it) that was lacking recently.
This will have similar results as building codes - increased bureaucracy, cost, complexity, time... but also, more safety. I think for critical things, we really do need it. Industrial controls, like those used for water, power (nuclear...), gas, etc, need it. Tanker and container ships, trains/subways, airlines, elevators, fire suppressants, military/defense, etc. The few, but very, very important, systems.
If somebody else has better ideas, believe me, I am happy to hear them....
Would you pay 10x (or more, even) for these systems? That means 10x the price of water, utilities, transport etc, which then accumulate up the chain to make other things which don't have criticality but do depend on the ones that do.
The thing is, what exists today exists because it's the path of least resistence.
At some point - maybe it already happened, I don't know - spending more on preventive measures and maintenance will be the path of least resistance.
if it's critical to your business, then yes; but you quickly find out whether or not it's actually critical to your business or whether it's something you can do without
In general, the biggest problem I see with late stage capitalism, and a lack of accountability in general, is that given the right incentives people will “fuck things up” faster than you can stop them. For example, say CrowdStrike was skirting QA - what’s my incentive as an individual employee versus the incentive of an executive at the company? If the exec can’t tell the difference between good QA and bad QA, but can visually see the accounting numbers go up when QA is underfunded, he’s going to optimize for stock price. And as an IC there’s not much you can do unless you’re willing to fight the good fight day in and day out. But when management repeatedly communicates they do not reward that behavior, and indeed may not care at all about software quality over a 5 year time horizon, what do you do? The key lies in finding ways to convince executives or short of that holding them to account like you say.
"Proper" engineering disciplines have similar systems like the Professional Engineer cert via the NSPE that requires designs be signed off. If you had the requirement that all software engineers (now with the certification actually bestowing them the proper title of 'engineer') sign off on their design, you could prevent the company from just finding someone else more unscrupulous to push that update or whatever through. If the entirety of the department or company is employing properly certificated people, they'd be stuck actually doing it the right way.
That's their incentive to do it correctly: sign your name to it, or lose your license, and just for drama's sake, don't collect $200, directly to jail. For the companies, employ properly licensed engineers, or risk unlimited downside liability when shit goes sideways, similar to what might happen if an engineering firm built a shoddy bridge.
Would a firm that peddles some sort of CRUD app need to go through all of this? If it handles toxic data like payments or health data or other PII, sure. Otherwise, probably not, just like you have small contracting outfits that build garden sheds or whatever being a bit different than those that maintain, say, cooling systems for nuclear plants. Perhaps a law might be written to include companies that work in certain industries or business lines to compel them to do this.
Even if it wasn’t wrong, that’s still the wrong reaction. We’re in this situation because so many companies were negligent in the past and the status quo was obviously untenable. If there is a problem with a given standard the solution is to make a better system (e.g. like Apple did) rather than to say one of the most important industries in the world can’t be improved because that’d require a small fraction of its budget.
Edit: Disclaimer: The quotes aren't mine, just retorts I've received from others when I suggest the R-word.
Yeah it makes no sense. Was the US not losing to China when we own-goaled the biggest cybersecurity incident in history?
Such a silly meme, too. Economics 101 China and USA would both benefit by halting the conflict and trading with each other
Software is pretty much always made cheaply and quickly. Even NASA will have b software blunders and have rockets explode mid flight.
They recently had a stroke at home just days after spending over a month in the hospital.
Then I remembered that they were originally supposed to be getting an important surgery, but it was delayed because of the CrowdStrike outage. It took weeks for the stars to align again and the surgery to happen.
It makes me wonder what the outcome would have been if they had gotten the surgery done that day, and not spent those extra weeks in the hospital with their condition and stressing about their future?
To illustrate:
If I were to do something horrible like kick a 3 year olds knee out and cripple them for life, I would be rightly labeled a monster.
But If I were to say... advocate for education reform to push American Sign Language out of schools, so that deaf children grow up without a developmental language? We don't have words for that, and if we did, none of them would get near the cumulative scope and harm of that act.
We simply do not address distributed harms correctly. And a big part of it is that we don't, we can't, see all the tangible harms it causes.
But as other posts on HN have discussed, anecdotes, especially your own, hit differently.
It makes me thankful the software I work on isn't involved in life and death situations... But then again, it causes me to better consider the things my work could be responsible for (banking). Rushed work that causes a loan application to fail or transaction to be held unnecessarily shouldn't kill someone outright, but there can be real consequences that affect real people just like Rita.
Their 'expert' on engineering process is a senior UX designer? Somehow, I doubt they were very close to the kernel patch deployment process.
In other words, it confirms our biases and we're willing to accept it at face value despite there being only a single anecdotal piece of evidence.
That’s why I said it was compatible: both these former employees and their own report showed an emphasis on shipping rapidly but not the willingness to invest serious money in the safeguards needed to do so safely. If you want to construct another theory, feel free to do so.
I bet my ass anyone working in low-level code don't ship the way you do in Cloud.
Their technical report says otherwise – and we know they didn’t adopt the common cloud practices of doing real testing before shipping or having a progressive deployment.
Much of the discourse around this topic has described ideal testing and deployment practise. Maybe it's different in Silicon Valley or investment banks, but for the sorts of companies I work for (telco mostly) things are very far from that ideal.
My view of he industry is one of shocking technical ineptitude from all but a minority of very competent people who actually keep things running... Of management who prioritize short term cost reduction over quality at every opportunity, leading to appalling technical debt and demoralized, over-worked staff who rapidly stop giving a damn about quality, because speaking out about quality problems is penalized.
I expect some ex-employees to be disgruntled and present things in a way that makes CroudStrike look bad. That happens with every company.
BUT, CrowdStrike has ZERO credibility at this point. I don't believe a word they say.
have never heard that word used is a non-negative way
A similar use of dis- as an intensifier apparently happened in "dismayed" (here from an Old French verb, esmaier , which meant to trouble, to disturb), and in "disturbed" (from Latin a word, turba, meaning turmoil). I haven't heard any one say they are "mayed" or "turbed", but people would probably see the same as "gruntled" if you used them.
https://www.ling.upenn.edu/~beatrice/humor/how-i-met-my-wife...
"$10 million a minute.
That’s about how much the trading problem that set off turmoil on the stock market on Wednesday morning is already costing the trading firm.
The Knight Capital Group announced on Thursday that it lost $440 million when it sold all the stocks it accidentally bought Wednesday morning because a computer glitch. "
Glitch. Oh...
Ihey were unwilling to push that button in the short time they had. If you read the reports to the SEC or the articles about it, you will note that. The follow-ups recommended that all firms adopt a big red button that is less catastrophic.
You would think so.
Cynical me.
But no. When money is at stake much more care is taken than when lives are at stake.
The TL;DR of Knight is that Knight had several things go wrong at the same time, and had no circuit breaker for the problem that did not stop trading for the whole firm for the day. Most trading firms have had things go badly, but the holes in the Swiss cheese aligned for Knight (and they were larger than many other firms). This all comes from a sort of culture of carelessness.
Your usage, in assigning blame rather than diffusing it, was novel to me.
It was really an uphill battle to convince everyone not to use Crowdstrike. Eventually I managed to but after many meetings where I had to spend a significant amount of time convincing different shareholders. I'm sure a lot of people just fold and go with them.
I've recently changed jobs and the new employer, a large company, obviously has to have an IT compliance / security update policy because everyone else has it so if they stand out from the crowd and don't do it and somehow get hacked, it's 100x worse than constantly annoying employees and top of the line computers working like a 1970s terminal.
It's rarely that a week passes without the obligatory update + restart. And at least once a month they update THE FUCKING BIOS! What the fuck can be so broken in those laptops that the BIOS is a constant security hazard?! And why would you buy software from someone who week after week after week tell you all you had so far was a hazardous piece of shit that cannot possibly function without constant pampering?
Ahh and of course they botch it. Had to have the OS completely wiped out and reinstalled after the laptop started to behave more and more erratically, 100% caused by faulty updates on top of faulty patches trying to patch the faulty updates. Worked OK for a while afterwards then updates started piling up and so far I only lost use of the web camera (before it was Wifi then display adapter).
There's literally no words how much I hate "the system" and the constant security update take it up the ass we're forced to put up with.
“It was hard to get people to do sufficient testing sometimes,” said Preston
Sego, who worked at CrowdStrike from 2019 to 2023. His job was to review the
tests completed by user experience developers that alerted engineers to bugs
before proposed coding changes were released to customers. Sego said he was
fired in February 2023 as an “insider threat” after he criticized the
company’s return to-work policy on an internal Slack channel.
Okay clearly that company has a culture issue. Imagine criticizing a policy and then getting labeled "insider threat".Reviewing tests is part of PR review.
--- and before anyone asks, this is my statement on CrowdStrike calling everyone disgruntled:
"I'm not disgruntled.
But as a shareholder (and probably more primarily, someone who cares about coworkers), I am disappointed.
For the most part, I'm still mourning the loss of working with the UX/Platform team."
At the same time, i feel like big profit-chasing software companies are all like how CrowdStrike is.
Many may be in the same type of company, but situations have not arisen that reveal how leadership really feels about employees.
Especially because that’s incredibly dumb. A true insider threat would play nice while you find all your confidential data leaking.
I know you're just quoting the phrase, but what a gross and dishonest way of phrasing "return to office". Implies working remotely doesn't count as work. Smacks of PR. Yuck.
Everyone must realize that crowdstrike has a captive audience with no alternatives that can meet corporate compliance.
On the plus side this should spur some disruptors into gear, assuming VCs are willing to pivot from wasting money funding LLM wrappers.
If it runs up a huge amount in the first half of the year and then the incident knocks off 30% of their market, that still means the incident was really bad.
C-Suite and investors don't seem to want to spend on quality. They should just price in that their stock investment could collapse any day.
Many competing platforms that can be a drop in placement for ClownStrike.
At some point you go past questions of laziness or discipline and it becomes a neurosis. Like an addiction.
This was the first place I saw standups. [Edit: this was the 1990s] They were run by and for the "meat", the people running the line. "Level 2" only got to speak if we were blocked, or to briefly describe any new investigations we would be undertaking.
Weirdly (maybe?) they didn't drug test. I thought of all the places I've worked, they would. But they didn't. They were firmly committed to the "no SPOFs" doctrine and had a "tap out" policy: if anyone felt you were distracted, they could "tap you out" for the day. It was no fault. I was there for six months and three or four times I was tapped out and (after the first time, because they asked what I did with my time off the first time) told to "go climb a rock". I tapped somebody out once, for what later gossip suggested was a family issue.
Sure most of the times I was tapped out I was distracted by personal thoughts. But one time I was just thinking about the problem. I protested "but I was thinking about the problem!" and they said "go think somewhere else!".
> every change needed to have an automated rollback process
How did you accomplish that?
Corps won't even put resource into anti-fraud efforts if they believe the millions being stolen from their bottom line isn't worth the effort. I have seen this attitude working in FAANGS.
None of this will change until tech workers stop being sadists and actually unionize.
The 3rd component of the CIA triad is often overlooked, yet the availability is what makes the protected asset—and, transitively, the protection itself—useful at the first place.
The disruption is effectively a Denial of Service.
Don’t find this particularly interesting news.
Business: Nothing - Windows Defender Advanced Threat Protection is built into the higher Microsoft 365 license tiers.
It amazes me people chose to pay money to have all their PCs bluescreen.
macOS: https://learn.microsoft.com/en-us/defender-endpoint/microsof...
It does iOS and Android too.
Again, if you're an organisation big enough to care about single-pane-of-glass-monitoring you probably already have access to this via the Microsoft 365 license tier you're on.
In house competence
/s
Its almost like there is a lesson for executives here. hmmmm
Reliability is a critical facet of security from a business continuity standpoint. Any business still using crowdstrike is out of their mind.
Anyone with access to your CS SIEM can search for GitHub, aws, etc creds. Anything your devs, ops and sec teams use on their Macs.
Only the Mac version does this. There is no way to disable this behaviour or a way to redact things.
Another really odd design decision. They probably have many many thousands of plain text secrets from their customers stored in their SIEM.
https://en.wikipedia.org/wiki/Security_information_and_event...
I guess it's pretty much the same principle.
On the other hand there was e.g. WannaCry in 2017 where 200,000 systems across 150 countries running Windows XP and other unsupported Windows Server versions had crypto miners installed. It shows that companies world-wide had trouble properly maintaining the life cycle of their systems. I think it's too easy to only accuse security vendors of quality problems.
I'm sure this is going to raise red-flags in my IT department.
Again, the plaintext is the problem.
These environment variables get loaded from the command line, scripts, etc. - CrowdStrike and all of the best EDRs also collect and send home all of that, but probably in an encrypted stream?
The plaintext part is not okay.
Another way to look at it is, the CS cloud environment is effectively part of your environment. the secrets can get scrubbed, but CS still has access to your devices, they can remotely access them and get those secrets at any time without your knowledge. that is the product. The security boundary of OP's mac is inclusive of the CS cloud.
In reality, you can take steps to prevent PII from being logged by Crowdstrike, but credentials are too non-standard to meaningfully scrub. It would be an exercise in futility. If you trust them to have unrestricted access to the credential, the fact that they're inadvertently logging it because of the way your applications work should not be considered an increase in risk.
If you want this redacted, it is a SIEM functionality not Crowdstrike's. Depends on the SIEM but even older generation SIEMs have a data scrubbing feature.
This isn't a Crowdstrike design decision as you've put it. any endpoint monitoring too, including the free and open source ones behave just as you described. You won't just see env vars from macs but things like domain admin creds and PKI root signing private keys. If you give someone access to an EDR, or they are incident responders with SIEM access, you've trusted them with full -- yet, auditable and monitored -- access to that deployment.
> Found out that the CrowdStrike Mac agent (Falcon) sends all your secrets from environment variables to their cloud hosted SIEM. In plain text.
So this person claims the data isn't send using TLS. I didn't verify the claim myself though.
Kidding, mostly, but wow that's a hell of a vulnerability.
One of the most famous examples can be seen in the NSA slide at the top of this article:
https://www.washingtonpost.com/world/national-security/nsa-i...
So there's this thing called "Threat model" and it includes some assumptions about some moving parts of the infra, and it very often includes assertion that a particular environment (like IDS log, signing infra surrounding HSM etc.) is "secure" (they mean outside of the scope of that particular threat model). So it often gets papered over, and it takes some reflex to say "hey, how we will secure that other part". There needs to be some conciousnes about it, because it's not part of this model under discussuon, so not part of the agenda of this meeting...
And it gets lost.
That's how shit happens in compliance-oriented security.
Realistically, secrets alone shouldn’t allow an attacker access - they should need access to infrastructure or a certificates in machines as well. But unfortunately that’s not the case for many SaaS vendors.
It's totally insane to send them to a remote service controlled by another organization.
1) employees are trusted with secrets, so we have to audit that employees are treating those secrets securely (via tracking, monitoring, etc)
2) we don’t allow employees to have access to secrets whatsoever, therefore we don’t need any auditing or monitoring
IMHO needing to be monitored constantly is not being "trusted" by any sense of the word.
Similarly employees can be trusted enough with access to prod, while the company wants to protect itself from someone getting phished or from running the wrong "curl | bash" command, so the company doesn't get pwned.
Factually, it is necessary for auditing and absolutely correlates with the extreme of needing to monitor the “usage” of “secrets”.
In a highly auditable/“secure” environment, you can’t give secrets to employees with no tracking of when the secrets are used.
Sending env vars of all your employees to one place doesn't improve anything. In fact, one can argue the company is now more vulnerable.
It feels like a decision made by a clueless school principle, instead of a security expert.
If you lean in the direction of keylogging all your employees, that's not only lazy but ineffective on account of the unnecessary noise collected, and it's counterproductive in that it creates a juicy central target that you can hardly trust anyone with. Good auditing is minimally useful to an adversary, IMO.
This does not seem to require regularly exporting secrets form the employee's machines though. Which is the main complaint I am reading. You would log when the secret is used to access something, presumably remote to the users machine.
Repeat after me: Security is not a bolt on tool.
Yeah. So you track them when they are used (which also gives you a nice timestamp). Not when they’re just sitting in the env.
It works the same way for biometrics like face unlock on mobile phones
Right, but doesn't that mean there is no risk from sending employee laptop ENV variables, since they shouldn't have any secrets on their laptops?
If you’re a PCI company then ending up with a credit card number in your SIEM can be a massive disaster. Because you’re never allowed to store that in plaintext, and your SIEM data is supposed to be immutable. In theory that puts you out of compliance for a minimum of one year with no way to fix it, in reality your QSAs will spend some time debating what to do about it and then require you to figure out some way to delete it, which might be incredibly onerous. But I have no idea what they’d do if your SIEM somehow became full of credit card numbers, that probably is unfixable…
You'd get rid of it.
They certainly wouldn’t let you keep it there, but if your SIEM was absolutely full of cardholder data, I imagine they’d require you to extract ALL of it, redact the cardholder data, and the import it to a new instance, nuking the old one. But for a QSA to sign off on that they’d be expecting to see a lot of evidence that removing the cardholder data was the only thing you changed.
This isn't realistic, it's idealistic. In the real world secrets are enough to grant access, and even if they weren't, exposing one half of the equation in clear text by design is still really bad for security.
Two factor auth with one factor known to be compromised is actually only one factor. The same applies here.
I think some configurability would be great. I would like to provide an allow list or the ability to redact. Or exclude specific host groups.
We all have different levels of acceptable risk
My mental model was that Apple provides backdoor decryption keys to China in advance for devices sold in China/Chinese iCloud accounts, but that they cannot/will not bypass device encryption for China for devices sold outside of the country/foreign iCloud accounts.
There's a triggering action that caused the env vars to be used by another ... ehem... Process ... that any EDR software in this beautiful planet would have tracked.
Why?
There is no need to send your environment variables.
Why are they only doing it for macs then?
Arbitrarily high levels of market penetration by sloppy vendors in high-stakes activities, far from being an argument for functioning markets, demand regulation.
Arbitrarily high profile failures of the previous two, far from indicating a tolerable norm, demand criminal prosecution.
It is recently that this seemingly ubiquitous vendor, with zero-day access to a critical kernel space that any red team adversary would kill for, said “lgtm shipit” instead of running a test suite with consequences and costs (depending on who you listen to) ranging from billions in lost treasure to loss of innocent life.
We know who fucked up, have an idea of how much corrupt-ass market failure crony capitalism could admit such a thing.
The only thing we don’t know is how much worse it would have to be before anyone involved suffers any consequences.
This silicon valley libertarian non sense needs to stop.