It's not just CrowdStrike – the cyber sector is vulnerable
ft.com
ft.com
Devs: Ugh...why?
Security: For safety!
Devs: Fine, we won't argue. Deploy it if you may.
A few moments later...
Devs: All of our VM's are slow as crap! Defender is using 100% of the CPU!
Security: Add another core to your VM's. ticket closed
Management: Why are our developers up 30% on their cloud spend!?
directly incentivised to make shit software
Apple had full control over the whole phone's software stack, in a very good way, meaning they built a good mobile OS that had good systems for power management and an app lifecycle that could actually kill apps at will to maintain efficiency, without disrupting the user.
With this, they decided to ship smaller batteries so they could make slimmer phones.
Except, they used garbage batteries. They were so small (1600mAh on the iPhone 6) that normal wear and tear of a few years degraded them to the point that the battery chemistry could not keep up with normal processor frequency and power ramping.
Apple started getting a lot of complaints because people were understandably upset that their 2-3 year old phone couldn't run for more than an hour off the charger. Apple didn't like increasing support load, even though they weren't covering anyone's battery replacement. Instead of putting out a press release that they had shipped sub-standard batteries in their phones, and offering free battery replacements with a new battery that wouldn't have the same problem in another 2-3 years, they included code in the new version of iOS to SIGNIFICANTLY slow down your 3 year old or less phone.
Apple made a product that deteriorated way too quickly, and then tried to hide it. That's batterygate. If LG sold a fridge that would die after five years because of compressor fatigue and then silently updated their fridges to not operate colder than 45 degrees F to extend the life of the compressor, I would hope you would be pissed at that, right?
A reminder that the iPhone 6 was also "Bendgate", which internal apple memos showed they knew was a serious problem before they sold it, and then claimed two years after release they only had 9 complaints of phone bending and that it wouldn't bend in normal use.
Apple sycophants are willing to put up with any bullshit from Apple. It is very tiresome to argue against blind faith.
Any other company would face incredible scrutiny if that happened. Imagine if MS did that to their surface devices. And this level of scrutiny from consumers is healthy.
And it's not like other vendors are not full of crap either. I had a Dell laptop with a clearly broken display that they never acknowledged or repaired - and many other problems of all kinds. Apple was always least (but obviously not zero) problems and best build quality.
Pff. I had a macbook battery (2018 model, brand new, issued by my employer at the time) that died in 1.5 years. Died in the sense that I couldn't use that crap unplugged for more than 10 minutes.
Since every place I work issues me a MacBook, I am very experienced with these luxury toys, and I wouldn't ever buy one for myself. I actually think Thinkpads are much better.
BTW you're saying it was 2018 model, and employer issued, so if I'm correct in assuming it was a top model Intel CPU, these really were chewing through the batteries because of the heat. It's very different with i5, less powerful i7 and Apple Silicon.
I really don't think anyone is claiming that Apple is perfect - it's just that the experience with other vendors is so, so utterly bad. For example ThinkPads - nice performance and cheap, I give you that. But the non-existent customer care (for consumers, not enterprise), the build quality, the bad sound and displays and the absolutely terrible touchpad make me avoid it. Also Windows - and I never got Linux properly working on a ThinkPad as well as MacOS does on a MacBook, even though they claim it's Linux certified.
The whole cybersecurity concept of installing third party mystery meat in the kernel controllable over the internet by a different company seems contrary both to good security practices and software quality assurance, immutable production architecture and repeatable builds.
Actually, even worse than that was we had to install AV on all the images that the ephemeral map-reduce/Hadoop clusters we spun up, but the way the AV stuff worked, the computes were gone by the time the registration for the new compute had gone thru whole process. And, in AWS accounts where there were maybe 63 IPs and say 600 EC2s/day they used IP as the primary key for the "list of compute in the VPC." So they stitched together totally unrelated stuff as if it was the same continuous compute. I guess it would be eventually fixed, but the bad security data was not a real concern of the devops team that was building out stuff as rapidly as possible, nor were the EC2 made in a sealed off VPC and only lived for a few tens of minutes or hours at best a serious security concern to the actual security people. Just a check list solution hitting a novel environment.
The disconnect is that companies are both (1) the only entity in control of their system and how it is tested and (2) not liable if a security breach does happen.
I believe we need to enable red teams (security researchers) to test the security of any system, with or without permission, so long as they report responsibly and avoid obviously destructive behavior such as sustained DDoS attacks.
A branch of the government, possibly of the military (the Space Force?) could constantly be trying to hack the most important systems in our nation (individuals and private companies too). The bad guys are doing this anyway, but hopefully the good guys could find the security holes first and report them responsibly.
Again, currently this doesn't happen because it would be embarrassing and inconvenient for powerful companies. We threaten researchers who do nothing more than press F12 (view HTML source) with jail time and then have our best surprised Pikachu faces ready for when half the nations data is stolen every week or major systems go down. Actually, we don't make faces at all, half the nation's data is stolen every week--no, actually we don't even take notice, we just accept it as the way things have to be. Because, after all, we can't expect companies to be liable, but we can trust companies to have exclusive control over the testing of their security. How convenient for them.
But if it wasn't a bug found in the wild, can you imagine the fights between the NSA red and blue teams on whether to alert Microsoft about it?
Of course they also crippled the key length to 56 bits...
sentinelone, tanium, guardicore, defender endpoint, delina
all running as root (or worse), sucking up absurd amounts of resources, often more than the software running on the machine (but advertised as "LOW IMPACT")
they also cause reliable software to break due to bugs in e.g. their EBPF
also often serialises all network and disk on the machine through to one single thread (so much for multi-queue NVMe/NICs)
the risk and compliance attitude that results in this corporate mandated malware being required needs to go
this software creates more risk than it prevents
(Just playing devils advocate. I hate Crowdstrike as much as anyone here :)
Page 53: “The evaluator will conduct penetration testing, based on the identified potential vulnerabilities, to determine that the OS is resistant to attacks performed by an attacker possessing Basic attack potential.”
That is the lowest level of security certification outlined in the standard. The elementary school diploma of security.
To see what that means, here is a sample of the certification report [3].
Page 14: “The evaluator has performed a search of public sources to discover known vulnerabilities of the TOE.
Using the obtained results, the evaluator has performed a sampling approach to verify if exists applicable public exploits for any of the identified public vulnerabilities and verify whether the security updates published by the vendor are effective. The evaluator has ensured that for all the public vulnerabilities identified in vulnerability assessment report belonging to the period from June 8, 2021 to July 12, 2022, the vendor has published the corresponding update fixing the vulnerabilities.“
The "hardcore" certification process they subject themselves to is effectively doing a Google search for: “Windows vulnerabilities” and checking all the public ones have fixes. That is all the security they promise you in their headline, mandatory security certification that is the only general security certification listed and advertised on their official security page.
When a company puts their elementary school diploma on their resume for “highest education received”, you should listen.
That is not to say any of the names in general purpose operating systems such as MacOS, Linux, Android, etc. are meaningfully better. They are all inadequate for the task of protecting against moderately skilled commercially minded attackers. None of them have been able to achieve levels of certification that provide confidence against such attackers.
This is actually a good sign, because those systems are objectively and experimentally incapable of reaching that standard of security. That they have been unable to force a false-positive certification that incorrectly states they have reached that standard demonstrates the certification at least has a low false-positive rate.
All of the standard stuff is inadequate in much the same way that all known materials are inadequate for making a space elevator. None of it works, so if you do want to use it, you must assume they are deficient and work around it. That or you could use the actual high quality stuff.
[1] https://learn.microsoft.com/en-us/windows/security/security-...
[2] https://www.commoncriteriaportal.org/files/ppfiles/PP_OS_V4....
[3] https://download.microsoft.com/download/6/9/1/69101f35-1373-...
Note how the intrusion detection system here only needs to do offline scans that are unaffected by security updates.
And whether you can see it or not, they're all still some form of dumpster fire, be it security, usability, price.
I think it's more like, security is heavily check mark based. Crowdstrike and friends have managed to get "endpoint security"[1] added as a "standard security best practice" which every CSO knows they must follow or get labeled incompetent. Therefore "endpoint security" must be installed everywhere with no real proof that it makes things more secure, an arguable case that it makes things less secure, and an undeniable case that it makes things less reliable.
[1] I also never understood how "endpoints" somehow are defined as "any computer connected to any network." I tried to fight security against installing this crap on our database servers with the argument that they are not endpoints. Did not work.
People and companies that hide behind this bullshit don’t deserve to be in leadership positions. Cowards
[1] https://www.datacenterdynamics.com/en/news/cloudflare-claims...
allow nothing and then gradually allow some activities that are deemed safe
do not allow software to be installed from arbitrary locations
app sandboxing and third-party vendors cannot break their sandbox
basically, iOS, Android, ChromeOS
50% of the people impacted today probably only need a browser
keep'er running...
Do you have more info about this ? I am very interested. Is it impacting SAN fc storage ?
Making systems hard to hack and robust to rare events:
* is really hard,
* costs a lot of money, and
* reduces earnings in the short term.
Faced with these inconvenient facts, many executives who want to see stock prices go up prioritize... other things.
It's still not enough.
To each individual company, it’s better to have the big single point of failure. That’s the problem.
Am I bitter at losing the business decisions that push ease of management by sending control to service providers? Not really. It's been dozens of times, and I lose every time.
I can raise the concerns to make sure the decisions are educated ones, and then let the decisions be made.
The look they gave was priceless.
In practice, failing closed (or crashing) is probably fine for most businesses, and lower cost than a breach, but the correct solution is automated testing across a broad spectrum of devices, staged and rolling updates to prevent entire fleets going down at once, and ensuring that there is an effective, tested rollback mechanism.
But that shit's expensive, so shrug :/
This kind of logic only works if you ignore any kind of possible nuances in the problem and just insist on throwing the baby out with the bathwater. Just because someone let you do automatic updates (or let's be real, you probably didn't give them much of an option) that doesn't mean you should use it for everything.
Automatic update of data (like virus definitions) != automatic update of code (like kernel driver)
And really, the only time you could justify doing automatic updates on other people's machines is when have reason to believe the risk of waiting for the user to get around to it is larger than the damage you might do in the process... which doesn't seem to have been the case here.
Even if this is not what happened, it is possible, and shows the data/code update separation does not prevent problems.
Sure they do? This is like saying seatbelts don't prevent injuries because people still die even while wearing them.
I never said that one weird trick would solve every problem, or even this particular one for that matter. What I was saying was that if you look for ways to add nuance... you can find better solutions than if you throw the baby out with the bathwater. I just gave two examples of how you could do that in this problem space. That doesn't mean those are the only two things you can do, or that either would've single handedly solved this problem.
The problem in your scenario is that kernel mode behavior is being auto updated globally (via data or code is irrelevant), and that should require a damn high bar. You don't do it just because you can. There's got to be a lower bar for user mode updates than kernel, etc.
You still need to drive carefully. To me, it looks like these people relied on the safety of seatbelts and drove really fast, and there was, predictably, a horrible crash with a lot of damage.
Crowdstrike themselves seem to have missed the nuance.
Oh, I agree - automatic updates are nuanced in many cases. Generally speaking, automatic updates are a good thing, but they offer trade-offs; the main trade-off is rapidly receiving security updates, at the risk of encountering new features, which can include new bugs. This is kind of a big reason why folks who buy systems should be requiring that updates offer a distinction between Security/Long Term Support, and Feature updates. It allows the person who buys the product to make an effective decision about the level of risk they want to assume from those updates.
> Automatic update of data (like virus definitions) != automatic update of code (like kernel driver)
Yep, absolutely, except for the case where the virus definitions (or security checks) are written in a language that gets interpreted in a kernel driver, presumably in languages that don't necessarily have memory safety guarantees. It really depends on how the security technology implements it's checks, and the facilities that the operating system provides for instrumentation and monitoring.
This whole incident would have not happened if just a basic deployment test was conducted. It's so widespread, it would have been impossible to miss detecting the issue.
But updates should be rolled out slowly and you need enough telemetry to detect problems as it's rolled out. Reboots, crashes, cpu/memory use, end user reports, etc should all be used to detect issues and pause the rollout.
Canarying is by now not a very new practice. This is like doctors not washing their hands.
Take ZScaler, which is a service that proxies all network connections of a computer to a central cloud proxy server, mitms it (decrypt, inspect, log, and encrypt), and then forwards it to the target server. Imagine that this is hacked, and this isn't immediately discovered. Hackers listening in and being able to tap off cookies, bearer tokens and other confidential information for weeks. That would affect so many companies. And if they would want to cause a DoS, many computers and servers would be left without an operational internet connection.
Is it just me?
OEM software is usually the worst offender here, all these installers and support utilities should be fully abolished. Drivers should stop loading if they aren't coming from Windows Update and haven't passed some quality control.
As I say every time this happens (and it will keep happening for the next decade or so)... Ambient Authority systems can't be secured, we need to switch to Operating Systems designed around Capability Based Security.
We need at least 2 of them, from competing projects.
I do not invest in cybersecurity companies, it is very risky IMO
But some companies are using this as an excuse to not care about the chances of an exploit at all, and just write code in a cheaper way.
We need a middle ground, where there is at least a reasonable effort towards security.
Right, and this should be the single deciding factor for most system programming and core infrastructure development. One doesn't throw away 20+-year-old battle-tested code simply because it's grown ugly bug fixes for edge conditions no one wants to worry about. The idea that it's possible to throw away, say 30-year-old font rendering code and replace it without revisiting a lot of the problems along the way is peak hubris.
And the same goes for choosing and building internal IT systems, KISS should rule those choices because each layer adds additional code, additional updating, etc. Monolithic general-purpose software is not only a waste of resources (having software that 9/10th is just taking up disk/memory/cache space because only 10% of its features are used), but it's a maintenance and security nightmare.
This is the problem with much of the open-source world, too. Having 20 different Linux filesystem drivers or whatever is just adding code that will contain bugs, exploits, and a monthly kernel update containing 80 KLOC of changes is just asking for problems. Faster processes, updates, and development velocity in projects that were "solved" decades ago are just a playground for bad actors.
So, to go back to Andrew Tanenbaum and many others, no one in their right mind should be writing or using OSs and software that aren't built from first principles with clearly defined modularity and security boundaries. A disk driver update should be 100% separate and compatible with not just the latest OS kernel but ones from 10+ years ago. A database update shouldn't require the latest version of python "just because".
Most software is garbage quality written by a bunch of people who are all convinced they are better than their peers. And yet another code review, or CI loop, isn't going to solve this, although it might stop a maintainer from throwing poorly tested code over the fence instead of subjecting it to the same levels of scrutiny they give 3rd party contributors.
People, companies, countries that do this, will be overtaken technologically by others that accept the brittleness and move faster.
I think the solution is to have a balanced approach, both to advance relatively fast and keep things relatively robust. Who knows, in the end, maybe this crash is a reasonable price to pay for all the security Crowdstrike has provided over some time. It's not at all easy to tell.
Creating exploit-free code is another matter - you have to be able to craft exploit-free specifications, and there's no real understanding what that might even mean. But bug-free software would be a start.
And these parts are the simple ones. Not even talking about operating systems, networking and so on... If even easy stuff is wrong, what hope is there for complex...
https://webcache.googleusercontent.com/search?q=cache:https:...
https://cc.bingj.com/cache.aspx?d=464600016483&w=i7yXBm7Gwof...
Practically speaking, that is all the end-user can do with Windows machines. My point is Windows is fundamentally unsecure. It is a dike with thousands of holes some of which are not even visible to Microsoft themselves. The reason of that is security has been after-thought. It is band aids/plasters put on top of other plasters.
There are people in that category who are not hands-on themselves but still have sufficiently deep understanding of technical details. But as one might guess, they are about as common as four-leaf clovers.
I know everybody hates the C-word but if I look at 27001 requirements or the CIS benchmarks, there is nothing in there that I do not want for myself. If you can keep a list of the products and services you are running, have actually put the time into implementing it correctly, and have an ongoing maintenance plan then you are probably in the top 1% of networks.
Right now, pretty much everyone is looking to outsource their "security" to a single vendor, disregarding the fact that security is not a product, but a process.
That... won't change! And incumbents will get less-awful about their impact on "protected" systems.
And yet, there's an opportunity here! Do you truly understand Windows? And whatever happens on that platform? And how to monitor that activity for adverse actions? Without taking down your customers on a regular/observable basis?
Step right up! There are a lot of incumbents facing imminent replacement...
Even these experts would probably want to consult with others instead of trying to just grow what's going on on their own.
I find it funny that adult security researchers still get away with identifying themselves with hacking monikers in public as if they were teenagers probing the local telco back in the 1980's.