Eliminating Memory Safety Vulnerabilities at the Source
security.googleblog.com
security.googleblog.com
This also implies that languages and tooling with robust support for integrating with unsafe legacy code are even more desirable.
It’s the only approach that has any chance of transitioning away from unsafe languages for existing, mature codebases. Rewriting entirely in a different language is not a reasonable proposition for every practical real-world project.
It's absolutely true that you need integration and compatibility to enable iterative improvements, but Rust historically has been hostile to anything besides unsafe C ABI FFI, which is not suitable for the vast majority of incremental development that needs to happen.
Luckily, this is starting to change.
This means that "solving this" requires partial integration of a C++ compiler; it's not a coincidence that the languages with the most success have been backed by the organisations that already had their own C++ compiler.
A much easier solution is to generate glue on both sides, which is what `cxx` does.
d does support c++ abi. It seems almost dead now but it is possible.
realistically there are two c++ abis in the world. Itanimum and msvc. both are known well enough that you can imblement them if you want (it is tricky)
It's not hostile to not commit resources to integrating with another language. It might be shortsighted, though.
> d does support c++ abi. It seems almost dead now but it is possible.
That was made possible by Digital Mars's existing C++ compiler and their ability to integrate with it / borrow from it. Rust can't take the same path. Additionally, D's object model is closer to C++ than Rust's is; it's not a 1:1 map, but the task is still somewhat easier. (My D days are approaching a decade ago, so I'm not sure what the current state of affairs is.)
> realistically there are two c++ abis in the world. Itanimum and msvc. both are known well enough that you can imblement them if you want (it is tricky)
The raw ABI is doable, yeah - but how do you account for the differences in how the languages work? As a simple example - C++ has copy constructors, Rust doesn't. Rust has guarantees around lifetimes, C++ doesn't. What does that look like from a binding perspective? How do you expose that in a way that's amenable to both languages?
`cxx`'s approach is to generate lowest-common denominator code that both languages can agree upon. That wouldn't work as a language feature because it requires buy-in from the other side, too.
I don't see Cranelift ever replacing them as the main backend.
On what basis do you make this flagrant claim? There is certainly an impedance mismatch between Rust and C++, but "kicking and screaming" implies a level of malice that doesn't exist and never existed.
Such a low quality comment.
As opposed to integrating with a whole C++ frontend? Gee, I wonder why more languages haven't done this obvious idea of integrating a whole C++ frontend within their build system.
Rust has integration with the C ABI(s) from day 1, which makes sense because the C ABI(s) are effectively the universal ABI(s).
> and its C++ compatibility is a joke
I'm not sure why you would want Rust to support interoperability with C++ though, they are very different languages. Moreover why does this fall on Rust? C++ has yet to support integration with Rust!
> compared to Swift
Swift bundled a whole C++ compiler (Clang) to make it work, that's a complete deal breaker for many.
> Rust historically has been hostile to anything besides unsafe C ABI FFI
Rust historically has been hostile to making the Rust ABI stable, but not to introducing other optional ABIs. Introducing e.g. a C++ ABI has just not been done because there's no mapping for many features like inheritance, move constructors, etc etc. Ultimately the problem is that most ABIs are either so simple that they're the same as the C ABI or they carry so many features that it's not possible to map them onto a different language.
I do agree that they could do more. I find it kind of wild that they haven’t copied Zig’s cross compilation story for example. But this stuff being just fine instead of world class is more of a lack of vision and direction than a hostility towards the idea. Indifference is not malice.
Also, you seem to be ignoring the crabi work, which is still not a thing, of course, but certainly isn’t “actively hostile towards non-c ABIs”.
It seems they have, which is all to the good: the work to make custom allocators practical in Rust is another example. Rust still doesn't have anything like `std.testing.checkAllAllocationFailures`, but I expect that it will at some future point. Zig certainly learned a lot from Rust, and it's good that Rust is able to do the same.
Zig is not, and will not be, a memory-safe language. But the sum of the decisions which go into the language make it categorically different from C (the language is so different from C++ as to make comparisons basically irrelevant).
Memory safety is a spectrum, not a binary, or there wouldn't be "Rust" and "safe Rust". What Rust has done is innovative, even revolutionary, but I believe this also applies to Zig in its own way. I view them as representing complementary approaches to the true goal, which is correct software with optimal performance.
I've known Andrew for a long time, and am a big fan of languages learning from each other. I have always found that the community of people working on languages is overall much more collegial to each other than their respective communities can be to each other, as a general rule. Obviously there are exceptions.
- C ABI is good enough for any compatibility
- "Rewrite it Rust" / RESF
- Prioritize language proposals, not fundamentals
I think the attitude as changed somewhat in the last 2-3 years, with more interest in cxx and crubit and crabi, although many of those things are at least partially blocked on stalled language proposals, and largely ignored by the "core" Rust community. I would say indifference and posturing does approach hostility.
There's also consistent conflation of "compatibility with C++" and "rich compatibility with all of C++". Swift is a great example of how targeting PODs and std::vector and aligning concurrency expectations gets you extremely far, especially if you limit yourself to just LLVM.
btw, I really appreciate your Rust book.
I think indifference is very different than hostility. That there's only limited time, resources, and interest and those must be balanced and that means some projects advance before others is a very different thing than a refusal to do something.
They are talking about actually safe languages, secury by design and secure by default. Which would be a lisp or scheme without an FFI, or a beam language (Erlang, Elixir). Or pony, without the FFI. Or Concurrent Pascal, not Go.
Or a safe scripting language
Such projects exist. That's fine.
Didn't something similar happened at Discord. It was Go if I recall.
Apparently there is unwillingness to have "high class talent", while starting from scratch in a completly different programming language stack where everyone is a junior, is suddenly ok. Clearly CV driven development decision.
Secondly, in many cases, even if all optimization options on the specific language are exhausted, it is still easier to call into a native library in a lower level language, than a full rewrite from scratch.
And totally agree about CV driven development. All these rewrite articles looks like they are trying to convince themselves instead if intelligent readers about rewrites.
But most places are small, and engineers optimize the time to market while remaining at an acceptable levels of resource consumption and support expenses, by using stuff like next.js, RoR, etc. And that saves the company money.
There is also a spectrum in between, with associated hassles of transition as the company grows.
My favorite example is that eBay rewrote their backend three times, and they did it not because they kept building the wrong thing. They kept building the right thing for their scale at the moment. Baby clothes don't fit a grown-up, but wearing baby clothes while a baby was not a mistake.
Ideally, of course, you have a tool that you can keep using from the small prototype stage to the world-scale stage, and it lets you build highly correct, highly performant software quickly. Let's imagine that such a fantastical tool exists. To my mind, the problem is usually not in it but in the architecture: what's efficient at a large scale is uselessly complex at a small scale. The ability to express intricate constraints that precisely match the intricacies of your highly refined business logic may feel like an impediment while you're prototyping and haven't yet discovered what the logic should really be.
In short, more precision takes more thinking, and thinking is expensive. It should be applied where it matters most, and often it's not (yet) the technical details.
Plenty of people keep putting GC languages on the same basket without understanding what they are talking about.
Then if it is a domain where any kind of automatic resource management is impossible due to execution deadlines, or memory availability, Rust is an option.
Though I like using effect systems like in Nim or Ocaml for preventing allocation in specific areas.
The point is not winning micro-benchmarks games, rather there are specific SLAs for resource consumption and execution deadlines, is the language toolchain able to meet them, when the language features are used as they should, or does it run out of juice to meet those targets?
The recent Guile performance discussion thread is a good example.
If language X does meet the targets, and one still goes for "Rewrite XYZ into ZYW" approach, then we are beyond pure technical considerations.
> The recent Guile performance discussion thread is a good example.
Sounds very interesting. Can you share a link?GC.TryStartNoGCRegion - https://learn.microsoft.com/en-us/dotnet/api/system.gc.tryst...
GCSettings.LatencyMode - https://learn.microsoft.com/en-us/dotnet/standard/garbage-co...
The latter is what Osu! uses when you load and start playing a map, to sustain 1000hz game loop. Well, it also avoids allocations within it like fire but it's going to apply to any non-soft realtime system.
Only pointing out the ones relevant in 2024, and not the whole CS history.
I wonder if there's a "bathtub curve" effect for very old code. I remember when a particularly serious openssl vulnerability (heartbleed?) caused a lot of people to look at the code and recoil in horror at it.
However, it is impressive that from 2019 to 2023 the issues went down by almost 60% without rewriting, just by adding new code in MSLs (Rust mostly, I guess, probably a bit Kotlin). 60% is a massive achievement. I wonder why there's a plateau from 2021 and 2023. The fact that they use extrapolated number for 2024 reveals their intention to show nice numbers, but it would have been more beneficial if they'd have spent the time to analyse and explain the rise from 2022 to 2023.
I'm a bit disappointed that they don't tell us how much of the new MS code is written in which language. Seems like they're using both Rust and Kotlin. Is it 95% Rust? Or "just" 50% Rust? If the Kotlin portion is significant, in what areas do they use it instead of Rust?
So the upshot of the fact that vulnerabilities decay exponentially is that the focus should be on net-new code. And spending effort on vast indiscriminate RiiR projects is a poor use of resources, even for advancing the goal of maximal memory safety. The fact that the easiest strategy, and the strategy recommended by all pragmatic rust experts, is actually also the best strategy to minimize memory vulnerabilities according to the data is notably convergent if not fortuitous.
> The Android team has observed that the rollback rate of Rust changes is less than half that of C++.
Wow!
It stands to reason, then, that it would be even better for security to stop adding new features when they aren't absolutely necessary. Windows LTSC is presumably the most secure version of Windows.
Obviously there’s a ton of variance in how practical this is any place, but it’s less common than it should be.
Congrats you're back to square 1!
Yeah, but are those bugs security bugs? Memory safety bugs are a big focus because they're the most common kind of bugs that can be exploited in a meaningful way.
Disabling entire segments of code is unlikely to introduce new memory safety bugs. It's certainly likely to find race conditions, and those can sometimes lead to security bugs, but its not nearly as likely as with memory safety bugs.
If the software is unusable, it doesn't matter if it has security bugs too. Or, to rephrase, the safest software is software nobody uses.
Only really practical if "features" are "plugins".
Having lots of knobs you can tweak is great for randomized testing. The more the merrier.
Even if features aren't necessary to sell your software, new hardware and better security algorithms or full on deprecation of existing algos will still happen. Which will introduce new code.
Not-yet exploited vulnerabilities, though, don't have that decay mechanism. They don't generate user unhappiness and bug reports. They just sit there, until an enemy with sufficient resources and motivation finds and exploits them.
There are more enemies in that league than there used to be.
Looking at vulnerabilities that were found from attacks, it looks different. [1] Most vulnerabilities are fixed in the first weeks or months. But ones that aren't fixed within a year hang on for a long time. About 18% of reported vulnerabilities are never fixed.
[1] https://www.tenable.com/blog/what-is-the-lifespan-of-a-vulne...
You can also uncover latent vulns over time through fuzzing or by adding new code that suddenly exercises new paths that were previously ill-tested.
Yes, there are some vulns that truly will never get exercised by ordinary interaction and won't become naturally visible over time. But plenty do get uncovered in this manner.
This should be true not just of vulnerabilities, but bugs of any kind. I certainly see this in testing of the free software project I'm involved with (SBCL). New bugs tend to be in parts that have been recently changed. I'm sure you all have seen the same sort of effect.
(This is not to say all bugs are in recent code. We've all seen bugs that persist undetected for years. The question for those should be how did testing miss them.)
So this suggests testing should be focused on recently changed code. In particular, mutation testing can be highly focused on such code, or on code closely coupled with the changed code. This would greatly reduce the overhead of applying this testing.
Google has had a system where mutation testing has been used with code reviews that does just this.
The bleeding edge is where many of the new vulns are. In general the oldest supported release is usually the safest.
The trade-off is when newer versions have features which will add value, of course. But usually a bad idea to take any version that ends “.0” IMO.
There is more than one possible and reasonable explanation for this correlation:
1. New code often relates to new features, and folks focus on new features for vulnerabilities. 2. Older code has been through more real life usage, which can exercise those edge cases where memory vulnerabilities reside.
I’m just not comfortable saying new code causes memory vulnerabilities and that vulnerabilities have a half-life that decays rapidly. That may —- may be true in sheer number count, but doesn’t seem to be true in impact, thinking back to the high-impact vulnerabilities in OSS like the heartbleed bug, and the cache-invalidation bugs for CPUs.
This essay takes some interesting data from very specific (and unusual) projects and languages from a very specific (and unusual) culture and stridently extrapolates hard numeric values to all code without qualification.
> For example, based on the average vulnerability lifetimes, 5-year-old code has a 3.4x (using lifetimes from the study) to 7.4x (using lifetimes observed in Android and Chromium) lower vulnerability density than new code.
Given this conclusion, I can write a defect-filled chunk of C code and just let it marinate for 5 years offline in order for it become safe?
I'm pretty sure there are important data in this research and there is truth underneath what is being shared, but the unsupported confidence and overreach of the writing is too distracting for me.
> stridently extrapolates hard numeric values to all code without qualification.
The sentence they quote as evidence of this directly qualifies that this is from Android and Chromium.
I concede this may not be the strongest example, but in my opinion, the language throughout the article, starting with the title, makes stronger claims than the evidence provided supports.
I agree with the author, that these are useful projects to use for research. I'm struggling with the lack of qualification when it comes to the conclusions.
Perhaps I missed it, but I also didn't see information about trade-offs experienced in the transition to Rust on these projects.
Was there any change related to other kinds of vulnerabilities or defects?
How did the transition to Rust impact the number of features introduced over a given time period?
Were the engineers able to move as quickly in this (presumably) new-to-them language?
I'm under the impression that it can take many engineers multiple years to begin to feel productive in Rust, is there any measure of throughput (even qualitative) that could be compared before, during and after that period?
I'm hung up on what reads as a sales pitch that implies broad and deep benefits to any software project of any scope, scale or purpose and makes no mention of trade offs or disadvantages in exchange for this incredible benefit.
> I also didn't see information about trade-offs experienced in the transition to Rust on these projects.
Yeah, I mean that's just not the topic of this particular post. But they have talked about it. Specifically this, from 2023:
https://opensource.googleblog.com/2023/06/rust-fact-vs-ficti...
> I'm under the impression that it can take many engineers multiple years to begin to feel productive in Rust, is there any measure of throughput (even qualitative) that could be compared before, during and after that period?
That is something that people say, but Google found differently. As that post says, "more than 2/3 of respondents are confident in contributing to a Rust codebase within two months or less when learning Rust. Further, a third of respondents become as productive using Rust as other languages in two months or less."
On this one:
> Were the engineers able to move as quickly in this (presumably) new-to-them language?
Regarding "presumably," 13% had Rust experience, but the rest did not.
They say "a third of respondents become as productive using Rust as other languages in two months or less" and "we’ve seen no data to indicate that there is any productivity penalty for Rust relative to any other language these developers previously used at Google."
My current company is an example of this. Early code written by founders. Some new code written by contractors under a tight deadline.
> The Android team has observed that the rollback rate of Rust changes is less than half that of C++.
I've been writing high-scale production code in one language or another for 20 years. But I when I found Rust in 2016 I knew that this was the one. I was going to double-down on this. I got Klabnik and Carol's book literally the same day. Still have my dead-tree copy.
It's honestly re-invigorated my love for programming.
People who care about this issue, especially in the last few years, have been leaning into a "memory safe language" vs "non memory safe language" framing. This is because it gets at the root of the issue, which is safe by default vs not safe by default. It tries to avoid pointing fingers at, or giving recommendations for, particular languages, by instead putting the focus on the root cause.
In the specific case of Android, the subject of this post, I'm not aware of attempts to move into other MSLs than those. But I also don't follow Android development generally, but I do follow these posts pretty closely, and I don't remember any of them talking about stuff other than Rust or Kotlin.
Don’t forget the old, boring one: Java.
I assume the reason that Go doesn’t show up so much is that most Android processes have their managed, GC’d Java-ish-virtual-machine world and their native C/C++ world. Kotlin fits in with the former and Rust fits in with the latter. Go is somewhat of its own thing.
Google also published their perspective on memory safety in https://security.googleblog.com/2024/03/secure-by-design-goo..., which also goes over some of the memory-safe languages in use like Java, Go and Rust.
Across Google, Go is used for some system software, but I haven't seen it used in Android.
For that reason I'd use OCaml as well even though it has GC, because it has sum types. That is, if I ever learn OCaml properly.
> memory safe languages
I would say anything that runs on JVM and CLR, and scripting langs, like Python, Perl, Ruby, etc.Edit: I forgot Golang!
Rust and Swift are the two most widely used.
Interestingly, Swift had interoperating with C as an explicit design goal, while Rust had data race safety as a design goal.
Now we have data race safety added in the latest version of Swift, and Rust looking to improve interoperability with C.
How much it gets there, depends on squizzing juice out of LLVM backend for Swift code.
Chris Lattner, the creator of both LLVM and Swift, has referred to Swift as “syntactic sugar for LLVM.” They are deeply tied together.
The amount of Swift code in Apple's operating systems has increased every year.
https://blog.timac.org/2023/1019-state-of-swift-and-swiftui-...
Lattner gave an interview looking at the advantages of reference counting over the sort of garbage collection used in languages like Java, C#, and Go while still avoiding error-prone manual memory management.
> ARC has clear advantages in terms of allowing Swift to scale down to systems that can’t tolerate having a garbage collector, for example, if you want to write firmware in Swift. I think that it does provide a better programming model where programmers think just a little bit about memory.
Apple has also optimized their custom ARM core to further reduce the cost.
> retaining and releasing an NSObject takes ~30 nanoseconds on current gen Intel, and ~6.5 nanoseconds on an M1
https://blog.metaobject.com/2020/11/m1-memory-and-performanc...
That said, Swift is working toward adding a future opt-in Rust inspired approache to memory management for those who need it.
https://forums.swift.org/t/manifesto-ownership/5212
https://forums.swift.org/t/a-roadmap-for-improving-swift-per...
Amazing, I've never seen this argument used to support shift/left secure guardrails but it's great. Especially for those with larger, legacy codebases who might otherwise say "why bother, we're never going to benefit from memory-safety on our 100M lines of C++."
I think it also implies any lightweight vulnerability detection has disproportionate benefit -- even if it was to only look at new code & dependencies vs the backlog.
It's far more common to look at recent commit logs than it is to look at some library that hasn't changed for 20 years.
The commits are just used for attribution. If there was some old lib that hasn’t been changed in 20 years that’s passed fuzzing and manual code inspection for 20 years without updates, chances are it’s solid.
For example: maybe the engineers over the last several years have focused on rewriting the riskiest parts in a MSL, and were less likely to change the lower risk old code.
Or… maybe there was a process or personnel change that led to more defects.
With that said, it does seem plausible to me that any given bug has a probability of detection per unit of time, and as time passes fewer defects remain to be found. And as long as your maintainers fix more vulnerabilities than they introduce, sure, older code will have fewer and the ones that remain are probably hard to find.
It wasn’t being look at as hard before either. I don’t think that’s changed.
They don’t give a theory for why older code has fewer bugs, but I’ve got one: they’ve been found.
If we assumed that any piece of code has a fixed amount of unknown bugs per 1000 lines, it stands to reason that overtime the sheer number of times the code is run with different inputs in prod makes it more and more likely they will be discovered. Between fixing them and the code reviews while fixing them the hope would be that on average things are being made better.
So overtime, there are fewer bugs per thousand lines in existing code. It’s been battle tested.
As the post says, if you continue introducing new bugs at the same rate you’re not going to make progress. But if using a memory safe language means you’re introducing fewer bugs in new features then overtime the total number of bugs should be going down.
My concern with it is more about legitimately old code (android is 20ish years old, so reasonably falls into this category) which was written using standards and tools of the time (necessarily)
It requires a constant engineering effort to keep such code up to date. And the older code is, typically, less well understood.
In addition older code (particularly in systems programming) is often associated with older requirements, some of which may have become niche over time.
That long tail of old, less frequently exercised, code feels like it may well have a sting in its tail.
The halflife/work-hardening model depends on the code being stressed to find bugs
I haven't found any language usage numbers for recent versions of Windows, but Microsoft is using Rust for both new development and rewriting old features [1] [2].
[0] Refer to section "Evolution of the programming languages" https://blog.timac.org/2023/1128-state-of-appkit-catalyst-sw...
[1] https://www.theregister.com/2023/04/27/microsoft_windows_rus...
[2] https://www.theregister.com/2024/01/31/microsoft_seeks_rust_...
• Use-after-frees are avoided by ARC
• Null pointer dereferences are usually safe (sending a message to nil returns nil)
• Objective-C has a great standard library (Foundation) with safe collections among many other things; most of C's dangerous parts are easily avoided in idiomatic Objective-C code that isn't performance-critical
But a good part of Apple's Objective-C code is probably there for implementing the underlying runtime, and that's difficult to get right.
Just to summarize the article, it shows that writing completely new code in memory safe language, while maintaining non-memory safe code, results in a steep reduction in memory safe errors overtime even though it results in an overall increase in unsafe code. It says that most memory safe vulnerabilities come from completely new code not maintained code and thus argues you can get the most of the benefits of memory safe code without rewriting your entire code base, which I think is the main takeaway from the article.
I’m not sure that’s totally happening in MacOS from reading your article, but it kind of is, so I think my hypothesis is correct that MacOS will likely have less vulnerabilities as it transitions many newer projects to swift although its important to note that important vulnerable projects such as webkit are still written in C++.
Do you have data for that? My impression is that a large fraction of Windows development is C# these days. Back when I was at EA, nearly fifteen years ago, we were already leaning very heavily towards C# for internal tools.
WinRT is basically COM and C++, nowadays they focus on C# as consumer language, mostly because after killing C++/CX, they never improved the C++/WinRT developer experience since 2016, and only Windows teams use it, while the few folks that still believe WinUI has any future rather reach out for C#.
If you want to see big chuncks of C# adoption you have to look into business units under Azure org chart.
So if this blog post describes the 4th generation, perhaps the 5th generation looks something like Lockdown Mode for iOS. Let users who are concerned with security check a box that improves their security, in exchange for decreased performance. The ideal checkbox detects and captures any attack, perhaps through some sort of virtualization, then sends it to the security team for analysis. This creates deterrence for the attacker. They don't want to burn a scarce vulnerability if the user happens to have that security box checked. And many high-value targets will check the box.
Herd immunity, but for software vulnerabilities instead of biological pathogens.
Security-aware users will also tend to be privacy-aware. So instead of passively phoning home for all user activity, give the user an alert if an attack was detected. Show them a few KB of anomalous network activity or whatever, which should be sufficient for a security team to reconstruct the attack. Get the user to sign off before that data gets shared.
The reduction of memory safety bugs to a projected 36 in 2024 for Android is extremely impressive.
...
In the final year of our simulation, despite the growth in memory-unsafe code, the number of memory safety vulnerabilities drops significantly, a seemingly counterintuitive result [...]
Why would this be counterintuitive? If you're only touching the memory-unsafe code to fix bugs, it seems obviously that the number of memory-safety bugs will go down.
Am I missing something?
It's not as if bug fixes haven't resulted in new memory bugs, but apparently that rate is much lower in bug fixes than it is in brand new code.
Instead they’ve shown that only using memory safe languages for new code is enough for the total bug count to drop.
So while not technically "requiring" C/C++, if your language cannot map exactly to C/C++ data structs & type definitions - it won't work.
Alternative is to, y'know, write a kernel in your language of choice and choose your own syscall specification suiting that language, and gain mass adoption. Easy!
If you add up all the JavaScript, C#, Java, Python, and PHP being written every year, that’s a lot of code.
Are we sure that all that combined isn’t more than C/C++? Or at least somewhat close?
It is like complaining a bullet vest doesn't protect against heavy machine gun bullets.
And all practical Pascal clones (e.g. Object Pascal) had to grow constructor/destructor systems to cleanup the memory on de-allocation (or they went the garbage collection route). So they were on-par with C++ for safety.
Ada is similar. It provided safety only for static allocations with bounds known at the compile-time. Their version of safe dynamic allocations basically borrowed Rust's borrow checker: https://blog.adacore.com/using-pointers-in-spark
We are beyond Ada 83, Controlled Types and SPARK exist, and Rust also does runtime bounds checking, so what.
ESPOL is one of the first recorded uses of unsafe code blocks, 1961.
They are eventually forced to transition to a new language, which makes the memory safety bugs moot. Without addressing the fact that they're still sub-par, or why they were to begin with, why they didn't use the memory safe functions, why we let them ship code to begin with.
They go on to make more sub-par code, with more avoidable security errors. They're just not memory safety related anymore. And the hackers shift their focus to attack a different way.
Meanwhile, nobody talks about the pink elephant in the room. That we were, and still are, completely fine with people writing code that is shitty. That we allow people to continuously use the wrong methods, which lead to completely avoidable security holes. Security holes like the injection attacks, which make up 40% of all CVEs now, when memory safety only makes up 25%.
Could we have focused on a default solution for the bigger class of security holes? Yes. Did we? No. Why? Because none of this is about security. Programmers just like new toys to play with. Security is a red herring being used to justify the continuation of allowing people to write shitty code, and play with new toys.
Security will continue to be bad, because we are not addressing the way we write software. Rather than this one big class of bugs, we will just have the million smaller ones to deal with. And it'll actually get harder to deal with it all, because we won't have the "memory safety" bogey man to point at anymore.
A lot of the recent moves here are motivated by multiple large industry players noticing a correlation between security issues and memory-unsafe language usage. "70% of security vulnerabilities are due to memory unsafety" is a motivating reason to move towards memory safe languages.
What do you believe the underlying cause to be?
If there is no process to eliminate the defects, then they will always appear. Based on my knowledge and experience of this kind of software development, I have noticed that a majority of the time, there is no process applied to the development which would eliminate these defects. We just expect, or hope, that the developers are smart enough and rigorous enough to do it by themselves without anyone asking. Well, in the real world, you can't just hope for quality. The results speak for themselves.
Even when there is a process applied to eliminate the defect, it is often not effective. Either the process doesn't go far enough (many still don't even use Valgrind), or the people applying the process either don't follow it, or not well enough.
Many in the industry have decided that the solution to these defect issues is to use a new system [programming language] which avoids the need for the process to begin with. That's not a bad idea in principle. But they've only done that for a single class of quality defect. The rest of the security bugs, and every other kind of defect imaginable, is still there. And new ones will appear over time.
So rather than play whack-a-mole very slowly designing brand new systems to eliminate one single class of defect at a time, my proposal is we change the way we do development fundamentally. Regardless of the system in use, we should be able to define processes, ensure they are applied correctly, and continuously improve them, to eliminate quality defects.
This would not only solve security issues, but all kinds of bugs. If this method was standard practice (and mandatory), the Crowdstrike issue never would have happened, because they would have been required to be looking for quality defects and eliminating them.
However, I don't see that as being in conflict with using tools that can also help you achieve higher quality within a certain stage of that process.
I think a lot of the arguments around C++ for example being 'memory unsafe' is a bit ridiculous because its trivial to write memory safe C++. Just run -Wall and enforce the use of smart pointers, there are nearly zero instances in the modern day where you should be dealing with raw pointers or performing offsets that lead to these bugs directly. The few exceptions are hopefully with devs that are intelligent enough to do so safely with modern language features. Unfortunately, this rarely gets focused on by security teams it seems since they are instead chasing the newest shiny language like you mention.
It is not unfortunately. That's why we see memory safety being responsible for 70% of severe vulns across many C and C++ projects.
Some of the reasons include: - C++ does little to prevent out-of-bounds vulns - Preventing use-after-free with smart pointers requires heavy use of shared pointers, which often incurs a performance cost that is unacceptable in the environment C++ is used.
I don't think that's really a rebuttal to what they're trying to say. If the vast majority of C++ devs don't follow those two rules, then that's not much evidence against those two rules providing memory safety.
For performance, you'll have to be more specific about your shared pointer claims. But I bet that it's a very small fraction of C++ functions that need the absolute best performance and can't avoid those performance problems while following the two rules.
Very bold claim, and as such, it needs substantial evidence, as there is practically no meaningful evidence to support this. There are some real world non-trivial c++ code that are known to have very few defects, but almost all of them required extremely significant effort to get there.
It is also far easier to apply formal analysis to code if you don't have to model arbitrary pointers.
Android has hard evidence that just eliminating memory safety makes a significant difference.
It's absolutely possible to write a reasonably usable language that makes injection/escaping/pollution/datatype confusion errors nearly impossible, but it would involve language support and rewriting most libraries--just like memory safety did. Unfortunately we are moving in the opposite direction (I'm still angry about javascript backticks, a feature seemingly designed solely to allow the porting of php-style sql injection errors)
You know what you have there? A new method for developing software.
If we just switch languages every time we have another class of vuln that we want to eliminate, it will take us 1,000 years to get them all.
Or we could just get them all today by fundamentally rethinking how we write software.
The problem isn't software. The problem is humans writing software. There are bugs because we are using a highly fallible process to create software. If we want to eliminate the bugs, we have to change how the software is created.
We should really stop putting the blame on developers. The issue is not that developers are sub-par, but that they are provided with tools making it virtually impossible to write secure software. Everyone writes memory safety bugs when using memory-unsafe languages.
And one helpful insight here is that the security posture of a software application is substantially an emergent property of the developer ecosystem that produced it, and that includes having secure-by-design APIs and languages. https://queue.acm.org/detail.cfm?id=3648601 goes into more details on this.
You can't rely on people being perfect all the time. We've been trying that for 50 years, and only got an endless circle of CVEs and calls to find better programmers next time.
The difference is how the language reacts to the mistakes that will happen. It could react with "oops, you've made a mistake! Here, fix this", and let the programmer apply a fix and move on, shipping code without the bug. Or the language could silently amplify smallest mistakes in the least interesting code into corruption that causes catastrophic security failures.
When concatenating strings and adding numbers securely is a thing that exists, and a thing that requires top-skilled programmers, you're just wasting people's talent on dumb things.
…which are?
I'm sorry, are we still talking about C here? Where the old functions are the ones that are explicitly labelled as too dangerous to use, like strcmp()?