Retrofitting spatial safety to lines of C++
security.googleblog.com
security.googleblog.com
Well, it's 2024 and remember arguing this 20+ years ago. Programs have bugs that bounds checking catches. And making it a language built-in exposes it to compiler optimizations specifically targeting bounds checks, eliminating many and bringing the dynamic cost down immensely. Just turning them on in libraries doesn't necessarily expose all the compiler optimizations, but it's a start. Safety checks should really be built into the language.
I still don't get why the standard library went the other way, other than starting the tradition of standardised wrong defaults.
Instead WG21 has very clearly (but without ever admitting it and that's important) taken the path of maintaining a legacy language. Even as debate carried on about whether in future C++ could end up like COBOL, the committee has acted exactly as though it is for some years now. Compatibility is King, no price is too high for compatibility, everything must be sacrificed to make that happen and that's how you end up like COBOL.
Three important opportunities to divert and pick other ways forward should be highlighted here. P1863 "ABI: Now or Never" by Titus Winters in 2020; P2137 "Goals and priorities for C++" also in 2020 but with a long list of authors and P1818 "Epochs" from 2019 by Vittorio Romeo.
In all these cases WG21 chose the "hope the problem goes away" path, preferring not only not to address the critical problem highlighted and take a new route forward, but to specifically ignore the problem and press on anyway.
"Hope the problem goes away" is also, quietly, the preferred strategy by WG21 for the safety problem.
There's a reason (albeit a terrible one) to prefer the C++ ISO document's language over the approach of Rust. These are both general purpose languages (I might also write separately in this thread about a non-general purpose language which Google should use more, if I have time) and so must wrestle with Rice's Theorem. Rust's solution is to require the compiler to be conservative. This is very difficult and indeed there are known bugs in the code doing this conservative check in the Rust official compiler. But C++ has a much easier (but IMO fatal) path, it says that's the job of the programmer and when the programmer writes C++ software which is nonsense as a result that's their fault, not the compiler's fault for failing to reject the program.
It would be extremely difficult to explain how a "standards conforming" Rust compiler can correctly accept all the programs Rust's actual compiler accepts and reject all those it rejects without essentially having a black box where the compiler implementation sits. We can explain the purpose of such rules without, but their detailed behaviour not so much.
Take borrow checking. All the easy scoped borrows (which is all that worked in Rust say eight years ago) can be explained without too much trouble, but today a lot fancier (but to a human obviously correct) borrowing will compile, because the checker is smarter - now, how do you express, not in Rust source code but in the English language, all the checks to be performed, and neither miss things out nor unknowingly accept programs a real Rust compiler will reject ?
C++ just needn't do that, in effect the ISO document says. "Don't do borrows that last longer than the thing borrowed, if you do, that's not C++ but your compiler won't notice so the result is arbitrary nonsense"
I think this is unintentionally stuck in the mindset of "the purpose of a language specification document is to enable armchair language lawyers to flame each other on Usenet about whether or not such-and-such degenerate edge case is technically valid". But a specification doesn't need to be written in English, it can be written as a formal proof, and indeed I would expect a theoretical Rust spec to specify the behavior of the borrow checker as just such a proof. Rust's borrow checking may no longer be as simple as the lexically-scoped model that existed as of Rust 1.0, but it's not like the extensions that have been added since then are ad-hoc; they're all still designed to result in a model that is provably sound.
Did you miss the part where the person you're responding to mentioned Rice's theorem? Do you know Rice's theorem and hence understand what they're implying?
The actual logic of gggp's statement also doesn't make any sense. We as humans also under and overestimate the soundness of programs.
Sometimes, a perfectly fine solution is massaged to better adhere to best practices because we can't convince ourselves that it's correct. Rust requires that we convince the compiler, and then we know it's correct via the compiler's proofs, instead of requiring us to do the proof all the time.
It doesn't need evidence; it is the null hypothesis.
Brains clearly compute, and it appears that computation is sufficient to produce the observed behaviour of brains. All our experience of the universe and physics suggests that there is no magic or metaphysics or souls or whatever.
So the onus is on you to show that there's something more going on. It isn't a 50:50 "is it heads or tails", it's more like "I claim that the tooth fairy exists" vs "I'm pretty sure it's your mum".
So even if the Turing machine model is correct (and we don't know that), it's overtly simplified.
I'm saying we don't obey axioms of Turing machine model. So Rice theorem nor Godel theorem can apply to unsafe code written by humans.
Even if borrow checker is limited by the Rice theorem, you can create either safe abstractions provably or unsound abstractions provably or potentially unsound abstractions, which humans can reject or accept.
This is on that continuum where it's definitely neither impossible nor easy enough that we can just let some bored grad student knock out the answer, and so now somebody who wants this must do lots of hard work.
I think a specification which says e.g. here's the semantic requirement, here's a rule for scoped borrows which works, you must do at least that, but you can do more however you must not allow anything which violates the semantic requirement - would be great, but if you had that rule in your standard then people can write conforming Rust programs which don't compile - they need a yet-to-be-written smarter compiler to figure out why they're legal, which is kinda annoying as a language feature.
You absolutely could do it, but it would be a ridiculous effort for nebulous benefit.
...and then you still had to argue with some circles of the C++ community why the game and engine code doesn't use the stdlib. It's crazy that it takes decades to convince some people that a bad idea is simply a bad idea.
I used to have all kinds of problems with array overflows. I didn't make them very often, but when I did, they took a long time to track down. They've been gone for 20 years now.
Note that it would be easy to add it to C/C++:
https://www.digitalmars.com/articles/C-biggest-mistake.html
It would be the most useful and cost-effective enhancement ever.
No, it's not the same. I never enable optimisations by manually passing in flags to the compiler. It's always a `cmake -DCMAKE_BUILD_TYPE=...`. There is no such easily accessible equivalent for bounds checking.
Lol then you don't use your compiler/toolchain correctly. How is that anyone's problem but yours?
Maybe I'm not getting what you mean.
You are saying you already run
cmake ...
So I am saying you can just change that to CXXFLAGS="-Dblah" cmake -U CMAKE_CXX_FLAGS ...
That genuinely seems pretty darn easy to me.In any case, any beef you have is clearly with CMake here. You'd have the same issue(s) with any other flag, for any language, if you use CMake.
It is not in the standard, it isn't neither portable, nor guaranteed to exist.
This turned out to be the right move.
All I was doing here was saying was that the fix for your "C's biggest mistake" (your T arr[..] proposal) is already in C++ and you can get it today: it's called std::span, and it was explicitly designed to let you get bounds-checking, with just a different syntax. It needs a compiler flag, and so do optimizations. You already pass one, so pass the other too, and get what you wanted.
That was all I was saying. But this being HN, everyone insisted on derailing this into an argument about whether safe-by-default is better than fast-by-default, when that had nothing to do with my point, and when I was certainly not trying to argue one is better than the other.
If you propose something with blatantly obvious flaws here, you'll usually get called out.
You suggested that people use an interface without bounds checking and jump through a hoop to enable bounds checking with it. Other people disagreed that this is a solution. You kept digging deeper after that while ignoring their responses, but that's on you.
The problem you don't seem to understand is that, with this being HN, if I'd told people to use gsl::span, then I would have had a similar barrage of people "calling me out" for it having the "obvious flaws" of (1) destroying performance for users who don't want it, and/or (2) being nonstandard and in no way equivalent to the dlang.org proposal, this is why C++ sucks, blah blah. I might as well have just told them to write their own configurable wrappers at that point.
So I proposed std::span because it was literally the standard solution that was explicitly designed to let people get bounds checking without those problems... so that they can have their cake and eat it however they want, without an immediate performance loss. I frankly thought that was obvious, but this being HN, I was greeted with people "calling me out". It's like it's impossible to tell people something useful here without writing a comprehensive dissertation on the general topic. Makes me regret trying to help people.
They care now (well they pretend at least) because Rust is going to take significant market share in domains where C++ is still king.
And then evolution will take care of which programming ecosystem are less expensive to result in lawsuits, or invalidation of insurance policies.
The question was "is it so hard to pass a command line flag". You said "yes" when you clearly don't see any difficulty with actually passing the flag. Instead you're apparently answering a totally different question: "why do people lack the motivation to do this." Which had nothing to do with the point you replied to.
It's not like opt-out vs. opt-in somehow changes the performance characteristics. People who want maximum performance will turn it off. People who want safety will turn it on.
You don't feel you're missing the point of the discussion?
The whole discussion started with: "if you want bounds checking in your own code". Notice the "if". That's the premise.... it by definition assumes you've already accepted the performance impact of getting the safety you want, and thus it's not a problem for you.
The only remaining question at this point is, how hard is it to get you that safety. Asking you "is it so much harder to pass -foo like the -bar you already pass" and expecting you to address the physical difficulty of adding a flag isn't taking an "overly literal" reading of the question, it's literally asking the most obvious and only remaining question.
If you want to go back to the premise and argue about the psychological hurdle of taking a performance loss, that's fine and all, but then you're completely changing the topic of the thread you replied to.
P.S. comparing passing an extra command-line flag to shooting someone is a rather insane comparison. Honestly, all this is really making me regret trying to share a tip to help people make their code safer.
No, I'm regretting it because having to spend hours replying to comments that ignore the premise is a complete waste of my time.
> you keep telling people to use an interface that explicitly was designed to not provide bounds checking
As a matter of fact it was very intentionally and specifically designed to allow bounds-checking to be configured at build time: "As an example, in the current reference implementation, violating a range-check results by default in a call to terminate() but can also be configured via build-time mechanisms to continue execution (albeit with undefined behavior from that point on)." [1]
Calling that "explicitly designed not to provide bounds checking" is quite a deceptively misleading way to paint it. It's not an accident that you can enable bounds-checking, it's very much by design and intended that you do so. They just didn't happen to standardize the flag name, just like they never standardized the optimization flag names.
> and claiming that this is the solution to make their code safer, while in reality you have to look up some nonportable flag to enable it for your STL if even offers the functionality at all.
Like I said, this is literally the same as optimization flags. Everybody passes them and nobody bashes C++ for it. You're making a big deal out of something incredibly tiny just to win an internet argument on the wrong thread.
[1] https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2018/p01...
It’s not hard to grok.
Nobody was ever saying that unsafe-by-default is somehow better. That just wasn't the question being asked.
Can your position not be summed up as “unsafe by default doesn’t matter, because changing the default is easy”?
If so, there’s an obvious flaw in that thinking.
No.
>> Nobody was ever saying that unsafe-by-default is somehow better.
I wanted to ask: did you ever consider what was missing from Dlang to achieve widespread adoption? Clearly it was not features, so I'm wondering what that would be from your pespective.
For example, Borland at one point decided to include the source code to some of its runtime library for free. At a compiler roundup in the magazine, this was hailed as a great advance forward by the reviewer. Meanwhile, Datalight C was also in the roundup, and had always included 100% of the runtime library source code. No mention was made of this.
EASTL has this as a feature by default, and unreal engine container library has the boundchecks enabled on most games. The performance cost of those boundchecks in practice is well worth the reduction of bugs even on performance sensitive code.
Maybe what really happened is that compiler technology has improved such that they are able to remove most redundant checks, such that it only costs 0.30% today. I can imagine things going the opposite direction 20 years ago, as in "we removed some bounds checks and gained X% of performance".
Meanwhile the guys on the standards committee thinks of fixed width RISC instructions being executed by jungle logic and the ALU.
No one that wants to emit vectorized code is relying on auto-vectorization to emit that code.
And as others note, bounds checking was the norm before the STL.
That's something
Might explain why they claimed 70% of exploits were memory related..
In most implementations of the standard library, safety checks can be enabled with a simple #define. In some, it's the default behavior in DEBUG mode. I wonder what this library improves on that and why these bugs have not been discovered before.
Most folks don't use those #defines, and many still haven't leaned about them.
Source: I worked on this apparently
gsl::span is
> Fast mode, which contains a set of security-critical checks that can be done with relatively little overhead in constant time and are intended to be used in production.
> Using std::span as an example, setting the hardening mode to fast will always enable the valid-element-access checks when accessing elements via a std::span object, but whether dereferencing a std::span iterator does the equivalent check depends on the ABI configuration.
The standard doesn't require any checks to begin with.
It also doesn't require optimizations.
But you originally implied using span was sufficient, you didn't mention LLVM's libc++ hardening. (You even mentioned iterators which, I just quoted, might not be bounds-checked on fast mode either.)
When I said "the standard doesn't require this" I clearly was not referring to C++26, which does not even exist yet. In any case, I'm not sure what the point of this pedantry is. I'm pretty sure the point was clear.
> But you originally implied using span was sufficient, you didn't mention LLVM's libc++ hardening.
Because this isn't LLVM-specific, every major STL has bounds checking. You just gotta enable it for your toolchain. Sorry I didn't list every single flag, I guess?
> (You even mentioned iterators which, I just quoted, might not be bound-checked on fast mode either.)
Which is why I had _LIBCPP_ABI_BOUNDED_ITERATORS, right? I'm not on HN to write comprehensive documentation for every toolchain, I'm just writing a quick tip for people to look into.
All this pedantic quibbling over "this isn't required by the standard by default" is just pointless arguing for the sake of arguing on the internet. For all the performance freaks who really care about this: no language I know of guarantees optimizations in the standard, so if you're relying on optimized performance, you're already doing nonstandard stuff.
And practically every major compiled language you love or hate has a way to enable or disable bounds checking, letting you violate their "standard" one way or another. D itself has -boundscheck, C++ has toolchain-specific flags, Go has -gcflags=-B, etc...
As for the bounds-checked accessors, I mentioned them because they already exist in current C++ for other collections, they're coming to the one you suggested using, and I thought them relevant to a discussion about C++ lacking spatial safety.
I literally said exactly that: "The standard doesn't require any checks to begin with."
> Defaults matter!
Sigh... nobody claimed otherwise. You're really missing the point of the thread.
All I did was give people a tip on how to improve their code security. The exact sentence I wrote was:
>> "If you want bounds checking in your own code, start replacing T* with std::span<T> or std::span<T>::iterator whenever the target is an array."
"BUT DEFAULTS MATTER!!!", you rebut! Well OK, then I guess keep your raw pointers in and don't migrate your code? Sorry I tried to help!
Switch to std::span and add 1 line to std::span::operator[] to check your bounds...
That's why I said add 1 line to std::span::operator[] to check your bounds.
I'm telling you to modify the STL header. It's a text file. Add 1 line to make it bounds-checked.
Use gsl::span or write your own bounded span.
std::span is not bounds checked...
> Attackers regularly exploit spatial memory safety vulnerabilities, which occur when code accesses a memory allocation outside of its intended bounds
Isn't that... 'out of bounds memory access'?
Spatial memory safety is a reasonably common term in the security / PL field. You can see examples of it being used at least as far back as 2009: https://scholar.google.com/scholar?hl=en&as_sdt=0%2C33&q=spa...
It's in contrast to temporal memory safety, which deals with object lifetimes (use after free, for example).
Here Google is probably also referencing a 2022 post of theirs with a very similar title, dealing with temporal safety: https://security.googleblog.com/2022/05/retrofitting-tempora...
The terms are also in Wikipedia: https://en.wikipedia.org/wiki/Memory_safety#Classification_o...
Note that there are some more heated takes on where these terms are being used. I tried to be as generous as possible in my description.
It's the part of memory safety that's just about bounds. You can also call it "bounds safety" and folks will understand what you mean, but "spacial safety" is the more commonly used jargon.
Rust advocates tend to turn stats like this into “40% of all security issues are memory safety”, which sounds very similar but is false.
You're right that it's false. Historically it's been a much more damning 70% of vulnerabilities that were rooted in memory-unsafety.
According to the Google Security Blog, in a post linked to from the OP:
We’ll also share updated data on how the percentage of memory safety vulnerabilities in Android dropped from 76% to 24% over 6 years as development shifted to memory safe languages. [...] The percent of vulnerabilities caused by memory safety issues continues to correlate closely with the development language that’s used for new code. Memory safety issues, which accounted for 76% of Android vulnerabilities in 2019, and are currently 24% in 2024, well below the 70% industry norm, and continuing to drop.
https://security.googleblog.com/2024/09/eliminating-memory-s...
OWASPs top ten security vulnerabilities are not memory safety.
People don't write web apps in C++, because they would have to deal with memory safety issues in addition to all the other issues related to auth, injections, etc.
Why are you offended at the idea that languages should be memory safe by default? What code are you writing that you constantly need memory unsafety, constantly available, without being able to write any sort of "unsafe" keyword? Who cares about whether or not it's the #1 problem in OWASP when it's clearly and undeniably been a massive problem for decades? It is sufficient, after all, that it crashes a program or produces incorrect results for it to be a problem worth pursuing, but it is also extremely well known to produce massive security vulnerabilities regardless of what some list says.
Why is this a hill you are willing to die on? What are you getting out of it? Is your programming life going to be easier? Are you better off when debugging something to not be able to just know that it's not a memory safety problem, and thus to still have to consider it?
What actual engineering benefit do those rare few of you who seem to be crusading against memory safety fear disappearing?
When I got into programming in the late 1990s, I was there to catch the last few holdouts of the "everyone should just write in assembler" opinion. I at least understood their arguments around performance and efficiency, and I understood their arguments around "not needing high level languages" even though I disagree with them both then and now. I think on the net they were wrong, but they did have some legitimate benefits to argue on their side, even if they were already outweighed by the costs then and even more so outweighed today.
But I don't get what you folk furious about memory safety are looking for. "Using" memory safety is already an invalid program. It's already pretty much automatically a bug, if not worse. You're not losing anything to simply have it, you're not gaining anything except bugs and sharp corners insisting on it. And when you absolutely, positively need it, which I'd call "exceptionally rare but definitely non-zero", it's still there in one form or another of "unsafe". I don't see any benefits at all.
(And let me reiterate and forstall the usual, memory safety does not mean "Rust". Memory safety is every major language on the market today except C and C++.)
Why are you okay with languages that are not overflow-safe, or unit-safe, or infinite-loop-safe, or safe against bit flips? Memory safety violations are a major chunk of bugs. Writing code to avoid them is about as hard as writing code to avoid other major classes of bugs. In either case, it’s failable. Static analysis and testing then gives confidence that the system is safe, by multiple metrics. Memory safety isn’t special enough to demand a different approach here — quality code requires a coherent approach to quality across multiple bug classes.
You argue like we live in some hypothetical universe where only some bizarre academic language has recently invented the idea in a world where nobody else has even heard of the idea, and it's solving a problem we don't generally have. But the truth is, we already have memory safety... everywhere except C and C++. Those languages stand alone now. They are the only ones where it's an issue. And they have demonstrated in as concrete an engineering way as it can be demonstrated that it is a problem, on numerous levels.
You're not arguing against some new fangled idea that has no evidence. You're arguing against something that is completely normal engineering practice in place almost everywhere, and the rest of us look at you arguing against it as if you're arguing against that source control is a stupid idea for people who can't keep track of the changes they've made, by gosh, just sticking random prefixes and suffixes on my files is enough for me and it ought to be enough for everyone. We're not hypothesizing about it. We've been living it for decades. We're not asking the world to change to be memory safe... it already has. Except C and C++.
We got too many C refugees that spoiled the soup.
Rust advocates like to muddy the water and make it sound like memory safety is the biggest issue in security. It isn’t.
Or advocates like NSA and FBI?
Security FUD, name calling Rust any time someone raises security issues, is quite impressive.