GCC's new fortification level: The gains and costs
developers.redhat.com
developers.redhat.com
I thought that hack was dead and buried. It won't work in debug modes where buffers are zeroed in "free".
Fear of copying time is usually misplaced. The PDP-11 is gone. Unless you're copying megabytes, the copy time of recently accessed in-cache data is very small.
It's the sort of arrogant attitude that's behind driving people away from C/C++
I hope the people behind Zig understand that in critical applications that's totally unacceptable.
If your point is that some software shouldn't crash, then yes, for sure. But that's on you to not make programming errors in your code.
In fact Zig does help you create software that doesn't crash like not many other programming languages do, for example by not having language features that rely on implicit memory allocations. This gives you the opportunity to always have a fallback strategy if a memory allocation fails.
1. Zig does not have more undefined behavior than C. Much like there are two kinds of complexity, accidental, and essential, C has multiple kinds of undefined behavior: accidental, and essential. Accidental UB is stupid shit like "if your file doesn't end with a newline, UB occurs". Essential UB is things like, if the memory of a local variable is changed by /proc/mem by another process, while a function is evaluated, UB occurs. Essential UB allows basic, essential optimisations to take place that everyone expects every language & compiler to be able to perform. C has a large amount of accidental UB; Zig has none.
2. A Zig application decides what to do when a safety check triggers by overriding the panic handler. Zig's default panic handler crashes with a helpful stack trace. This is a killer feature.
3. I see a lot of people talking ignorantly about safety critical applications. Let's talk about Level A Clearance. This is software that is licensed to run on airplanes and other safety critical components in the United States. Here's how it works: you have to test every error condition and every branch at the machine code layer. This makes a simpler language such as C or Zig more well-suited than language with hidden control flow such as C++ or Rust, because it causes problems for testing every branch at the machine code level. Furthermore, such components are redundant, so that when one fails, the readings of the others are used. So, crashing or otherwise indicating a faulty reading is absolutely what you want safety-critical software to do, as opposed to giving a well-defined, incorrect reading due to, for example, an integer overflow, which can happen in "safe" Rust.
The abuse of UB is the way how those 1980's C and C++ compilers that generated lousy code easly outperformed by Assembly coders on 8 and 16 bit home computers, finally improved their code generation quality due to the lack of strong type information for the compiler.
So where we are, trying to escape bad decisions from the past.
void Foo(Bar * bar) { if (bar == nullptr) { println("Bar is null!"); } return bar->ComputeFoo(); }
E.g., one pass may say "the println code is unreachable unless a null pointer is dereferenced", and the next may say "the only code that is reachable is `return bar->ComputeFoo()`", so just compile the function as that.
Imagine instead of a println the code is actually some large body of code, and that the function is inlined into another where bar is null. In that case you'd want the compiler to avoid compiling the code, but it can't without those passes.
1) embedded or OS hackers: they want code generation to be predictable and dislike surprising optimizations. To them, undefined behavior is behavior outside the standard specific to certain compilers that they expect to not suddenly change.
2) application and especially HPC developers. They want the compiler to exploit every trick in the literature to improve performance. In return, they are aware that undefined behavior cannot be relied on.
Any compiler that is popular with both crowds would have to strike a compromise.
Of course ideally we'd move away from provenance-based UB and Rust's third-class aliased mutability, to simpler conservative aliasing semantics by default, and allow programmers to opt into non-general optimizations by manually saving unchanging values like vector::size() into locals, or have the optimizer and programmer interactively explore code invariants and optimizations, making the optimizer a performance-focused pair programmer rather than a black box.
void *oldptr = malloc(73);
void *newptr = realloc(oldptr, 42);
if (memcmp(&oldptr, &newptr, sizeof oldptr) == 0) {
// not changed
}
The value of oldptr is indeterminate if it is used as a pointer; it can still be accessed as an array of bytes. void *ptr = malloc(73);
void *old = ptr;
void *ptr = realloc(ptr, 42);
bool realloced = old != ptr;Not comparing the values is the point. Your code uses the pointer that was passed to realloc, and that is undefined behavior according to ISO C.
The game you're playing with the identifiers is pointless; there is no difference between what you're trying to do and just:
void *oldptr = malloc(73);
void *newptr = realloc(oldptr, 42);
bool realloced = oldptr != newptr; // undefined behavior
That's what the article is referring to, and that I'm specifically addressing with the memcmp. Accessing the pointer as an array of bytes doesn't use its value as a pointer. Bytes cannot be indeterminate; they are not allowed to have trap representations.(There could be a false negative: the address didn't change, but the pointer bit pattern did. On 64 bit systems, the C library could easily put a tag into the upper bits of the pointer, and have realloc change the tag even if the address is the same.)
The problem described in the article is that the compiler generated a false negative even when the pointer didn't change, due to the undefined behavior. The idea is something like that since oldptr was passed to realloc, it is garbage. The newptr is good, and we need not compare garbage to non-garbage; we can just declare them to be unequal.
That's what you might get if your follow your advice of "compare the values, so that the compiler knows about it".
void *oldptr = malloc(73);
uintptr_t savedptr = (uintptr_t)oldptr;
void *newptr = realloc(oldptr, 42);
if (savedptr == (uintptr_t)newptr) {
// not changed
}I'm not sure it's also OK, because the memcmp suggestion you are replying to seems a bit suspect.
There's still the problem that comparing the uintptr_t's is not guaranteed to yield the same result as the pointers they were cast from. But that's merely implementation-specific behaviour, not undefined.
(An unsigned type can only have trap patterns if it has padding bits. Every combination of the value bits is a valid value according to the pure binary encoding.)
int *oldptr = (int *)malloc(42);
int *newptr = (int *)realloc(oldptr, 73);
if (oldptr == newptr) {
newptr[70] = 10; // ok because the "object" newptr can contain 73 ints
oldptr[70] = 10; // SIGABRT here because the "object" oldptr can only contain 42 ints even though the memory block oldptr points to can hold 73
}From the previous article linked at the footer of this one:
- https://developers.redhat.com/blog/2021/04/16/broadening-com...
It says that this is available in LLVM
GCC support for __builtin_dynamic_object_size or equivalent functionality is in progress. At the moment this is available only when building applications with LLVM. There are some unspecified corner cases with __builtin_dynamic_object_size that may result in avoidable performance overheads. We hope to iron those out with the GCC implementation and feed it back into LLVM, thus making both implementations consistent and performant.
Does this mean I can set this flag in Clang and it'll work?I can imagine enabling this in programs that directly interact with the internet, e.g., web browsers, email clients, and resolver libraries. I can imagine Debian, Ubuntu, and Red Hat enabling it. It would especially make sense for Tails, at least in some cases. But that doesn't mean it will happen. I'd love to hear more.
There's a bug with the patch to enable it, tracking issues in other packages, so they seem to be in the process.
For folks who prefer reading text using large, complex, graphical browsers released by organisations that seek to profit directly or indirectly from the proliferation of online advertising^1
curl https://developers.redhat.com/articles/2022/09/17/gccs-new-fortification-level|sed -n '/./{/article-content/,/<\/div>/{/<\/div>/,$d;p;};}' > 1.html
firefox ./1.html
1. Apple, Microsoft, Mozilla, Google, Brave, etc."The most surprising new adman is Apple. The iPhone-maker used to rail against intrusive digital advertising. Now it sells many ads of its own. As sales of smartphones plateau, the company is looking for new ways to monetise the 1.8bn devices, from smartphones to smart earphones, it already has in circulation. So far it is only dabbling in ads and does not report sales figures. But Bloomberg reported recently that Apple's ad business was already generating sales of $4bn a year, making it about as big an ad platform as Twitter. Apple executives believe there is much more to be had."
FROM quay.io/fedora/fedora:38
RUN dnf install -y <bunch of tools>
Got any useful tips or flags to enable that're bleeding edge?Of course if you only want to build your application with _FORTIFY_SOURCE=3, you can do it right away even on Fedora 36. The Fedora change will be to build the distribution (or at least a subset of packages) with _FORGIFY_SOURCE=3.
https://stuff.mit.edu/afs/sipb/project/bounds/src/gcc-2.7.2/...
I played with that; it worked.
In the 90's I reached for Bruce Perens' Electric Fence; that did a good job for me.
We've Valgrind for some two decades now?
Call me unexcited ...
What we miss is people actually using these tools in any meaningful way.
Instead we have to go for hardware memory tagging and sandboxing, because there is no advocacy that changes their ways.