memset tend to be optimized out in the dead code elimination pass and highly relies on SSA. Compilers adhere to the abstract machine specs. If you want memset not to be optimized out then use memset_s because the spec explicitly doesn't allow that.
if an optimization relies on a UB and surprised the programmer then the programmer was also relying on a UB in the first place.
The semantics of memset and memset_s are defined by a standard. They're not unknown functions that could do anything, and they're breaking a rule by treating them differently, like you're suggesting they are.
Optimizing memset, in particular, can reap huge benefits, for example for sizes of 4 or 8 bytes in tight loops or for zeroing cache line sized blocks.
Compilers do have an intrinsic equivalent but that has nothing to do with whether it gets optimized out or kept.
Following SSA rules, if you are writing to a non-volatile memory and not reading from it for the remaining duration of its lifetime then it's dead code.
use memset_s was a suggestion if your goal is to scrub memory because it guarantees the write will happen. The guarantee comes from the C standard:
> Unlike memset, any call to the memset_s function shall be evaluated strictly according to the rules of the abstract machine as described in (5.1.2.3). That is, any call to the memset_s function shall assume that the memory indicated by s and n may be accessible in the future and thus must contain the values indicated by c
Which essentially means treat the memory as-if it's volatile.
If you pass a pointer to a volatile or atomic object to another function the compiler cannot see into, the compiler knows that any specified operations might be observable, so must act accordingly. Then, operations on that object will not be optimized away. Or, anyway, not until someone builds the program using Link Time Optimization ("-flto") and the called function body becomes visible.
Calling C11's memset_s, or MS's SecureZeroMemory, or FreeBSD's explicit_bzero, has the effect, to the compiler, of making the object observable, so that operations on it will not be optimized away. There is no way to express this locally within the language; you must rely on facilities provided by a Standard or your platform, or on the dodgy linkage trick cited elsewhere.
However, memset's declaration takes a void* which prevents you from passing a volatile void*. casting or not won't change anything in this case because what matters is what memset takes and not what the pointer's fully qualified type is prior to calling memset.
From the compiler's perspective, you're passing a non-volatile pointer to memset which makes it a candidate for elimination if the pointer's lifetime ends after memset call.
[1]: https://lore.kernel.org/lkml/Pine.LNX.4.64.0607060856080.124...
Many people mentally model atomic as akin to volatile, about which they also frequently harbor superstitions.
I knew a FreeBSD core committer who insisted (and convinced our boss!) that volatile was entirely sufficient to serialize interactions between threads.
Someone else expected a line with a volatile or atomic operation never to be optimized away, and was appalled that it happened. Assignment to a volatile stack object whose address has not escaped and is not used will, in fact, be optimized away, regardless of the "clear language of the Standard".
Often, not performing what would appear to be a pointless optimization would block a more substantial optimization. Compiler writers optimize aggressively in order to break through to the substantial ones, which are most usually dead-code elimination that can then propagate backward and upward.
If you really need for a volatile object and operations on it not to be optimized away, you need to allow a pointer to it to escape to somewhere the compiler cannot see. Have a function in a separate ".o" file, and pass the pointer to that. Or, better, use a compiler intrinsic or other facility provided specifically for the purpose, such as memset_s or explicit_bzero.
> Assignment to a volatile stack object whose address has not escaped and is not used will, in fact, be optimized away
I don't believe either of these statements. https://godbolt.org/z/8cqEdY9oE is a simple case where this doesn't happen. Can you post a Godbolt link or something that demonstrates when it does?
Keep in mind that the standard says "Accesses to volatile objects are evaluated strictly according to the rules of the abstract machine." In particular, it doesn't say that accesses to non-volatile objects through volatile-qualified pointers will be handled in any special way. (This looks like it's going to be changed in C2x, but it hasn't been changed as of C17.) Also, it says "If an attempt is made to refer to an object defined with a volatile-qualified type through use of an lvalue with non-volatile-qualified type, the behavior is undefined."
Talk to any core compiler implementer, and they will say that the optimization is allowed. No compiler performs all permitted optimizations, in any release, but the set of those each does do, varies.
Can you provide a different program, compiler/version, and command line where it will be optimized away?
> Talk to any core compiler implementer, and they will say that the optimization is allowed.
Are you just guessing that they'll say this, or have they actually said this? If the latter, can you link to where they did?
> No compiler performs all permitted optimizations, in any release, but the set of those each does do, varies.
My Godbolt link is a really obvious, easy place that the optimization you claim is possible could be performed. I think the fact that none of the 4 major compilers optimized it away is decent evidence that it isn't allowed.
They have said it face-to-face, in person. It is possible that, with enough pushback, they have since disabled the optimization. (They all work at large corporations, nowadays, and must obey management direction.) Or, more likely, disabled certain cases of it. I will ask again.
For documentary evidence, you may consult the definition of C11 memset_s, which is specified not to permit the optimization. It is hard to imagine such a requirement being perceived as needed, absent cases where calls to memset have actually been elided.
I will also experiment further with Godbolt. You might have failed to tickle the compilers in just the right way.
If you do ask again, can you do so somewhere like a public mailing list, so I can read their response myself?
> For documentary evidence, you may consult the definition of C11 memset_s, which is specified not to permit the optimization. It is hard to imagine such a requirement being perceived as needed, absent cases where calls to memset have actually been elided.
The point of memset_s is to prohibit the optimization when the object isn't volatile.
99% of C/C++ programmers are not bored language lawyers with time to memorize thousand page specs. They use the language because they're pragmatic. Messing up peoples code with incredibly over-eager optimization is the opposite of pragmatism. I'd rather spend a day writing inline assembly for a hot path than two days dealing with the compiler breaking my code and making users data insecure becuase "usually" its ok to ignore a memset.
Just because you don't understand something fully doesn't mean it's not common sense.
It's very common that programmers don't bother reading the standard or attempt to understand the concept of an abstract machine.
Programmers that know what they are doing tend not to be surprised by such nuances.
"In theory, theory and practice are the same. In practice, they are not."
C11 and C17 are supported, minus C99 features made optional in C11.
ASan is now in the box.
And Visual Studio definitly provides a better experience of coupling static analysis with code being written.
Second, Microsoft considered C done on Windows, and green field systems programming should be done in C++, there were a couple of blog posts done about the subject, which apparently hater kept forgeting to read about.
So no, it wasn't about when it was convenient, rather to the extent required by ISO C++ compliance.
In fact, as per Microsoft Security Center advisory, C# or other managed languages, Rust or C++ with Core Guidelines, should be the main languages to use for new applications.
Most likely they ended up backtracking with the goal to drop C support from Visual Studio due to pressure from high profile customers, WSL and FOSS code on Windows, or even the culture change brought by Satya.
This is true, but an OS with no C compiler is horrifically bad. If Microsoft didn't want their C++ compiler to also be a C compiler, then they should have made a separate C compiler, instead of forcing people to rely on something like MinGW.
There is hardly any reason to keep writing new C code in 2021 other than UNIX like codebases or embedded development that keeps resisting C++ adoption.
I was actually disappointed with Microsoft's decision to backtrack on their decision.
Was there a separate C compiler? I thought Microsoft had a C++-only compiler but not a corresponding C-only compiler, and that the lack of a separate C compiler was the problem.
> There is hardly any reason to keep writing new C code in 2021 other than UNIX like codebases or embedded development that keeps resisting C++ adoption.
> I was actually disappointed with Microsoft's decision to backtrack on their decision.
You seem to be saying that C is deprecated or something, but it isn't.
C was deprecated from Microsoft point of view regarding systems programming on Windows, apparently complainer keep forgetting to actually read Microsoft employees blog posts.
Here, have a go at them, to spare you some Google-fu.
https://herbsutter.com/2012/05/03/reader-qa-what-about-vc-an...
https://msrc-blog.microsoft.com/2019/07/18/we-need-a-safer-s...
But alas, some manager decided to backtrack on the decision and support C after all.
BTW, go look at how many optimizations are explicitly disabled to compile the kernel without it breaking and then come back to me with how these are a good idea. One of the most critical pieces of software infrastructure in the world looks at how broken compilers are and says yeah, leave that junk out.
By "working code", do you mean code that the standard and/or target implementation promise will do what you want, or do you mean code that just happened to do what you want the last time you tested it on your system? If you mean the latter, then yes.
If you are stuck with UB, would you prefer a compiler that does the safe thing, or a compiler that unleashes dragons?
Maybe we should focus on better warnings, static analyzers, and UBSan, so that people with UB can fix their code, rather than not letting compilers optimize anymore.
> If you are stuck with UB, would you prefer a compiler that does the safe thing, or a compiler that unleashes dragons?
I'd very much prefer that my compiler unleashes dragons, so that I notice the bugs and can fix them. If my compiler did what you want, then when someone else compiled my code with a different compiler, they'd find them instead. (And it's obviously totally unrealistic to expect that literally every compiler in the world would do what you're asking.)
This isn't possible, since "breaking user programs" in this context means "violating the intuition of the programmer", who is an unknown person who has never written their intuition down to be observed by the compiler authors.
If you are actually talking about breaking user programs, both C and C++ already have ABI stability as a design goal, which hugely limits changes to the languages but enables people to keep linking against binaries that were compiled ages ago.
I believe MSVC developers do put more effort than GCC/Clang developers to maintain source-level backwards compatibility. Alternative development strategies are certainly possible, but I understand that the GCC/Clang developers do face pressure from a significant subset of users to maximize performance at any cost.
This wouldn't just be expending effort. It would be making all code GCC/Clang ever compiled be slower.
> Alternatively, they could disable the relevant optimizations at -O2 (and maybe -O3) level.
Ditto.
The sane thing in a realistic scenario is to ban all usage of memset and add a diagnostic to use memset_s, like any responsible programmer should do with strcpy. Then however, was it really wise to make this machinery, that the people who need it the most are most likely to not have, necessary? I really, really don't think so.
\Edit: Here is some code written by a very smart person, Niels Fugerson. I think (not sure) it tries to wipe sensive material from memory and gets it wrong. Really not sure, tho.
https://github.com/MacPass/KeePassKit/blob/master/TwoFish/tw...
It demonstrates that there are many different meanings for "smart", and even for "smart programmer". Nobody qualifies for all of them.
Compilers won't optimize out memset if it changes the program's behavior within the boundaries of abstract machine's specification.
A call to memset to clear a memory region after you're done with it because a password was stored there can be optimized out because it doesn't change the behavior of the program.
The code you shared is incorrect if your definition of correctness here includes K being wiped out.
Well, that does change the behavior of the program. After all, the behavior of the program was to wipe the memory. The programmer specified it. The compiler is clearly wrong by any reasonable measure. It's removing an intentional, clearly specified instruction. It's just wrong. It's making an "interpretation" of something the programmer intended as a fact, and breaking things in the process. Why does anyone think this is desirable? Any good programmer knows that correctness comes before speed.
I've been programming C++ since 1998 and I didn't even know this till reading this thread. Blaming the user here is counterproductive. Compilers should not make this optimization, plain and simple. memset should set the memory. If you think otherwise, I invite you to take responsibility for all the users data you have placed in harms way.
This is simply not accurate. memset_s provides that guarantee. memset does not. It would appear that you have been under the impression that it did. However, it did not.
There's a big difference between you not liking something and it being wrong. You want to make the common case slower for everyone so that you don't have to learn one extra rule for your uncommon case.
> it strikes me as incredibly fucked that you need a separate memset_s to mean "I'm not fucking around". When I learned C, memset just meant, uh, set this memory, it was pretty straightforward.
I don't think it's ever worked the way you want it to work.
I personally prefer to write code that is more readable to humans that will get translated to the optimal instructions over having to learn some tricks known only by some elite C hacker that will actually be inferior, since CPUs are finicky. Even with very good understanding of the CPU one can not really guess which instructions are faster to execute without benchmarks.
Nonetheless, there are tiny C compilers that do next to no optimizations, feel free to compare the speed of the resulting binaries.
The compiler cannot read minds, if the programmer clearly thought the code should behave in a certain way it should have told the compiler in a language it understand. Volatile would probably be part of that solution.
BTW, I think memset_s is also problematic: there is nothing preventing a compiler from storing a copy of your data anywhere else. At the very least is likely to be in registers, possibly partially spilled on the stack.
It can't, so don't remove my memset. Why is this hard. Compilers are dumb, yes, so don't delete the thing I wrote on purpose.
You are right about the definition of memset. It's unreliable. My point is why is that ok with anyone sane. If I clear 128 bytes on a stack frame that's returning, sure, it's not optimal, but the compiler doesn't know why I did it and I surely did it for a reason. The definition of memset is bad.
I don't use memset much, but I use memcpy and certainly do want (and expect) the compiler convert memcpys to register moves when possible and even to eliminate them when the result is not actually used.
1) Yes, I know there are rules and all that are precise and allow that.
When people write C, they often do so for specific domains. Crypto, embedded, device drivers. They choose C for these domains because they think, and they are taught that C is some kind of macro assembler, because in these domains they care deeply about the "how". If they wouldn't, they would write java or any current top 20 programming languages. So almost all optimization off looks ok. Reasonable compared to this, anyhow.
If people truly still think C is a macro assembler then that's the problem. Besides if you care deeply about how modern high performance architectures won't cut it anyway.
Compiler writers are stuck in a trap. C++ compilation model is broken. Which makes compiling C++ programs very slow. Which means compiler writers care a lot about how fast compilers are. So they desperately add more optimizations to speed up the compiler program. But which adds more work to the process of compiling a program.
Most of the other users, especially C programmers don't care about speed nearly as much. Notably because while the compilation model for C is also broken, it's not nearly so. So their programs compile in a few seconds, not minutes to hours.
If they really thought that speed was the most important thing wouldn't that abandon those languages for C++.
Explain yourself.
FreePascal only does safe optimizations. (unless it has bugs where it does work at all)
So your argument is "people not using C don't care about the speed of C"? What is this supposed to be proving?
> compiler writers care a lot about how fast compilers are. So they desperately add more optimizations to speed up the compiler program. But which adds more work to the process of compiling a program.
Compiler writers are stuck. The more optimizations they add. The more work the compiler has to do. And the slower it runs. Defeating the purpose of those optimizations.
If you write other types of programs the compilers speed and your programs speed aren't the same at all. People that write other types of programs care more how fast the compiler is and less how fast their programs run. People that write programs that process entrusted data care vastly more about no surprises than they care about speed. Even more so, in modern code bases the hot sections are very small parts of the total code base. For 99% of the code, CPU speed doesn't matter.
The whole thing is made worse because of the disconnect between modern CPU's and machine implemented by C++.
Literally everything you ask for here is fulfilled by not turning on compiler optimizations. That's the default. I don't understand what your complaint is.
And since nobody stated it plainly to you yet: This whole world view is completely bogus. Entire branches of the industry require C++ compilers that produce the fastest possible code (to name a few: Games, HFT, image processing). I don't know how you convinced yourself that compiler writers are the only consumers of compiler optimizations, but it's dead wrong. I don't claim your experience or what tradeoffs you seek in a compiler are wrong, but they are extremely far from universal.
> How likely do you think is it that a program that contains memset and has this optimisation applied how diverts from the indented behaviour and now has a serious flaw in it?
What if this optimization happens after three levels of inlining and some other dead code elimination? Is that "diverting from intended behavior"?
Almost every time.
Yeah, the times when it leads to problems are more important individually. And I have no idea of the overall picture. But looking purely at the frequency, those you enumerate basically never happen.
C really needs a "I mean this!" marker that one can use on those cases. Rust also would gain from some kind of predictable region similar to unsafe, but C would have to gain it first.
The article is about C++ which people choose because they want to go fast. Which is also the dominant reason people choose C (embedded & device drivers both absolutely care about performance & efficiency, too, after all).
You can't have both "C is fast" and also "compilers can't optimize." That's not how that works. Similarly if you're just starting out in crypto and you're reaching for C you're already in a bad place and atomic optimizations are the least of your concerns. C is a terrible language for crypto.
Maybe people using Java or various interpretted languages are fine with that, but most C programmers want their code to be a very close mapping to what they specified.
The flag that may allow things to break is -Ofast.
Eliminating dead stores (including calls to memcpy whose result is never read) seems like one of the most basic and least objectionable optimizations I can imagine. So, if you’re against eliminating dead stores, what optimizations _are_ you okay with?
It has no idea! Which is why it should settle down.
What you seem to want is "do all the optimizations, except the ones that contradict what I 'obviously' mean", which is just not possible with current technology.
Even that isn't enough, unless you go with a CPU that doesn't have any branch prediction, speculative execution, or out-of-order execution. I'm not aware of any such processors.