Being Sneaky in C
codersnotes.com
codersnotes.com
"Also in his Turing Award lecture, he described how he had incorporated a backdoor security hole in the original UNIX C compiler. To do this, the C compiler recognized when it was recompiling itself and the UNIX login program. When it recompiled itself, it modified the compiler so the compiler backdoor was included. When it recompiled the UNIX login program, the login program would allow Thompson to always be able to log in using a fixed set of credentials."
I'd be surprised if these changes haven't made it into FreeBSD, but afaik Linux doesn't work this way (by default, anyway).
And of course, some programs use replacements like dlmalloc and do all their own allocation management anyway.
Yeah. I wrote my own allocator in C++ a long time ago. I wouldn't be surprised if there weren't quite a few other bits of software out there doing the same thing.
Either way, "don't write your own allocator" is a good lesson to learn.
Unless, of course, you're doing it for fun. In which case, efficient heap management really is a neat exercise.
Do you have a URL to a git ... hahaha, pardon me, CVS repo?
And some context around malloc and the various techniques OpenBSD implements to make exploitation of bugs more difficult: http://www.openbsd.org/papers/dev-sw-hostile-env.html
The tl;dr is : if your code compiles and runs on OpenBSD, chances are it will fine on other Unices. The opposite is not necessarily true though.
I would argue every language has that property. But with C/C++ being so closely tied to the ABI of the machine perhaps they are more underhanded than others. But to me, this branding does feel a bit unfair.
Still, a fun contest and an interesting read.
Do not attempt to write secure software with C / C++. It is a Bad Idea(TM). Because what you wrote is not what gets run.
For instance: there is no way to write a portable secure memset in portable C / C++. (You can write a "secure" memset that works in all current compilers, but that is not the same thing. What doesn't get optimized now can and will be optimized tomorrow.)
I think that C or C++ could, without too much effort, support semantics that would allow for this sort of thing. Something as simple as a "secure" keyword that could be applied to variables (where it means "leak as little as possible about this variable when it goes out of scope") or functions (where it means the same, but for the function itself and all locals of the function).
It's dependent on the function attributes of memset (e.g. __attribute__((pure))) - it won't always be optimized out.
https://www.cs.auckland.ac.nz/references/c/gcc4.7/Function-A...
The compiler is allowed to optimize away that function call regardless of if it is memset or your own alternative.
Note that that is not the same thing as saying if it does currently optimize away that function call.
Really, you can print a c program on piece of paper and ask some slave to "execute" the program in his head given some input x, how he "implements" memset will surely be different than what a computer would, and if you only ask for the output y he will surely see that this memset doesn't affect y at all and skip doing it.
There is no way to ensure that something is actually overwritten, because under the memory models of C and C++ you cannot ever read that memory again, even though in actuality you can.
I mean it when I say you cannot.
Edit: unless your point is a temporary copy can be spilled in memory and this copy will stay in memory and won't be overwritten?
Alternatively, use a library that the C compiler doesn't know so much about that it will attempt to remove calling functions in them.
If you copy your standard library's memset in a separate DLL that is not the standard C library, the compiler will not even see the code during compilation, so it has to compile a function call.
The linker (or a JIT in your C runtime) is allowed to remove calls to the function, if it can prove that it doesn't have side effects. However, to prove that, it has to look at the assembly of the function; it cannot use the way simpler heuristic "it came from <memory.h> and is called memset"
Although I'd like things to behave this way, I don't think this is true. The C standard library was incorporated into the language spec for C89. The behaviors of the named functions within it are specified, and the compiler is allowed to inline it's own version (ignoring your custom code) and then optimize out the inlined portion.
So while it's possible that the external linkage approach still works with certain compilers, it's not portable. I believe you are OK with the external approach if you use a non-standard name (my_secure_memset_pretty_please()), but that just shifts the problem to forcing the compiler to generate your external function without making the same dangerous optimizations.
In the end, I fear you are left with three options: blind faith, non-standard language extensions, or switching to a more secure language (likely assembly). If there are other options, I love to hear about them.
(IoT security is doom for other reasons though, mostly UI, updatability and cloud services)
And then that code you wrote now silently becomes deadly a few years down the line.
(That said, I think TheLoneWolfling is being too strong with his/her claims. You can get modern compilers to avoid dangerous optimizations; it's just not for the faint of heart.)
I am saying that it's impossible to do so and remain in the realm of portable C / C++.
There is a distinction.
And as for the second part... Meh. I don't see any optimizations that hard-coding calling something named "memcpy" (or whatever) does that cannot be enabled by looking at the actual code that gets linked. Albeit with more difficulty.
Remember: there is nothing that specifies that C / C++ needs to be compiled.
Also: you could have said the same twenty years ago about many of the optimizations that currently compilers do.
W.r.t. 1, the compiler's definition of no optimization today is not the same thing as it was last version, or will be next version. For instance, on IA-64 there are things the compiler has to do that are typically considered optimizations.
W.r.t. 2, you have to make sure there is no link-time optimization happening.
However, that is not an inherent restriction - that is only a restriction on current compilers. It is entirely possible for a compiler to read the assembly of things being linked and optimize based on that.
For instance, when someone runs it in an emulator for backwards compatibility purposes. Or when someone runs it in a JITter. Or even just if the compiler decides to special-case for the existing link target.
* Use a specific compiler and verify.
* Don't use C / C++.
* Panic.
> memset may be optimized away (under the as-if rules) if the object modified by this function is not accessed again for the rest of its lifetime. For that reason, this function cannot be used to scrub memory (e.g. to fill an array that stored a password with zeroes). This optimization is prohibited for memset_s: it is guaranteed to perform the memory write.
The compiler can and will copy things around, and it is not required to memset_s said cop(y)/(ies) away.
Of course there is, you just use the volatile keyword. volatile guarantees that all read/writes have corresponding memory accesses and cannot be optimized away.
It's not going to be as fast as memset but it's definitely portable and it won't be THAT slow. Then for platforms that have memset_s defer to that instead, otherwise fallback to the totally portable volatile + for loop.
Colin Percival of FreeBSD disagrees:
http://www.daemonology.net/blog/2014-09-04-how-to-zero-a-buf...
Namely that C / C++ allows temporary copies of variables that are not cleared afterwards. The most obvious case of this being things being temporarily copied into registers / stack, but there are other examples as well.
But that does not work. Full stop. The compiler can optimize in ways that still leak the contents of the thing that was supposed to be memset-ted away.
But I'm pretty sure you're just aggressively anti-C/C++ so whatever.
C / C++ are very good languages in all sorts of ways. However, there are components that currently have... flaws. This being one of them. As such, I complain about said flaws, in the hopes that someone will take notice, and/or someone will point me in the direction of things that contain the good parts of C / C++ without said flaws.
I have already learned a fair bit about bounds checking, SIMD instructions, etc, etc from this. And I always want to know more.
*
And no, it is the same problem. Namely, that the memory models of C and C++ doesn't match with the underlying hardware, and the mismatch is such that things that are trivial to do on the underlying hardware are literally impossible to do with C and C++.
Part of this is for compatibility purposes, but there are ways to keep the compatibility that don't present this sort of problem.
The only place that keywork should be used is a qualifier for member functions or when used in an embedded sense. It's not well defined outside of that scope.
There are some ISAs that pretty much require optimizations, for instance.
The compiler can (and will!) just propagate the values through directly and skip the memset.
I think some systems have a "secure memset" function that can be used for things like this - i.e. one that's guaranteed not to be optimized out.
The compiler can, and will, make copies of data behind the scenes. And not erase said copies.
What we really need is a keyword / modifier that says that when X passes out of scope no state related to X may be leaked. Ideally, that can be applied to a function / block as well as a variable.
(Or rather, not necessarily no state. Read "as little state as possible", preferably with modifiers that panic unless the compiler can ensure specific things.)
On the other hand, something as simple as a keyword marking a variable as "as secure as possible given hardware constraints (read: should wipe any temporary copies and the variable itself after it goes out of scope, should attempt to prevent it from being written to non-volatile storage, that sort of thing)" (sort of like how inline works), with compilers required to bail if the constraint cannot be done to the level specified, would be a massive step in the right direction.
There is, quite literally, no way to ensure data is not leaked (namely, that data is zeroed / etc) in portable C / C++.
With C it's far too easy to get the data back.
Also, you missed the worst example: CPU cache.
The compiler would just optimize that set right back out.
memcpy(filter->buffer, output->piu_text_utf8, sizeof(output->piu_text_utf8));
1. memcpy is less safe than memmove and strncpy. strncpy should be used.2. The two character arrays should use the same constant in defining their length, and that constant should be used both in the struct definitions and here in the copy operation.
3. The code is written in C in spite of it being 2014 at the time.