Interestingly enough, if you were to incorrectly use stack memory in this scenario, the magic checks should trigger, pointing out improper memory usage.
Interestingly enough, if you were to incorrectly use stack memory in this scenario, the magic checks should trigger, pointing out improper memory usage.
I'm pretty sure you can get similar optimization behavior when you mark functions with special attributes, I'm not 100% sure on that point though. So for the Linux Kernel I'm not sure that kind of optimization would ever be done since obviously it's not using the standard defined functions and the compiler might not be given enough information to know the functions have the same semantics as 'malloc()' and 'free()'.
I personally was able to get the exact behavior described here by compiling the below code using `-O2`. The volatile write just ensures the first constant write and the malloc itself cannot be optimized out (and that does happen! Take out the volatile and there are no calls to malloc in the result), but gcc is still free to do whatever it wants with the other write and for me it's completely gone in the resulting program.
int main()
{
int *p = malloc(sizeof(*p));
*(volatile int *)p = 20;
printf("p=%d\n", *p);
*p = 30; /* this is gone from the -O2 compiled code */
free(p);
return 0;
}
The relevant part of the assembly looks like this, I don't think I'm missing anything in that the 30 assignment is completely gone: push $0x4
call 460 <malloc@plt>
movl $0x14,(%eax) # Assignment of 20
mov %eax,%esi
pop %eax
lea -0x1930(%ebx),%eax
pop %edx
pushl (%esi) # push the 20
push %eax
call 440 <printf@plt>
mov %esi,(%esp) # load address to pass to free(), no assignment between printf() and free()
call 450 <free@plt>Edit note: allocators are user space libraries (stdlib) and not part of the C spec. Use-after-free is extremely unsafe, however its completely valid C.
2nd note: running all programs below give me the expected result, writes to pointers are live, regardless if they are freed. So please provide more concrete steps to reproduce your results.
I'm not sure why you think this, but the C standard includes a whole section defining the behavior of the standard library functions, including malloc and free. And it includes this note as undefined behavior:
> The value of a pointer that refers to space deallocated by a call to the free or realloc function is used (7.20.3)
So it is undefined behavior to access memory passed to free, or IE use-after-free is undefined behavior if you're using the standard library.
> 2nd note: running all programs below give me the expected result, writes to pointers are live, regardless if they are freed. So please provide more concrete steps to reproduce your results.
I would check that you're compiling them with `-O2`, and also check the assembly output. However as someone linked, you can already see from the online compiler output that clearly gcc is capable of optimizing the assignment out. Here's my second program in the same compiler, notice how printf is called twice but the 30 assignment (which is what is supposed to set your magic number) is still completely gone: https://godbolt.org/z/9reErcbGK
Edit: Sorry, I previously included the wrong quote from the standard (it was about a double-free being UB, slightly different), I have the right one now. I got it from here if you want to look at it, it's a draft but practically the same as the actual one: http://www.open-std.org/jtc1/sc22/wg14/www/docs/n1256.pdf
memset_s is the only function that would be defined to work precisely because this was needed for crypto so it is solvable but just take it on faith that if it’s not explicitly called out as forcing the compiler to not do dead store elimination, it will happen.
I think that's your best bet but volatile does bring its own set of problems with it (how it works differs from compiler to compiler). That said it should probably work fine, this is a pretty simple use of volatile. You might check the documentation for your particular compiler though if you have one in mind, it should tell you what it does and if it does anything undesirable (or if there's a better way to do this). Also the volatile cast like I did may be a better approach than actually making the member itself volatile.
FWIW, the whole idea is that whether this happens or not has no impact on your program, so in theory you shouldn't ever notice this happens. Detecting use-after-free like this is not really standards compliant so that's a big reason why it's problematic to implement.
Also, if you're not using the standard allcator then most of this logic doesn't apply because the compiler won't know your special allocator has the semantics of 'free()'. There are `malloc` attributes in gcc that might trigger similar optimization behavior, but you'd have to be using them, and even then I'm not really sure as I haven't looked into what all they do.
Trying to structure code to trick the compiler is a bad idea. The compiler authors know the standard better than you and eventually the compiler will exploit your misunderstanding.
I think there's been godbolt links posted in this thread showing this assumption is wrong. Dead store elimination applies just fine to volatiles.
As context, the Linux Kernel uses volatile to ensure loads and stores happen, that's ultimately how READ_ONCE and WRITE_ONCE work[1]. If that's actually broken in such a simple case I think they'd like to know xD
[0]: https://gcc.gnu.org/onlinedocs/gcc/Volatiles.html
[1]: https://elixir.bootlin.com/linux/latest/source/tools/include...
Edit: To be clear, I looked for the example you mentioned but couldn't find it. I'm somewhat wondering if you were thinking of the example I posted, since I used volatile to get gcc to not optimize the store out :P
I’m genuinely amazed at the response. There’s literally an API defined that has the contract you want and your response is “yeah, but I want to write it a totally other way the standard doesn’t allow”. Just use memset_s. It’s a compiler builtin so the generated code is as efficient (more so) as compared with a volatile version except actually safe. Volatile has a totally different purpose and isn’t suitable to try to write a value before calling free.
I’ll leave writing a godbolt example of writing to a volatile right before a free in the same compilation unit at O3 for you to try out.
[1] http://www.daemonology.net/blog/2014-09-04-how-to-zero-a-buf...
As far as that article goes, the example for `secure_memzero` works and you will not find any compiler that will 'optimize that out', it would be a bug. And as I linked, the gcc documentation says as much. With that, memory allocation is not as special as you're making it out to be, normal memory can be volatile in perfectly valid situations (even ones mentioned in the standard), and just because it's related to a free() does not mean the compiler is now allowed to remove a volatile dead store - and even if you think it does, gcc will not do that.
Here's an example of such a case[0]. A signal handler is able to view the object being set right before the free() call, and a signal could trigger at that point, but the compiler still optimizes it out (which is correct). Using volatile on the variable to ensure all loads and stores actually happen (and are visible to the signal handler) is the suggested way, and if you do that then the code does set the value before the free().
As for your suggestion of writing to a volatile right before a free(), I'm not sure if you tried but it works just fine as expected, look[1]. I am perfectly confident in saying you will never find an example where the volatile store doesn't happen. With that, if it was willing to make such an optimization in the first place, don't you think my original example that used it to avoid dead store elimination and memory allocation elimination wouldn't have worked in the first place? ;)
[0]: https://godbolt.org/z/WWPz5Gqjo [1]: https://godbolt.org/z/anc1cfnPs
volatile bool doCheck = false;
if (doCheck)
{
// code I want to enable at some point during debugging
}
The idea is that I attach a debugger, and then only at a certain point enable doCheck.I was baffled to learn that MSVC will happily constant-fold the false into the if, as long as the variable is function-local. The variable still exists and I can change it in the debugger, but it doesn't actually impact control flow as intended. The "solution" is to move it to e.g. global scope (this is a debugging hack, remember).
Not an exact match for what you asked, but I think a good reminder that optimizers work in mysterious ways, and sprinkling in volatile may confuse the programmer more than the optimizer...
To show this even more, if you add an extra printf to print the value of p after the `free()` (so add a literal use-after-free to check the 'magic' value) the 30 assignment is still gone. It prints zero for me because free() clears that memory, but the assembly does not include the 30 assignment even though I'm clearly reading the value it assigned after the `free()` statement, which is exactly the behavior you're attempting to catch. If `free()` didn't touch the value I'd still see 20.
With a little finessing I got my code to print 20 both times (the larger struct gets malloc() to leave the magic value alone) even with the 30 assignment still in the code:
struct foo {
char bar[35];
int p;
};
int main()
{
struct foo *p = malloc(sizeof(*p));
*(volatile int *)(&p->p) = 20;
printf("p=%d\n", p->p);
p->p = 30;
free(p);
printf("p=%d\n", p->p); /* Prints 20 for me, the 30 assignment is optimized out completely */
return 0;
}If you have the function
void free_ws(struct ws *ptr) {
ptr->magic = 0xdeadc0de;
free(ptr);
}
there is no possible circumstance, in valid C, where ptr->magic could be read and have its value equal to 0xdeadc0de. That object is freed immediately after that write.If you do a read in some other function, say
char read_from_ws(struct ws *ptr) {
if (ptr->magic != WS_MAGIC) {
exit(1);
}
return ptr->s[0];
}
it is impossible for the pointer passed into read_from_ws to be a pointer that has passed through free_ws, i.e., it is impossible for ptr->magic to have been set to 0xdeadc0de by free_ws. Therefore, free_ws doesn't need to actually do the write.You're right that the correct magic check is not dead, and cannot be eliminated by the same logic. But the effect there is that the magic check always succeeds!