How to zero a buffer
daemonology.net
daemonology.net
If you don't use an the result of some logic it will be optimized out. One way to prevent this is to route it to a pin.
If logic is fed by a constant, it will be optimized out right up to the point where the result of the logic is mixed with some external input. (early tools could not use the dedicated reset net due to this- reset for each flip-flop had to be routed to a pin or the reset net was optimized out which means the initial state of your flip-flop is lost).
If you have identical logic, one copy is optimized out due to aggressive CSE. This is often bad for performance (routing in an FPGA is as slow as logic, so it's better to regenerate identical results in multiple places), so you add "syn_maxfan" constraints to prevent the "optimization".
On the other hand, an input flip flop will be duplicated if the fanout limit is exceeded- but this prevents the use of the dedicated I/O cell flip flop which then causes external timing to be messed up. So you use syn_maxfan=infinite for this case.
I presume this is for some incremental development work. Like testing and seeing the number of gates/pins used for a design but you need to feed the logic with constant placeholders.
static void * (* const volatile memset_ptr)(void *, int, size_t) = memset;
I've written some C but that is utter gibberish to me.The "void *" before the first parenthesis is the return type of the function. The stuff in the first parenthesis applies to the function pointer variable. The second parenthesis lists the types of the arguments to memset.
http://cdecl.ridiculousfish.com/?q=static+void+*+%28*+const+...
Learning the right-left rule helps here. You'll still need to know what the keywords mean.
So, first, the name:
memset_ptr
Then, try going to the right, but aha! before we can really gain speed, a parenthesis is immediately blocking us. We shrug off the bruises and turn to the left: (* const volatile memset_ptr)
"Hey guys, memset_ptr is a volatile const pointer..." Now, we hit left parenthesis, so we're again allowed to go right, yay!... (* const volatile memset_ptr)(void *, int, size_t)
"...to a function taking such-and-such arguments..." Uh, oh, the equal sign, so no more to read to the right; disappointed, we turn back to the left for the final run: static void * (* const volatile memset_ptr)(void *, int, size_t)
"...returning a pointer to void! Hah, got you! Simple, really. No arrays, no pointers to pointers, not even a function pointer returning a function pointer, meh. Uh, oh, aaaaand, yes, by the way, the variable is static, so, like, file-local, um. Yeah, yeah, I saw it from the beginning, oh, go away, you're just picky. And, and, you wouldn't recognize a function returning a pointer to an array of pointers to functions returning anonymous struct even if it hit you in the face, pfff!"Some of this may be technically incorrect. This is my own mental model of the C language which is sometimes incomplete.
Um; then there's the equal sign, so this is not only a declaration, but a definition too; but definitely not a call.
A call is further down in the original blogpost, in the below line:
(memset_ptr)(p, 0, len); asm ("" : : "m" (&key));
just before or after the memset, effectively telling the compiler that the address of "key" escapes the scope of the function.The best answer is really GCC's __asm__("" : : "m" (&key)), or perhaps something like __asm__("" : : "r" (key) : "memory"), after the memset. It generates no extra code, just ensures that the memset won't be removed.
For other compilers (in practice only MSVC, since clang is gcc-compatible), you could pass the pointer to a dummy assembly function instead of using inline assembly; even link-time optimization can't know what happens within a function written in assembly. Or, for better performance, create in assembly a "safer_memset" which is a single instruction: a jump to the real memset function.
This might cause side effects if /dev/null does not exist or is not the null device.
You don't even need to assume some sort of crazy evil compiler to have to worry about this - speculative inlining of function pointers guarded by a safety check is something that FDO builds will actually do.
(memset_ptr)(p, 0, len);
> can be replaced by: if (memset_ptr == memset) {
memset(p, 0, len);
} else {
memset_ptr(p, 0, len);
}
> Which in turn can be optimized using the other tricks noticed above into: if (memset_ptr != memset) {
memset_ptr(p, 0, len);
}
I'm no expert, but this seems like a believable defeat of the technique in the post.Of course, using `dlsym()` isn't exactly portable...
If you've put a semaphore, or mutex lock, or whatever around your calls through memset_ptr(), the transformations will all take place inside the lock, and data races should not be an issue.
memset_ptr is a const (not changed by this program... theoretically) volatile (allowed to be changed by the system, theoretically) pointer to memset. In THIS PARTICULAR CASE, memset_ptr points to memset. The compiler however doesn't know that it won't change due to another processes, but we do. So the compiler shouldn't be able to optimize out the call directly to the function pointer because it introduces a possible race condition: the program reads that memset_ptr points to memset, then the pointer changes (due to some other process changing it), but the program still calls memset, and not memset_ptr. The optimization allows for a possible race condition to occur.
Because a JIT may have enough knowledge of the underlying system to know that the pointer is not pointing to a memory-mapped / DMA'd / etc area, and as such can be assumed to remain constant.
This begs the question of what is "observable behaviour" - execution time, which is definitely "observable" and the basis of timing-based attacks, can certainly change depending on what the optimiser decides to do.
I think this and similar cases of "fighting the optimiser" should really be solved with per-function (or even per-statement) optimisation settings; both GCC and MSVC support #pragma's to do this, although it's nonstandard.
This "Performance at all costs, including safety and predictability" thing may be appropriate in video games, but for security-critical applications that philosophy is downright negligent.
Just like how a language that doesn't do array bounds checking and permits pointer math (read: a language where buffer overflows are a damn feature) isn't appropriate for security critical tasks.
Just like how a language that must be actively fought using clever hacks in order to prevent it from undermining your attempts to guard against common exploit vectors isn't appropriate for security critical tasks.
The biggest downside I see is that it's non-portable, but the reality is that there's not all that many architectures out there to port to anyway (x86, ARM, MIPS probably covers 90%+) and for truly security-critical code having that level of control could be worth it. (This also avoids the "trusting the compiler" problem - an assembler is far easier to verify correctness of than even the simplest C compiler...)
... sort of. Chips themselves perform some optimizations.
http://blog.cryptographyengineering.com/2012/10/attack-of-we...
More recent stuff in this vein:
https://eprint.iacr.org/2014/435.pdf
(Yikes!)
There's a entire class of timing attacks that rely on the CPU cache - they would not exist if CPUs didn't do the optimization of storing local copies of limited sections of RAM.
In most cases that distinction is nit-picky. It can be perfectly reasonable for the language to throw up its hands, shout "undefined behavior", and just assume that whatever random uncontrolled thing happens won't be too terrible, assuming whatever your program does isn't too important. But for security-critical applications it's a really stinking important distinction, because the range of possible behaviors found in the "undefined" category includes things like Heartbleed.
That said, I realize that allowing data that's hypothetically disappeared forever into the free() black hole to come back into the universe through white holes such as malloc() and buf[buf_length] aren't the only reasons why you'd want to make for sure that you can clear out unused memory. Which is why it would be nice if the C spec also included some way to securely clear up memory that the compiler isn't allowed to defeat. No, an optional feature in a spec that's only 3 years old and mostly not supported isn't good enough. If it isn't ubiquitous it's not a whole lot more useful than any of the platform- and architecture-specific fixes that already exist.
sbuf = mmap(..,..,MAP_PRIVATE|MAP_ANON); // &etc.I can't agree with this comparison at all. It may be true if speed is what you want but that's not the only reason.
C (and to a lesser extent, C++) are indispensable in lots of situations where you're working close to the metal. There are very few viable alternatives when working with kernel space code, micro controllers or embedded applications as well as crypto primitives.
Rust is perhaps the only language that can be used instead of C and C++ in these applications.
I'd liken C more to a heavy duty vehicle, something that most people never need but there's no replacement for the tasks it is intended for.
Honestly, in my eyes, this is a place where the language is specified in a way that it cannot ever guarantee what you're trying to achieve and in that case you're best off not relying on the language, but on other things that can make such guarantees.
memset(key, 0, sizeof(key));
if (key[0]) // we are using key, so you can't skip memset()
dropDead();
Unless the compilers "understand" memset and still optimize away the last two lines? I would hope not... Does anyone know how aggressive the C optimizers are these days?They do. That's the whole "problem" - the compiler knows what memset is and what it does, since it's specified in the standard.
i.e. declare the storage volatile but running your crypto code on a non-volatile ptr to it (obtained via cast) to get your performance back?
If the compiler then generates enough smarts to work out that the non-volatile ptr you've passed into your crypto code is referring to volatile storage, then you keep security but get a (noticeable in testing?) performance hit.
I guess that's not as good as your solution though.
mlock() or mlockall() is useful in this case, but those are POSIX functions, not C.
If your user lacks the capability, you have to run the program setuid root. Allocate the sensitive buffers at the start, call mlock() on them and only then drop the privileges.
There's also the 10kg fine-tuning hammer, mlockall(2). That makes ALL the memory for the calling process to become unswappable. As it can lock either the "currently held" memory, or "all the memory to be allocated during process lifetime", it can provide for some additional amusement under memory pressure.
You don't zero sensitive buffers. You randomize them, then free() them.
I wonder if this trick could also be used to solve the double-checked locking problem.
From the quintessential DCLP paper (http://www.aristeia.com/Papers/DDJ_Jul_Aug_2004_revised.pdf):
Consider again the line that initializes pInstance:
pInstance = new Singleton;
This statement causes three things to happen:
Step 1: Allocate memory to hold a Singleton object.
Step 2: Construct a Singleton object in the allocated memory.
Step 3: Make pInstance point to the allocated memory.
[...]
DCLP will work only if steps 1 and 2 are completed before
step 3 is performed, but *there is no way to express this
constraint in C or C++*.
But Colin's pattern here seems to be a way of indeed guaranteeing this. The volatile function pointer is a barrier against inter-procedural optimization: if the function must be called, then step 3 cannot possibly be performed before steps 1 and 2.(There might still be necessary hardware barriers that are missing, and the lack of a memory model for pre-C11/C++11 probably makes it all technically undefined behavior anyway. But the key sequential ordering constraint that was claimed inexpressible in C and C++ appears to indeed be expressible with this trick, if indeed the trick works for guaranteeing a call to memset).
So just to clarify, there's no way the compiler could do "1, 3, 2" instead of "1, 2, 3"? It seems a naive implementation of a compiler could store the pointer to the allocated memory in the pInstance variable before calling the constructor, rather than using a temporary location for the pointer (e.g. a register). Does C++11 and later specify otherwise?
static void secure_memset(void *, int, size_t) __attribute__((weakref("memset")));But no, I'm not talking about VM paging.
To a C dev, that's the same as communism to a US Republican.
You can still side-step the problem by calling a function that's not defined in the standard (so it cannot be inlined by the compiler), usually something like SecureZeroMemory on Windows and its equivalent on other OSes.
On another note, searching for memset_s and openbsd yielded this hit from 2012:
https://mail-index.netbsd.org/tech-security/2012/07/22/msg00...
Which seems to point back to:
https://mail-index.netbsd.org/tech-userlevel/2012/02/25/msg0...
So I guess the "trick" outlined in the (very lucid) post has been known for a while.
foo->bar->baz[i].oof = foo->bar->baz[i].durb + meep;
vs what *tmp = foo->bar->baz[i];
tmp->oof = tmp->durb + meep;
EDIT: I'm not asking for a link to this:https://gcc.gnu.org/onlinedocs/gcc/Optimize-Options.html
I'm asking if there is advice about it. Any overviews with common pitfalls, advice on when to use -O1 vs -O2, specific optimizations to turn on/off, etc.
foo->bar->baz[i].oof = foo->bar->baz[i].durb + meep;
This is fine, no need to "optimize" anything. This kind of common subexpression elimination should be done by any modern compiler (for any language!) and the algorithm behind it is taught in university classes too.Most of the time it's safe to use -O3. If you're doing numerical code with floating points -ffast-math is also pretty safe if your code is correct (ie. no NaN/Inf bugs). Almost the only reason to turn off optimization (-O0) is when higher optimizations make using a debugger harder.
Here's a pretty nice article with some specific optimizations that GCC can (or can't) do. It's pretty old, though, the examples were done with GCC 4.2.1, current version is around 4.9.
http://ridiculousfish.com/blog/posts/will-it-optimize.html
These days Clang can be as good or better than GCC most of the time. The exceptions are in more exotic code like kernel space stuff or micro controller programming.
There's no room for guesswork if you actually want to optimize code, so spend some time reading the assembler output from your compiler as well as benchmarking the results. I usually use objdump -d objfile.o to look at assembly output.
You can also compile to assembly with -S. I think it's clearer that way.
I personally find code like that harder to follow. The first version is clearer than the second (and you forgot to take the address of foo->bar->baz[i]).
Ha, I was afraid of that. :-) Still re-learning when I need that with arrays and when not.
Nope -- shouldn't hurt at all.
It's interesting to me that LuaJIT recommends not using temp variables like this because they can hurt optimization for LuaJIT. That's obviously very different than C in almost every way, I just mention it because it was so surprising to me that there is a situation (in any optimized language) where a temp variable could hurt optimization.
It's useful to add some variables to be inspected in the debugger.
Thanks for the clarification. The obvious caveats of -ffast-math are well documented and most applications shouldn't use that flag.
There are exceptions to this however, I tend to work on such problems. For example, game physics, 3d graphics and some scientific algorithms that have built-in numerical inaccuracy (so -ffast-math doesn't help but doesn't hurt either) but high perf requirements. I also tend to have extensive testing for the most crucial parts of my programs that should catch any problems with this (but many people don't do this with game physics, etc code).
Thankfully, -ffast-math is easy to disable if you start suspecting problems that are caused by that flag.
Game physics because of lockstep networking and replays, scientific algorithms because... well, you want your results to be reproducible. Scientific method and all that.
Often times you don't mind if it isn't accurate, but that isn't the same thing as precision. You want it to be precise, i.e. reproducible.
Yes, there are cases
foo->bar->baz[i]
unless "meep" has side effects.To address the question, these some of the guidelines I try to follow:
- Compile with "-Wall -Wextra" (and "-pedantic" if feasible)
- Modularize your code. You can always mark functions "static inline."
- Don't try to be clever; "Everyone knows that debugging is twice as hard as writing a program in the first place. So if you're as clever as you can be when you write it, how will you ever debug it?" In fact, Kernighan has a lot of good advice: https://en.wikipedia.org/wiki/The_Elements_of_Programming_St....
- Be careful with signed integers. Overflow can do weird things to your program. You can make signed integers act like unsigned integers on overflow using -fwrapv, but if that behaviour is correct, you probably should have used an unsigned integer outright.
- Be careful with pointers; specifically, the requirements of any pointer passed to a function should be explicitly documented: whether it is allowed to be null, whether it's an "in" parameter or an "out" parameter, whether it points to one object or an array, etc. If a pointer points to an array, carry a length parameter with it; null-termination is really easy to foul up.
- Don't optimize until the program needs to be faster. When it does, profile and target the low-hanging fruit. Personally, I usually use either -O0 or -Ofast, depending on whether or not I'm debugging something (-Og is a good one if you need speed while debugging).
- Speaking of optimization, don't underestimate the power of inlining. It's easy to go overboard with it, but it can make a big difference in the right situations.
- Your compiler probably has a peephole optimizer. Replacing "i / 16" with "i >> 4" is probably not an improvement to the quality of either the source code or the object code.
- If you find yourself reimplementing something that C++ knows how to do, consider using C++ to do that. It's not always politically feasible, but remember that you can link C and C++ code.
However, if you're ever in doubt, I recommend compiling very short functions and viewing their output.
typedef struct {
int oof;
int durb;
} baz_t;
typedef struct {
baz_t *baz;
} bar_t;
typedef struct {
bar_t *bar;
} foo_t;
void f(foo_t *foo, int i, int meep) {
foo->bar->baz[i].oof = foo->bar->baz[i].durb + meep;
}
$ gcc -O2 -c -o test.o test.c
$ objdump -d -r -M intel test.o
test.o: file format elf64-x86-64
Disassembly of section .text:
0000000000000000 <f>:
0: 48 8b 07 mov rax,QWORD PTR [rdi]
3: 48 63 f6 movsxd rsi,esi
6: 48 8b 08 mov rcx,QWORD PTR [rax]
9: 48 8d 34 f1 lea rsi,[rcx+rsi*8]
d: 03 56 04 add edx,DWORD PTR [rsi+0x4]
10: 89 16 mov DWORD PTR [rsi],edx
12: c3
You can see here that it followed the chain of pointers only once.The one thing to watch out for though is things that gcc isn't allowed to optimize because of C. For example, if a pointer escapes the function (to another function that the optimizer can't see), gcc cannot assume that the pointed-to memory remains unchanged, even if the called function takes a const pointer! Because the function could always cast away const. For example, this variant will have to follow the chain twice:
int g(const foo_t *foo, int i);
void f(foo_t *foo, int i) {
int x = g(foo, foo->bar->baz[i].durb);
foo->bar->baz[i].oof = x;
}
Generally people always use at least -O2. The main difference between -O2 and -O3 is that -O3 is more aggressive with unrolling and other optimizations that increase code size, so sometimes -O2 is faster because of icache pressure. I generally use -O3 on my tightest loops and -O2 (or even -Os) on everything else.The subtlety is that while the compiler considers it a dead store, you don't, because you're going behind the compilers back to examine memory afterwards.
In general this isn't a big deal. It's only a big deal here because we're working on the assumption that you might have an issue in your code that invokes undefined-behavior and access contents of memory that you're not supposed to be looking at anymore, and the compiler's just assuming that your program won't ever allow that.
void
dosomethingsensitive(void)
{
uint8_t key[32];
...
/* Zero sensitive information. */
memset((volatile void *)key, 0, sizeof(key));
key[0] = key[1] + 1;
}
Would that thwart the optimizer, or would it also see through that usage and eliminate it as well? void
dosomethingsensitive(void)
{
uint8_t key[32];
/* ... */
int i;
for (i = 0; i < 32; i++)
key[i] = 0;
key[0] = key[1] + 1;
}
Assuming the optimizer is sufficiently smart, then it'll remove that 'for' and it'll remove the addition after it in the same fashion.For instance:
if (something)
do_something();
Elsewhere: #ifdef CONFIG_RUNTIME_ENABLE_SOMETHING
bool something = false;
/* and some means of changing something at runtime */
#else
const bool something = false;
#endif
If you don't define CONFIG_RUNTIME_ENABLE_SOMETHING, and thus "something" cannot change at runtime, then the compiler should recognize the if as always false and throw away the call to do_something(). If that was the only call to do_something(), it should throw away the code of do_something().Since there's no accounting for shenanigans like using a debugger to look at variables that are never accessed anymore or dumping the entire stack to a file at arbitrary points etc, the compiler is free to consider code unworthy of being executed if it provably doesn't contribute to further observable behavior of your program.
-O0 // Do Not Optimise
This sounds like it should do the trick (but I've not done C coding for quite some time, so I don't know if there's a nuance as to why it wouldn't), but would also presumably kill any other optimisation.
A better option may be to combine that with pragmas (https://gcc.gnu.org/onlinedocs/gcc/Function-Specific-Option-...) to switch optimisation levels within the code.
Anyone who's played with GCC recently know whether this would work or not?
(discussed on HN 8 days ago: https://news.ycombinator.com/item?id=8233484)
How is the case for modern C++? Are there `vector` or smart pointer alternatives that reliably zero the memory in the destructor?
/* implemented in another translation unit */
void zero_for_sure(void *data, size_t size);
void func(void)
{
char securedata[42];
/* ... */
zero_for_sure(securedata, sizeof securedata);
}
The key here is that our zero_for_sure is an external function in a separately translated file. In the absence of a stunningly advanced global optimization that peeks into other previously compiled units, the compiler has no idea what zero_for_sure does, and so it has to earnestly pass it the given piece of memory.In turn, zero_for_sure is just this:
void zero_for_sure(void *ptr, size_t size)
{
memset(ptr, 0, size);
}
The compiler has no idea where ptr might come from since this is an external function, and so it cannot optimize away the memset.Only if the compiler could consider the whole program together could it still optimize this.
In fact, you don't even need this function, just a dummy external function:
void zero_for_sure(void *ptr, size_t size)
{
char securedata[42];
/* ... */
memset(securedata, 0, sizeof securedata);
commit(securedata);
}
Of course, commit is a noop which just returns. But the compiler doesn't know that because commit is in another translation unit.The only optimization card that the compiler could pull here is since securedata is going away (so that it is illegal for commit to stash a pointer to it), it's okay to call commit with a pointer to some other block which contains zeros, and not actually securedata.
With any trick like this, you should inspect the object code to make sure it's doing what you think it's doing.
Oh, and sizeof doesn't require parentheses when the operand is an expression; they are required when a type name is used as an operand.
You'll be surprised, but this stuff exists since late last century, known as Link-Time Code Generation (LTCG):
A C program consists of translation units which may be preserved in translation form. That happens in translation phases 1 through 7. Multiple translation units may be linked, which is translation phase 8. Phase 8 only consists of resolving references; the last semantic analysis takes place in phase 7.
An example under 5.1.2.3 gives the range of adherence between actual semantics and abstract semantics. Though it is just an example, and not normative, it is very clear from the wording that the locus of valid optimizations is the translation unit.
Selected citations:
5.1.1.1 Program Structure
A C program need not all be translated at the same time. [...] After preprocessing, a preprocessing translation unit is called a translation unit. Previously translated translation units may be preserved individually or in libraries. [...] Translation units may be separately translated and then later linked to produce an executable program.
5.1.1.2 Translation Phases
[...]
7. White-space characters separating tokens are no longer significant. Each preprocessing token is converted into a token. The resulting tokens are syntactically and semantically analyzed and translated as a translation unit.
8. All external object and function references are resolved. Library components are linked to satisfy external references to functions and objects not defined in the current translation. All such translator output is collected into a program image which contains information needed for execution in its execution environment.
5.1.2.3 Program Execution
8. EXAMPLE 1 An implementation might define a one-to-one correspondence between abstract and actual semantics: at every sequence point, the values of the actual objects would agree with those specified by the abstract semantics. The keyword volatile would then be redundant.
9. Alternatively, an implementation might perform various optimizations within each translation unit, such that the actual semantics would agree with the abstract semantics only when making function calls across translation unit boundaries. In such an implementation, at the time of each function entry and function return where the calling function and the called function are in different translation units, the values of all externally linked objects and of all objects accessible via pointers therein would agree with the abstract semantics. Furthermore, at the time of each such function entry the values of the parameters of the called function and of all objects accessible via pointers therein would agree with the abstract semantics. In this type of implementation, objects referred to by interrupt service routines activated by the signal function would require explicit specification of volatile storage, as well as other implementation-defined restrictions.
I don't think that's clear at all. As you said, it's just an example.
> Phase 8 only consists of resolving references; the last semantic analysis takes place in phase 7.
Why do you think that the only place you can optimize is during the "semantic analysis" phase?
To me the phrase "All such translator output is collected into a program image" (from step 8) is vague enough to not rule out optimization during "collection".
"Translator output" suggests that translation is complete and we just have its output to link together. Optimization is "semantic analysis"; you cannot optimize without reasoning about meaning, and optimization also implies that translation is still going on: the output of earlier translation is still being tweaked, with regard to the meaning of the original source.
2. You bring this on yourself; it's not enabled by default by ordinary optimization options like -O2 or -O3. You have to ask for it, and so you must know what you're doing.
3. Under gcc, it looks like only those object files compiled with -flto are prepared for this optimization. You can arrange through your makefile or whatever not to apply -flto to sensitive modules that cannot be inlined or optimized away into other translation units. Those object files won't then contain the GIMPLE bytecode and whatnot needed to be able to peer into their internals at link time.
4. I don't think the dynamic linker in libc (ld.so) does this optimization, so putting code into shared libs may be another good way to hide it.
So, basically, the external function approach is still a very good tool for defeating unwanted inlining and dead code elimination, provided you don't stupidly use some advanced features that bend the standard translation model of the C language. In security-critical code, to boot. External functions are expressed using the standard language; the approach will work under pretty much any compiler.
Thanks to the performance requirements of dynamic linking it's going to be a really long time until we have dynamic linker peeking into .so files and checking what a func does.
If you write crappy code and expect the compiler to fix it for you, you should maybe consider another language. I can only imagine how hard it is to write reliable system software in a language that does these things.
Specially fun when trying at home your UNIX homework and vice-versa.
You don't expect to write sound-looking code and have it broken for you by the optimiser.
This seems rather to be a case of writing good code and having the compiler break it for you.
BTW, I gave you an upvote by accident :D
You could try do low level optimizations in your C or assembly, but for most programs this will eventually backfire. So letting the compiler do its job is actually a good thing.
He also runs http://www.tarsnap.com/ which is arguably the most secure (and cost effective) back up solution in the market.
(I'm in no way affiliated with Colin and/or Tarsnap. Just a fan of his work and humble attitude.)
I'm guessing you haven't seen the "comeback of all time" thread...
But I think it's super funny and cool that you (yourself) are pointing it out.
All the best with you.
Edit: just read the "comeback of all time". That was really funny. Nice nod from PG as well. For those of you unaware like me: https://news.ycombinator.com/item?id=35083 Colin is our resident mathematical genius :)
I'm not looking for you to take sides on the matter. I just enjoy reading history and would like to read a recent review of the two BSD now 11 years later and how they compare.
I've looked and looked for the past few months and can't seen to find a good solid length article on it ... which is why I ask.
On the other hand, I wouldn't want to run DragonflyBSD in production... because skunkworks projects often don't work.
1. The FreeBSD base system is developed intact, so there are far fewer kernel/library versioning issues.
2. FreeBSD, possibly because of its academic heritage, tends to take a more careful "let's study this problem and make sure we come up with the right solution" approach. This means that FreeBSD development is often slower, but once a feature is added it is more likely to actually work. (This is somewhat self-reinforcing: FreeBSD's stability attracts companies building servers and appliances, and when those companies contribute back they care a lot about having things continue to not break.)
3. Linux is far more popular, especially on desktops, so it tends to get drivers for new hardware faster (especially for consumer hardware).
4. FreeBSD is BSD licensed, which makes it available for a lot of companies which wouldn't want to get anywhere near Linux.
[1] https://www.facebook.com/careers/department?req=a0IA000000Cz...
It's still not FreeBSD, but the difference is pretty small these days.
#include <string.h>
void doSecure(void)
{
/*char key[32];*/
char *key = (char*) malloc(sizeof(char)*32);
memset(key,sizeof(char),32);
}
int main(void)
{
doSecure();
return 0;
}
-- key on stack
main:
.LFB13:
.cfi_startproc
xorl %eax, %eax
ret
.cfi_endproc
-- key on heap
main:
.LFB13:
.cfi_startproc
subq $8, %rsp
.cfi_def_cfa_offset 16
movl $32, %edi
call malloc
movabsq $72340172838076673, %rdx
movq %rdx, (%rax)
movq %rdx, 8(%rax)
movq %rdx, 16(%rax)
movq %rdx, 24(%rax)
xorl %eax, %eax
addq $8, %rsp
.cfi_def_cfa_offset 8
ret
.cfi_endprocYou have zero guarantees it will work in another compiler or even between releases of the same one.
Never take a compiler behaviour for the standard. That's the beauty of standards.
memset(key,sizeof(char),32);
Note that you are not clearing 32 bytes with NUL (byte value 0). You are clearing 32 bytes with byte value 1 == (int)(sizeof(char).That's why memset that works 8 bytes at a time fills memory with 72340172838076673 == 0x0101010101010101
If you force it to execute this code it only benefits the very rare security program, yet every single program will run slower.
That's a bad tradeoff. Better to make the security program jump through hoops and let everyone else run fast.
So: this is not something you can rely on always working. Yes, it works currently, but it is not guaranteed to always do so.