How to dismantle a compiler bomb
codeexplainer.wordpress.com
codeexplainer.wordpress.com
It was particularly strange to set a breakpoint at the beginning of main() with gdb and see the program never got there. Oddly enough, he never actually used the 8GB array, even though he had no problem allocating the array on the POWER workstation he was using.
That’s why it worked on his machine. It didn’t work on yours because of the lack of address space.
Not if you tell it not to. This is a common configuration on servers and other applications where requiring overcommit to function is considered a bug.
(There are some surprising exceptions such as PIC hardware stacks where you might be allowed exactly 8 call frames, and your whole program's call stack must be a DAG with no recursion)
Even if that weren't true, the only way overcommit ever helps is if you don't initialize or otherwise touch the memory you allocate.
I'm not sure what happens if you try to write zeroes in overcommit allocated memory, but I would bet it breaks because it'll pagefault and linux will allocate it.
https://codegolf.stackexchange.com/questions/69189/build-a-c...
All are short bits of code in a variety of languages that expand to massive files.
Also, TIL according to the C standard, 'main' is not a reserved identifier! (https://stackoverflow.com/questions/34764796/why-does-declar...)
If anyone can clarify: I assumed that the gcc read literals directly into their 'target' type, but it seems like some literals (such as '-1u') are read as signed integers first then typecasted to the the target type?
That's not quite true.
Main is not required to be defined in a freestanding environment, but in a hosted environment:
> 5.1.2.2.1 Program startup
> 1
> The function called at program startup is named
> main. The implementation declares no prototype for this function. It shall be defined with a return type of int and with no parameters:
int main(void) { /* ... */ }
> or with two parameters (referred to here as argc and argv, though any names may be used, as they are local to the function in which they are declared): int main(int argc, char *argv[]) { /* ... */ }
> or equivalent; 10) or in some other implementation-defined manner.main can take on differing types, but it then becomes undefined-behaviour, which allows the compiler to do whatever it wants.
(5.1.2.2.1 Program startup, C11 Standard http://www.open-std.org/jtc1/sc22/wg14/www/docs/n1570.pdf)
In fact, main() is just a convention of libc. You can have C without libc. (Such as when writing a kernel!)
Now, attempting to link a standalone executable without a '_start' symbol, on the other hand...
No, it's not just a convention. 'main' is defined as the execution entrypoint in at least the C11 [0], C90 [1] standards. Both have both these forms defined:
int main(void) {}
int main(int argc, char* argv[]) {}
You don't have to follow that convention, but then it becomes implementation-defined behaviour.C without libc can still expect to have a main. One doesn't imply the other. It's just without a main, you also have to manually link to _start as well.
[0] http://www.open-std.org/jtc1/sc22/wg14/www/docs/n1570.pdf
[1] http://www.open-std.org/jtc1/sc22/wg14/www/docs/n1256.pdf
GNU ld and gold both support --entry to change the entry point. Alternatively, you can specify it in your linker script.
https://gist.github.com/prattmic/1f3618025aab52b0e90e88445b6...
$ gcc -nostdlib -ffreestanding -Wl,--entry=foobar -o main main.S
$ ./main; echo $?
42I would have expected a type error.
Likely there's some important code out there that relies on this strange behavior.
> Otherwise, if the new type is unsigned, the value is converted by repeatedly adding or subtracting one more than the maximum value that can be represented in the new type until the value is in the range of the new type.
https://stackoverflow.com/questions/50605/signed-to-unsigned...
Signed integer underflow/overflow is UB, on the other hand.
You can even do sparse initialization if we're talking C99 here, like so:
int array[] = {1, 2, 3, [99] = -1}
And you'll get an array of length 100 {1, 2, 3, 0, 0, 0, ..., 0, -1}
Any missing initialized element is always 0 if initialization is done at all, though.
But I agree that it would make sense to fill the array (or have an easier method to do that), it may just be an argument of speed. I think OSes generally hand over uninitialized memory zero'ed out to prevent reading of memory from previous programs - so it's a case of allocating the memory space and then continuing, as opposed to setting values for each position.
int array[5] = {X};
Could be be used to set all values of the array to X, but have never used it for anything other than 0. That's somewhat surprising behavior.> int array[5] = { 0 };
which takes advantage of this missing element initialization behaviour.
In this particular case, the article clearly says:
> The array will contain 4294967295 integers, each with the size of four bytes, taking up 17179869180 bytes in total.
Later, the article again stresses again that you have to multiply by four:
> having the size of 10000 integers, that is, 40000 bytes
It's really hard to imagine how somebody who actually did read the article could have missed that.
Anyhow, this was more of a short blog post than a full article. It was a 2 minute read (not exactly a novel). And the thought experiments at the end are interesting, and illuminate some important points about how the language operates. Definitely worth the read.