The Linux Kernel Is Now VLA (Variable-Length Array) Free
phoronix.com
phoronix.com
[1]: Just by way of example, one that I ran into recently: https://elixir.bootlin.com/linux/latest/source/include/uapi/...
void f(void *buf) {
// Clear bytes 10-19
memset(buf + 10, 0, 10);
}
Making void pointers behave like char pointers avoids pointer casts in cases like this.It would be better for the C spec to just define sizeof(void) == 1.
Both of those have good reasons, following them makes sense, and if you're at all lucky there's no conflict.
While that might have been the case, this announcement says that one of the many reasons to stop using VLAs is to allow the kernel to be compiled with clang. The announcement reads like being able to compile the kernel with other compilers is a very desirable property, that has taken many years of hard work to achieve due to the incorrect assumption that "GCC is the C standard".
So while that assumption might have made sense back then, it does not appear to make sense now. If you treat one compiler as "the standard", chances are that that's the only compiler that you will ever be able to use. That's a bad strategic decision for a big project like the Linux kernel.
So it's just the question of who is active on the committee.
{
float *x = malloc(n * sizeof*x);
...
free(x);
}
you do simply {
float x[n];
...
}
In image processing, you often need large temporary images, but it is dangerous to distribute code such as the above unless you play with the stack limits from outside your program.For variable sized arrays, there is already a bit of overhead in sizing the allocation. Would it really have been impossible to move large allocations to the heap with automatic free at exit of the scope?
C++ has had a proposal for std::dynarray floated for a while, which would basically have semantics allowing stack allocation, but wouldn't mandate it, leaving it as a quality of implementation issue (so it would be legal to implement it on top of vector, but it was assumed that compilers would go for optimizations given the opportunity). It didn't pass, unfortunately.
How about std::array ?
My point is only that runtime-sized arrays may have made it into the C standard, but they aren’t in the C++ standard.
Not a compiler writer here so I'm not sure. I think it wouldn't be too difficult, but then this goes very much against the exact control you normally expect from C. And violating expectations like this is usually not good.
It would also produce bigger functions in a somewhat opaque manner, because you would need the code to deal with malloc and free and stack allocation at the same time.
This is not a property of the language, it's a property of the runtime its running on.
In a normal userspace program under Linux (at least), you can stack allocate megabytes and the stack will dynamically resize (note that shrinking the stack might not happen in a timely manner).
But in kernel space, there's a conscious decision to keep the stack small and not dynamic. While it's not impossible to have dynamic stack in kernel space, that has a lot of implications.
"Many years later we asked our customers whether they wished us to provide an option to switch off these checks in the interests of efficiency on production runs. Unanimously, they urged us not to--they already knew how frequently subscript errors occur on production runs where failure to detect them could be disastrous. I note with fear and horror that even in 1980, language designers and users have not learned this lesson. In any respectable branch of engineering, failure to observe such elementary precautions would have long been against the law."
-- C.A.Hoare on his Turing's award speech
C and B designers had another point of view regarding security, or lack of thereof.
But yes, it was probably too naive to expect that to work out.
_cleanup_free_ float *x = malloc (...);
Since cleanups are supported by both GCC and Clang it's not a real problem to use them on Linux, BSD and MacOS. We recently added them to libvirt, have used them in libguestfs for a long time, and as I mention above they are used in systemd for years. They are also applicable to many other cases such as automatically closing files at the end of the current scope.It wouldn't work, like a fixed size array wouldn't.
Languages like Ada do have them, but then again the language also protects against stack corruption.
If the requested amount if bigger than the stack size defined at compilation time, an exception is thrown.
While C just overwrites the stack.
But even as implemented today, in practice, it all works fine if you're not working with really large arrays. In a lot of scenarios, you do not. This isn't really any different than using recursion when you reasonably assume that recursion depth is not going to be large. Sometimes it's just easier and clearer all around.
Which was the main theme at Linux Kernel Security Summit 2018.
Just for info, if not wanting to bother watching Google's talk, 68% of Linux kernel exploits were caused by out of bounds errors, including those caused by VLAs misuse.
3: Use verifiable loop bounds for all loops meant to be terminating
[1] https://lars-lab.jpl.nasa.gov/JPL_Coding_Standard_C.pdfAccording to the article, “there may be a couple more VLAs hiding in hard-to-find randconfigs.”
Those configs mean there are massive combinations of what is actually set, and which code you're actually trying to compile. I imagine that's similarly difficult for static analysis; due to the build system and what code you run it over. I think a simpler approach with grep might actually work better.
void foo(int n)
{
int m[n];
}
The downside is if n is large you use a lot of stack, but you wouldn't notice if it's mostly called with low values. So a bit of a ticking time bomb.I guess if you have a really large n, like more than a few pages worth, then certain values of m[i] might wind up in some other page allocation where it shouldn't.... I seem to recall a vulnerability like this a year or two ago. (Normally a stack can grow on a page fault by hitting a guard page, but if the n above is absurdly large...)
if (n > MAX){
return;
}
Would not help?* the code is bigger, and people say is meaningfully slower
* super easy to accidentally blow the stack - in kernel land that is probably going to end in a panic
* using vLAs used to disable stack canaries for those frames, not sure if it still does
For a lot of cases where you are truly going to need variable length you’re unlikely to be penalised heavily by a heap allocation due to the relative cost of the operations.
really ? there's a difference but unless you keep allocating in a hot-loop this should hardly ever have observable costs : https://gcc.godbolt.org/z/SN9Ois
A really nice use for a VLA would be in the middle of a struct. You could make something like a PNG image file chunk as a struct, with the CRC at the end coming after the VLA. You get a struct of the correct size with the right fields, despite one part of the struct being of variable size.
For example, a parameter n is passed to a function, and then a struct is defined based on n.
struct {
uint32_t header;
uint32_t array[n];
uint32_t footer;
}foo;
The function fills in the struct. Maybe it then sends the struct via a packet-oriented protocol, perhaps over a datagram socket, so the length is transmitted implicitly in the packet size.https://gcc.gnu.org/onlinedocs/gcc/Variable-Length.html
There is no problem cleaning the stack because the compiler has numerous ways to do that. One way is to keep a copy of the original n, as passed to the function. Another way is to save the original stack pointer.
Making this work is not much harder than supporting the alloca function.
If you prefer, it's an Ethernet packet, with a CRC16 on the end. It is outbound.
The send() and write() functions are happy to accept a void pointer.
Of course whether any of this matters will depend on the specifics of the matter, kernel code being the most conservative by necessity.
uint8_t is effective protection, given the normal assumption of a stack with 4096 to 16384 bytes of space and a call stack that isn't insane.
If you wish to make a formal proof of correctness, feel free to make worst-case assumptions.
Same issue, you just need larger values. In both cases you need a guard.
Hence why they were dropped in C11. Being an optional annex means that actually most C vendors won't bother.
What I don't understand is: after all these decades of stack smashing vulnerabilities, why would such a shortcut have been taken?
Convenience costs performance, safety costs performance ... at some point you just have to suck it up and pay the performance penalty to make things safer and to improve programmer productivity.
Did they profile before optimising[1] and identify it as a problem?
He actually said it is all those things in comparison to doing fixed allocation. The reason that is relevant is the kernel stack is so small you can only use a VLA when you know in advance the upper bound on the size is small, but if it's small you may as well use the fixed upper bound, and if you do that it is always faster, smaller and less fragile than using a VLA.
He's wrong in the general case, because user land C programmers will replace a VLA with:
if (!(array = malloc(n * sizeof(array[0]))) fatal("I'm out of memory");
which compared to the VLA generates more code and is slower. In both the VLA and malloc() case if you run out of memory the program will die nice and deterministically (unlike the kernel), but in the malloc() case that will only happen if you remember to do the check whereas in the VLA case the compiler ensures it always happens.Does this mean that mainline now builds with llvm?
>>> - VLAs within structures is not supported by the LLVM Clang compiler and thus an issue for those wanting to build the kernel outside of GCC, Clang only supports the C99-style VLAs.
one less vendor lock in (but since it's GCC, it makes me sad).