Understanding thread stack sizes and how Alpine is different
ariadne.space
ariadne.space
The wording sounds as if it is trying to assign blame for the problem. What, then, is the guaranteed thread stack size? A developer would obviously need to know this (and other things such as the amount of stack size required by variables, parameters and frames) to not fall into this trap of writing non-portable programs.
Yes. It sound like someone that have been in a million discussion about why "that your program crashes is not a bug in our operating system".
Many people, even engineers, expect things to behave as they usually do instead of behave as specified. So, I can imagine how all that discussions go.
If you need a big stack, you can just ask pthread to give you one.
Thanks, posix.
Yes. Sadly, people often write incorrect programs.
> What, then, is the guaranteed thread stack size?
I can't be bothered to look up the POSIX guarantees, the gist of it anyway is that it depends on the application developer and system administrator.
> A developer would obviously need to know this (and other things such as the amount of stack size required by variables, parameters and frames) to not fall into this trap of writing non-portable programs.
If you're not going to calculate the exact requirements (probably unnecessary), guess/measure a number of bytes and allocate a stack that's 50 or 100 times greater than that. That's better than ignoring the existence of the stack, anyway.
But in general, the difference in image size is negligible because of shared layers, and I just don't think enough testing happens on Alpine / musl in any given stack. Even if your app runtime is tested this way, how many dependencies are?
Come to think of it, I'm not even sure why there was a push for Alpine-based Docker images at some point. Maybe it was just hype.
This broke teams that rely on python and on node, but the docker image guidelines come from a team whose ideal language is now go (and most of whose legacy code is in java), so they are not really sensitive to those concerns. Ironically we tried to move to distroless as implemented by google[1], but that's based on debian which includes glibc, so the un-nuanced CVE checker freaks out again. That effort was quietly dropped.
(I'm not actually disputing the proposition that alpine is better for security under certain circumstances, but I think a lot of "the push" comes from what might uncharitably be described as cargo culting, or with more insight as interpretations that make sense in one context [everything is a static binary, little to no reliance on traditional userland tools] being unquestioningly extended to other contexts.)
The continuing push is due to the smaller footprint and better security properties. And no amount of sharing makes up for the difference between a single-MB image and a GB image.
Any application can just dictate its own thread stack size. What is discussed here is a default.
And so the mountain must come to Mohamed. Increase Alpine's default stack size to something more in-line with the big boys.
Diversity helps discover what is fundamentally a broken and fragile assumption that a dynamic property will always have some value. An assumption that can fail anywhere, including on the OS that was initially targeted, and will fail the moment another OS is targeted.
The developer should fix their broken assumption, but is entirely free to do so by taking control of the value at link time.
I guess make sure your required stack size is not a function of input, and test against a minimum stack size.
But in my experience you don't really compute a "guaranteed" stack size, you use your experience and knowledge of the program to make an educated guess, and then you apply a reasonable multiplier to give you some security margin.
If you don't use (or severely limit) recursive calls you can usually just check that your deepest call stack fits within the bounds. Although finding the deepest call stack in the first place can be tricky given that compilers can aggressively inline function calls.
Keep in mind that this is just for thread stacks - you can set the size for them yourself, so ideally you'd always do it. Then a guaranteed minimum size becomes irrelevant.
How does a normal working programmer calculate the size of each of their stack frames? I'm a compiler researcher and I'd struggle to do that. How are application developers going to do it?
And how do you design a program to have a deterministic maximum call stack depth?
I don't think these things are as easy as you're making out.
I tried googling for how SPARK Ada provides assurances against exceeding stack-size limits, but I couldn't find a decent answer. I presume it does so, though.
edit: forgot about alloca
edit 2: Turns out the AdaCore folks have a tool specifically for static analysis of stack-space requirements of Ada/C/C++ code: https://www.adacore.com/gnatpro/toolsuite/gnatstack
In those leaf functions you can check &local_var and compare it to pthread_attr_getstack(pthread_getattr_np()). (Of course that's not precise for many reasons.)
> And how do you design a program to have a deterministic maximum call stack depth?
If you're running only your code - don't use recursion, or alloca. If you use external libraries, you have to research what they do and add some extra in case of updates.
Bounded stack size is also a common issue if you're targeting small microprocessors.
For non-critical apps it should be pretty easy to figure out the needed stack size. For cases when you want to guarantee it... that gets more tricky.
Edit: just learned that clang has the option -fstack-usage which should help a lot.
Or you can try to figure out the maximum number of times it'll recurse: for example, the height of a red-black tree with less than 2^64 nodes is less than 128, iirc.
I remember one day in the 90s counting out like max address len and max zip code len and so and trying to figure out how long to make my target stack allocated buffer, and i was like fuck it, I have more important things to do, all my stack buffers are hence forth 65536 bytes long.
I know that in the small-embedded world, people do work on such things.
For total depth, keeping you program simple and predictable helps. People certainly manage to do it even for large programs like Linux itself, where stack size is like 16KiB or so. https://elixir.bootlin.com/linux/v5.2/source/arch/x86/includ... and less on other archs. 8 KiB on arm https://elixir.bootlin.com/linux/v5.13-rc7/source/arch/arm/i...
If I tell you as a compiler writer that this loop body from this function, but with this branch and this branch outlined, but only when called from this context, takes n bytes... I don't get what most working programmers are going to usefully do with that information.
If the language is complicated and has generics or whatever, the programmer will have to do more work to understand it.
It's not a huge issue in C.
If you ask a compiler how much stack a function will use the answer for a non-trivial compiler for a complicated language is always going to be 'it depends...'
The only use case is for real-time embedded code -typically in aerospace- where the coding guidelines prohibits you from using any recursion at all and you have to prove the highest stack usage fits into the chosen microcontroller.
On exit just scan from the maximum stack to minimum looking for non-zero.
If you have tests it should be easy to get within a few bytes of max stack used, which is probably just as good as instrumenting everything.
Sure, that could happen.
But what the other guy was saying about being a compiler developer and being unsure how to calculate the maximum depth is that there are many, many ways to arrive at the wrong result. Resursion, argv/envp, varags, alloca, and so on. So unless you are going to spend a great deal of energy proving maximum depth you're going to be using an estimate of some sort. Thus, 'probably just as good'.
Is Alpine using some new kernel? No, it's a Linux distribution that uses the Linux kernel, albeit with some unusual defaults.
Does Alpine not have any of the GNU userspace tools? Also no, there are plenty in the Alpine package repository.
Look, I get that GNU/Linux and ”GnU pLuS LiNuX" is a loaded term and has a lot of baggage, and that everyone would like to just be rid of the whole mess, but the characterization used here had me thinking that there was some other "Alpine kernel" experimental OS project that I had missed that had nothing to do with Alpine Linux.
The word "Linux" never once follows the word "Alpine" in this article, and it discusses overcommit mode as if it's a uniquely "GNU/Linux" thing. WTF does kernel overcommit have to do with GNU?
Please just call it what it is.
Kernel overcommit has nothing to do with GNU, but default stack size has a lot to do with which libc you use. Musl has a different default than glibc. Overcommit is mentioned because it is the justification for glibc having a large stack size by default. Musl has defaults that make fewer assumptions about how the system is configured.
I still think only referring to it as "Alpine" and not once calling it "Alpine Linux" is weird.
That seems pretty strange to modern ears, but maybe the underlying point was that it isn't possible to statically know how much stack size would be required.
I suppose it wouldn't have been obvious then that if you fudge the issue for the first thirty years or so, everyone will just accept that this is the way the world is.
Still, it's a bit of a shame that there are still widely-used systems where if you exceed the available stack space you're likely to face a "weird crash" rather than a clean error message at runtime.
I remember reading that the call stack was invented as an implementation device for recursion as introduced in Algol, but I can’t recall how that claim was sourced.
Given the time frame you're talking about, remember to be thinking in terms of kilobytes, not gigabytes. And potentially low-single-digit numbers of kilobytes. It could be in the range of hundreds of bytes dedicated to stacks at the time. Even if you could compute your maximum size it's easy to imagine people balking a the results of such a computation and think it's not worth it to even consider the possibility because you'd blow your stack so quickly it's not like there'd be any benefit to it.
It's not just that - if it were merely that you couldn't know the size statically, you could use dynamic memory allocation. The problem with that though, is that now every (not provably nonrecursive) function call can now fail with a memory allocation error, and if your language doesn't surface that to the caller, and (correctly) doesn't allow spurious errors to appear out of nowhere (cough every modern programming language cough cough), there's no way to handle that error.
By your own table, OpenBSD uses 512 KiB unless you're on an ancient version. Among the listed, Alpine is the lone outlier.
As far as I can tell this is something they think in principle should be changed.
#include <threads.h>
void some_function(void) {
thread_local char scratchpad[500000];
memset(scratchpad, 'A', sizeof scratchpad);
}
As an important note, thread-local storage through this keyword still isn't supported on OpenBSD. It's a serious PITA.[——— also, copied from a reply I posted below: ———]
The autofree macro is wrong since __attribute__((cleanup)) expects a function that takes an additional level of pointer. In this case, it'll call "free(&scratchpad);". Which doesn't get you a compiler warning in C because passing a char ** as a void * is perfectly fine. But your heap is f*cked after this.
Correct way to do it is
void free2(char **p)
{
free(*p);
}
#define autofree __attribute__((cleanup(free2)))It's also wrong since __attribute__((cleanup)) expects a function that takes an additional level of pointer. In this case, it'll call "free(&scratchpad);". Which doesn't get you a compiler warning in C because passing a char ** as a void * is perfectly fine. But your heap is f*cked after this.
Correct way to do it is
void free2(char **p)
{
free(*p);
}
#define autofree __attribute__((cleanup(free2)))Seems like it would be more sensible if the stack space could just grow when required. Surely not that difficult?
For the main thread, a system can also try to keep other allocations far from the stack without committing to any particular size (heap grows upwards from the bottom of the address space, stack downwards from the top). But this doesn't scale to multiple threads and leads to an unpredictable maximum stack size, so I prefer the fixed reserved space approach.
If you use huge pages to alleviate that, you'll waste a lot of physical memory.
No sense in clinging to old world ideals when they no longer make sense.
Storing small data, like function-local ints, pointers, etc, on the stack is beneficial due to L1$ prefetching semantics, but storing a 512KB scratchpad on the stack (which is what the article is about) will totally trash your L1$ and you'll have MORE cache misses than you would if that scratchpad was not on the stack.
But in the contemporary world, the trend is increasingly to transform functions to "async" forms where much of the functions' local state including return address is stored in heap-allocated space instead.
future languages will be able to suss this out at compile-time: https://github.com/ziglang/zig/blob/2ac769eab9b7dba4cd38e5de...
Nowadays we have at least 48bit virtual address space available; what's the harm in giving each thread a full GB of stack?