Stack size invisibility in C and the effects on "portability"
utcc.utoronto.ca
utcc.utoronto.ca
The customary way to handle stack overflow gracefully in userland is to install a SIGSEGV handler and check the faulting address to see if it is near the stack segment end, handling the signal on an alternate stack. Except that determining if the access was in the stack segment requires knowing the stack segment start. The signal API doesn't allow querying that; you'd have to query the process's mappings by other means. It's also not robust because a program with big frames might just stride too far and confuse the kernel's automatic growing, causing spurious SIGSEGVs when in fact, there is stack remaining.
I found it was just easier to allocate my own stack segment and immediately switch to it and just leave the loader-provided stack segment alone. Virgil does this in the first 50 or so instructions upon entry.
This also makes stack overflow very deterministic (only depends on the program) and controllable--there is a compiler option to set the stack size in KiB. There's no need for the compiler to insert strided accesses either, as the whole segment is mapped, and you can have more than a single guard page at the beginning (think: 1MiB guard region).
Stack segments in C (and Linux) are madness if you use very much and want anything to be robust to overflow.
This is also how the kernel grows the stack. When there's a fault, it compares the faulting address to the stack pointer register. This way, big frames don't confuse the automatic growing. (On Windows, by contrast, stack growth is detected using a single 4k guard page, so compilers must be careful with big frames and insert strided accesses)
I only experimented briefly[1] with getting Virgil threads to run on Linux. You need to call SYS_clone, which forks a lightweight process (i.e. shares all virtual memory), and specify a stack segment. The segment in question would of course, be mapped by Virgil and not have MAP_GROWSDOWN, so it could be made robust to stack overflow along the lines above.
[1] https://github.com/titzer/virgil/blob/master/apps/Multi/Mult...
Also I see x86_64 Linux targets are supported now and that's awesome. Do you plan to support ARM architectures?
Nice to hear someone has read the code. Glad you like it. Cheers :)
The language has no builtin classes, interfaces, or functions. In contrast to say, Java, where the language (and compiler) know about java.lang.Object, java.lang.String, etc. The compiler knows how to compile classes, but it doesn't know the name of any built-in ones.
Virgil is really designed to build the lowest level of software. Thus all my focus on low-level things like calling the kernel and mapping stacks, etc. The only thing that depends on more than just the kernel is the test suite, which uses bash and a little bit of C.
For others reading this thread, to clarify the "no builtin functions" claim, although Virgil doesn't have any builtin values which refer to methods, it does have some built-in operations: ==, !=, +, -, *, /, %, <<, >>, >>>, &, |, and ^, as well as data structuring facilities like arrays and tuples and operations to access them.
Your choices are:
* Make a "stack" yourself, in heap memory, and implement your algorithm to manually push/pop from the stack you implemented (often not hard to do, but feels like reimplementing for no reason).
* Spawn a new thread, where in most languages you can set the stack size. This is what I tend to do, although it does seem like a waste to have an extra thread for no reason other than to get more stack.
You can (in principle) set the stack size when you compile a program, but how to do it seems to differ between OSes and compilers, so I prefer the new thread strategy as an in-language option.
The problem here is that once it is used once it is never freed. (Other than being swapped out). So you effectively cache all of the space your program needs forever.
There are some solutions like trying to mark the stack below you freeable but it is hard to do this in a way that doesn't cause huge performance concerns if you happen to be crossing whatever boundary is picked often. IIRC Golang used to do this but switched to copying, growable stacks because no matter how hard they tried they always ended up in scenarios where they were allocating and deallocating stacks in a hot loop.
They used to do segmented stacks. The bad overhead is not in allocation, but in switching between segments when stack pointer is oscillating around the boundary — but if you want to allocate really many stacks that could grow large, you don't have much choice: either this or tracking pointers that need to be updated on reallocation. Normal contiguously mapped stack, but with dynamic growing and shrinking of actually mapped pages, is an entirely different proposition.
(If you can take the address of a local variable, but you have garbage collection and can handle memory allocation failures by exiting, you can heap-allocate whatever local variables have their address taken.)
Segmented stacks are a particularly efficient way to implement call/cc.
All the space your program uses at the largest stack size, you mean, which is different. Until you touch it, the huge stack is purely virtual. And it only lasts until you restart the program.
Yes, if you default to a 10 GB stack, then recursively tree-search over 9 GB of data, then close the data and spend the next week with 1 MB data that only touches the bottom of the stack, you'll have 8.999 GB swapped out, which does suck.
But that only happens if you run the program on that 9 GB of data. The alternative was to fault when you overflowed the smaller stack when that data was imported.
It does mean if there's an infinite recursion bug, instead of crashing quickly with a stack overflow, you have to wait while it pages out other programs and fills the whole 10 GB of stack...
Yes, if most of your program is using roughly the same amout of stack space then it is great. However I think if you are worried about running out then you are likely using far more than the average part of your program.
So yes, it isn't intrinsically bad, but context important to evaluate the decision.
If it's a balanced tree, then consuming 9GB of stack to do a search would imply that the data structure is what, on the order of 2^1000000 in size? Yikes :-)
If you're up to 10GB worth of nodes in a syntax tree, you're probably doing something elaborate enough that you'd want to squeeze out every ounce of performance, and you're possibly a domain expert. Maybe you're making a static analysis tool? On the other hand, a lot of static analysis tools are painfully slow, so maybe there's a lot of recursive searching on deep data structures going on!
Likewise, if you're trying to solve an NP-hard problem and you're recursing through 10GB of anything, you're probably already asking the computer to do something unreasonable by many orders of magnitude, and all a larger stack would do is delay an eventual out of memory error. Or, you're doing something clever and knowingly pushing the computer to its limits, and so you once again would want to squeeze out every ounce of performance and go with an iterative solution.
What do you mean? It will get freed once you pop that stack-frame won't it?
It was virtual memory before it was used, but once you use it, it's mapped into physical memory and marked as dirty. So it may be swapped out or live there, those are the only two options.
It would take your OS being fully aware of your usage schema for it to be able to free the memory.
https://github.com/ziglang/zig/issues/1639
They're also looking at static analysis of the maximum stack depth which might help in some cases.
However, there are plenty of interesting cases where if you don't use recursion, you end up having to implement your own stack as a resizable vector.
The question "why should the busy work of hand-rolling your own stack be necessary?" is the kind of question that you don't want to waste too much time on as a working programmer, but unless some people think deeply about such questions, the field as a whole stops progressing.
As others point out in this discussion, with a 64-bit address space one could organize the stack(s) in such a way that recursion could use all of physical memory if necessary.
You can also request whatever stack size you need for additional threads when you create them, this has always been a good idea.
It's only the default stack size for non-main threads, if you don't specify it, which may be considered reasonably limited.
The former would not merit consideration if it was automatically increased.
> it grows until it runs into another allocation
Which it won't because vmem.
> and these days there are guard pages and such
That just makes the access segfault if the stack allocation is not large enough to straight jump over the guard page.
Apart from recursive algorithms, stack size can be determined by static analysis.
Which reminds me: what is a good tool for analyzing / viewing a C++ programs call tree?
As already mentioned, perf can capture call stacks, but those are statistical samples, so not necessarily complete.
callgrind (https://www.valgrind.org/docs/manual/cl-manual.html) by contrast can capture all call-stacks.
kCacheGrind/qCacheGrind (https://kcachegrind.github.io/html/Home.html) can be used to view the output from either perf or callgrind. Sources for kcachegrind, including converters from perf, oprofile, etc., can be found at https://github.com/KDE/kcachegrind
Yes, this is true for comparison/search/sort/bucketing algorithms. But in other applications (e.g. graph theory, networks, sparse matrices) a depth of n, for n elements, can be perfectly legitimate. Or it can even be an edge case, but one you don't want to handle by just segfaulting.
> Which reminds me: what is a good tool for analyzing / viewing a C++ programs call tree?
Linux perf can collect the necessary data. Then for frontends, I am not sure. KDAB's "hotspot" ( https://www.kdab.com/tag/perf/ ) can do flamegraphs, but this may be more performance-oriented than what you're looking for.
Edit: perf can collect the data statistically. It will miss many calls, the ones that are the least time-consuming.
But in many contexts you have to be robust against a bad tree.
> Apart from recursive algorithms, stack size can be determined by static analysis.
Not always, in the presence of dynamically sized arrays (or alloca and friends).
In C++, std::vector::emplace_back() and pop_back() are in the standard library for a decade now.
yes, kilobytes.
As long as you use less than 1 page per activation record though, most modern libc implementations behave safely under unbounded recursion as they will allocate a guard-page around each stack.
[edit]
It would also be useful if the compiler could warn about activation records larger than 1 page. Absent alloca (and deprecated c99 VLAs), that should be statically knowable.
Tail Call Optimization (TCO) is an optimization, yes;
but if Tail Call Elimination is required in the implementation of a language as it is in Scheme, then programmers in the language can rely on it (along with lexical scope) to write iterative loops instead as recursive calls. That is a change in program language semantics.
see Guy Steele's "Debunking the 'Expensive Procedure Call' Myth, or, Procedure Call Implementations Considered Harmful, or, Lambda: The Ultimate GOTO" http://dspace.mit.edu/handle/1721.1/5753
This is also supported as a non-standard extension in C, C++, and Objective-C as of Clang 13: https://clang.llvm.org/docs/AttributeReference.html#musttail (I implemented this).
> see Guy Steele's "Debunking the 'Expensive Procedure Call' Myth
Yes! :) I cited this paper in my blog article describing how we are using musttail: https://blog.reverberate.org/2021/04/21/musttail-efficient-i...
EDIT: "My program used to fail with a stack size exhaustion error, but now it works -- I'm so angry about this!" -no one ever.
TCO might cause your program to avoid one stack exhaustion fault only to reveal a different fault (possibly a different stack exhaustion fault). This is not a semantics change. And it's not a bug either.
So I guess that means that an imaginary API to query these things should not talk about "stack", since that is making an assumption that is a bit out of line. That does not make naming any easier ... and, as pointed out by vasama, it would be hard to use an API that queries about the current function only. Perhaps something like "autosizeof f" where f is a function could return, at compile-time, the projected "automatic storage size" of the given function (perhaps making it into a compilation error if f() uses VLAs, or calls functions that do), and a runtime "autoleft()" function could be used to determine how much space is left.
Uh, no, probably not. These things are hard, and I'm making assumptions about code analysis capabilities that are probably not realistic.
I'm not aware of any way to programmatically query the rules laid out in the psABI. The implementation of the corner cases is often based on mutual agreement (convention) between the compiler and the link-editor.
You could probably write a tool that maintains a representation of these rules and returns the results in some sort of parse-friendly EBNF, but you'd need a convention linter to make sure that the implementation of your tool tracks the change in convention in the toolchain components.
On-the-other hand, you could write an extension to the compiler that warns you when certain C code results in behaviors around the stack that might be undesirable, i.e., if you're passing too many parameters to a function and it's going to spill those to the stack.
The standard way to avoid unknown behavior in this regard is to explicitly tell the compiler what you want to accomplish by not relying on auto-allocation or parameter passing rules (as stated in the top thread). This means explicit allocation on the heap, rather than than allowing the compiler to auto-allocate on the stack, and using pass-by-reference to avoid parameters spilling to the stack. All of this is slower since accessing values in the stack is often a single instruction operation (small number of cycles), vs multi-instruction loads from far-away addresses (large number of cycles).
If you have neither, then you still have to worry that fixed sized automatic allocations might be too large, but at least this can be checked statically.
Given either alloca() or VLAs, those can be allocated on a stack or on a heap, and the spec does not say either way. The OS might over-commit memory, too, and if the C run-time has fixed-sized stacks there may or may not be a guard page at the end.
However, even with all of that, using memcpy() or similar (see elsewhere in this thread) to touch every byte or every 4KB's worth of bytes in what might be the direction of stack growth (if there's a stack, otherwise the direction is irrelevant), is something you can do to make sure that the memory allocated is committed, and that any failure due to over-committing (bounded stacks look like over-committed unbounded stacks) is detected as early as possible and without corrupting memory (e.g., by wildly exceeding stack bounds). Failure may ensue, but it should be behavior defined by the machine rather than undefined behavior.
May I ask how you'd do this ? I need this for optimizing a C project.
#ifndef STACK_FRAME_UNLIMITED
#pragma GCC diagnostic error "-Wframe-larger-than=4096"
#if __GNUC__ >= 9
#pragma GCC diagnostic error "-Walloca-larger-than=1024"
#pragma GCC diagnostic error "-Wvla-larger-than=1024"
#endif /* GCC 9+ */
#endif /* STACK_FRAME_UNLIMITED */I fail to see the relation between your first measure (which seems to act as a guard?) and the second. Also their sizes don't match.
Though I'm afraid I'll encumber this thread with my naive questions.
This stack affair of `local arrays are super fast! but keep it small, 1024 should be okay` is very puzzling to aspiring C lovers like me.
I'm not entirely sure it was a standards-compliant compiler though. I think it may have denied recursion?
Mind you, I also think people have unreasonably high expectations of the C standard; in order to get things done you inevitably have to make compromises and use features that aren't available on every platform. Which leads to autoconf hell, but if you're still developing on C you've accepted its imperfections.
Obviously those restrictions mean you have to write your code without recursion, but in practice the uses that SPARK is put to don't tend to need it. A statically allocated stack in your code and a loop can solve a lot of problems (along with suitable logic for gracefully handling that stack filling up).
Granted, it is maddening that C does not offer any great ways to reason about what stack size is needed, but it is simply nonsense to imply (as people seem to be inferring here) that there's no way for an application to handle running into musl's stack size limits.
The nice thing about C++ is that the preprocessor is a monster. We #define'd all references to alloca in the ODE codebase away into calls to our own function that returned, instead of a bare pointer, an object with the right pointer-dereference semantics to act like a bare pointer but that was really a reference to a slice of a colossal chunk of heap memory. When the context of the alloca call closed, the returned object was destroyed, and its destructor moved our imaginary "stack" pointer in the big chunk of heap memory back into place.
Ugliest thing I ever wrote. ;)
Also, this isn't limited to C... The real problem is not that the stack size is invisible to the program, or that you can run out. The real problem is that you can make a stack allocation so large that you might miss the end of the stack and scribble over some other memory. This is why you need to touch a byte ever guard-page-size bytes (in the direction of stack growth) in large stack allocations before using those allocations.
What's specific to C is alloca() and C99 VLAs. Not that other languages don't have those, but that in their absence it's extremely unlikely that any function would have a frame so large as to run past the guard page. A maximum frame size can be computed by the compiler in the absence of alloca() or VLAs.
Also, the compiler could (and arguably should) generate code to do this touching of a byte every guard-size-bytes thing for dynamically sized stack allocations, which would make the problem disappear. (You could still run out of stack space, but more safely.)
> you need to touch a byte ever guard-page-size bytes (in the direction of stack growth) in large stack allocations before using those allocations
> Also, the compiler could (and arguably should) generate code to do this touching of a byte every guard-size-bytes thing for dynamically sized stack allocations, which would make the problem disappear.
That's precisely what MSVC does since 1995 or so: inserting the calls to the _chkstk() function that touches the guard pages. I believe gcc also does this when it targets Windows.
https://docs.microsoft.com/en-us/windows/win32/devnotes/-win...
https://metricpanda.com/rival-fortress-update-45-dealing-wit...
The only way you can reliably do this, I've found, is to allocate your own stack segment.
The cost is slower function calls and an ABI break.
What's NOT fine is for a large stack allocation to cause the program to silently scribble over another stack or heap.
It's more that you can't expect anything better, having an actual stack overflow error would be a lot more helpful.
And yes, portable C has finite size_t. Therefore portable C must fail in some way when limits are exceeded.
How do you do this in pure C?
You can determine stack growth direction by having one function call another with a pointer to a local variable in the former and then compare that to a pointer to a local variable in the latter. To do this correctly you have to cast to uintptr_t.
And, yes, pure C does not even assume a stack, much less a bounded stack. In pure C it's possible to have every call frame allocated on the heap, which means that assuming that they are allocated contiguously is not safe. Moreover, even with stacks, the run-time could allocate a second stack when the first is exhausted and push new call frames on the second rather than die. In pure C the only assumptions about limits have to do with the size of size_t. The real world intrudes though, and in practice all C run-times allocate call frames on bounded stacks.
That's the point - you can't do it without sacrificing portability.
That'd be a completely standards-compliant C implementation.
Your code wouldn't run on it. Your code would not be portable to my standards-complaint implementation. Therefore your code is not portable.
It seems to me like it is impossible to do poking in a way that could not be elided by the compiler, because removing poking never alters the semantics of the program (stack overflow or guard pages are not part of the semantics of C, I believe).
Of course, in practice you'll probably get the right behavior, but I'm personally always interested in language lawyery questions. I do believe that there are cases of security issues introduced by compilers eliding memset operations on password buffers, because the compiler can tell that the result of memset is never used and so it doesn't actually clear the memory. These sorts of things are tricky to get right.
You can cast to volatile, but casting away volatile-ness and dereferencing the resulting pointer is UB.
If you need that part, and that part is not portable, then your solution is not portable is it!
What happens when you access memory which has not been successfully allocated?
Strictly less (not more than every 4KB), but, yes :)
> What happens when you access memory which has not been successfully allocated?
Portable C has finite size_t. Actual machines, virtual or otherwise, have finite resources. At some point you may run out of resources. With a dynamic stack allocation checker you'll fail with a segmentation fault when you run out. Without a dynamic stack allocation checker (but with dynamic stack allocations) you may well fail in random ways due to corruption of other stacks, heap, mmaps, etc.
If your C run-time doesn't have stacks then a dynamic stack allocation checker may be pointless... unless your OS is over-committing memory, then it may cause failure to happen sooner than it might otherwise.
Who says you will?
Nothing in the C specification says that.
Real life has no machines with 64-bit ISAs and 2^64 bytes' worth of resources.
Real life isn't portable. Who knew.
If there's no guard pages, no segmentation faults, no OOM killers, etc., then an allocation checker will cause failure in much the same kinds of undefined ways as not having the checker.
A segmentation fault is what you can expect in real systems, and it's absolutely better than undefined behavior.
You can do this portably, in C. It may not do you any good, and it may waste CPU cycles and memory bus usage. But it shouldn't harm. If the alternative is to not use alloca() or VLAs, and to make it so you can statically know your program's resource needs (which is very limiting, and not really attempted outside contexts where that such limitations are very strongly desired), then do that.
We all know what you're saying works in practice on most machines.
But it's not pure C and it's not portable if you follow the spec.
Also, in the absence of stacks, there's no "direction in which the stack grows", and any attempt to determine that direction will yield a nonsensical result, but doing memset() in one direction or the other will also then be equivalent anyways.
EDIT: There is one non-portable thing here: the assumption that machines are more limited than size_t implies. However, that's a distinction without a difference for any 64-bit ABIs or larger because all such machines must be more limited than size_t implies. That there is any failure mode when resources are exhausted is a non-portable assumption is uselessly true. Not being able to assume defined resource exhaustion failure modes would be destructive -- that the C spec does not is irrelevant.
This is undefined behavior.
See n1570.pdf, Annex J page 560, J.2 Undefined behavior
Pointers that do not point to the same aggregate or union (nor just beyond the same array object) are compared using relational operators (6.5.8).
So such a simplistic stack growth detector may get the wrong answer, but it's not undefined behavior:
char stack_grows_up(uintptr_t callers_local) {
char my_local = (uintptr_t)&my_local > callers_local;
return my_local;
}
EDIT: Ah, and uintptr_t is an optional type. If the platform doesn't have it (e.g., because it's a segmented memory architecture) then this won't work, so it's not portable, but aside from that it's not undefined behavior, it's just that the computation of stack growth direction may well be wrong.Nope. Just today I've seen a stack corruption bug that instead of producing a segmentation fault caused the process to go into uninterruptible sleep, somehow.
There's no real way to answer this question at that layer of abstraction; the minute you add any library calls that deal with a stack at all (such as alloca), we're no longer talking about pure C.
for (size_t i = 1; i < WHATEVER; i <<= 1) {
char vla[i];
vla[0] = 0;
vla[i - 1] = 0;
...
}
that is, you can still exhaust your stack trivially with VLAs.Sure, the two things have a different "API", but as far as allocating arbitrary amounts of memory goes, they both let you do it.
In the rare cases where VLAs might prove useful, e.g. a vector math library (IIRC, a major reason for VLAs was to help FORTRAN programmers transition to C), this difference absolutely does matter in reality.
I don’t think you can portably (1) do that. You have to allocate before you can touch those bytes.
Also, if you do before, I don’t think there’s a guarantee that you won’t hit some volatile memory that interacts with hardware, and of course, there’s no portable way to get the stack pointer, in the first place (if only because it, conceptually, doesn’t exist in C)
(1) specific compilers can safely do that because they know how they are implemented, and many compilers are fairly similar, so one can make this “quite portable”, but not really portable.
If you can assume a linear stack (which is kinda the whole point of this discussion) then you can call a function that will repeatedly alloca() in small chunks and touch a byte and return on success or SEGV on failure, and then when it returns it will release that memory but then you know you can safely alloca() that much.
Excellent point being made about C, though. Amazing how we've lived with this limitation for so long. I get that C is supposed to be portable assembler but the cases where stack wasn't used for return address and local var frames are rare so some support should have been added early on.
cat <<EOF>demo.c
#include <sys/resource.h>
#include <stdio.h>
#include <unistd.h>
int main(int argc, char *argv[], char *envp[]) {
struct rlimit rlp = { 0 };
getrlimit(RLIMIT_STACK, &rlp);
printf("stack size=%llu\n", rlp.rlim_cur);
if (rlp.rlim_cur < rlp.rlim_max) {
printf("setting stack size to %llu\n", rlp.rlim_max);
rlp.rlim_cur = rlp.rlim_max;
setrlimit(RLIMIT_STACK, &rlp);
execve(argv[0], argv, envp);
}
}EOF
make demo
./demo
Similarly, there is pthread_attr_setstacksize(3), which does the obvious thing for threads.
One of the projects I was to go in the future is implement boringcc [1], and this sounds like something I would want in it.
The question then becomes: would it even be possible to do?
[1]: https://groups.google.com/g/boring-crypto/c/48qa1kWignU/m/o8...
It could disable some (many?) optimizations, but I don't see it as a particularly hard thing to do (at least conceptually).
The C specs do not define any resource exhaustion failure modes. But actual operating systems, ABIs, platforms -- they do.
If you are on a system that gives you a SIGSEGV but no way to catch it on an alternate stack -or if you don't catch or ignore it- then your process will crash. This is fine. It's better than memory corruption.
But if you progressively write into the alloca()ed area in the direction of stack growth then you cannot corrupt other mappings, and you can only segfault if you exceed your stack's maximum size. Of course, there can be platforms with no SIGSEGV or whatever.
#include <stdio.h>
int main(int argc, char*argv[]) {
printf("ERROR: resources exhausted!");
return -1;
}I believe Scheme basically only uses the heap (which is necessary because of call/cc iirc).
Haskell uses a variation of a G machine (looks like a spineless tagless g machine to be specific). The way that works is that they convert your program into a giant lambda term and then they work on reducing it. This of course looks nothing like a traditional call stack, so you can replace that whole thing with the G machine. Recursion (at least of the non-tail recursive variety) will look like a lambda term slowly getting bigger in heap space.
And then I also heard of someone (I forget which language) talking about setting up some logic to look for stack overflows and then throw all the current stack onto a heap and then start over.
https://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.54...
If the language is based on recursion, you can't just hope it happens commonly enough, it has to happen all the time, even for corecursive functions.
Except I've never been super clear on Haskell having TCE or not, because laziness means there are limited opportunities for it to even fire.
It's surprising how hard it can be to track down a stack-caused bug when you've got a bunch of interrupts banging on it.
Generally it's a little unnerving to just pray that you've got enough stack given it's hidden usage and the possibility for things like recursion. Kind of a shame considering what a slick way it is to deal with memory leak avoidance.
It's worth noting that the limits are adjustable on a per-package basis in Alpine based on the compiler-time link options.
(okay, by "clean" I do mean littered with `.await`, but it's still structurally clean)
That said, you can make kernel-aware user-controlled context switching much faster than it is currently: https://www.theregister.com/2020/08/10/google_scheduling_cod...
Also having the addressable method arguments in the same space as the return address, allows a whole host of other bugs and cracks.
I have long thought, the return stack should be a memory region un-addressable by the code except via call and return.
Local method arguments could then revert to ordinary allocations or some other check-able allocation, and have normal error recovery.
If you’ve formally verified that function X uses at most Y pages of stack, and the remaining stack is smaller than that, you could return an OUT_OF_STACK error instead of calling it. Most libraries/functions wouldn’t need this, but there are C libraries which place an extraordinary importance on reliability, and it seems like it would be a useful feature for them.
There might be other reasons why these functions shouldn’t be in the standard (or are impossible on some platforms), but ”the information is worthless” is not one of them. Of course this information can be useful.
There are more hacks in Heaven and Earth, Horatio, than are dreamed of in your philosophy.
Moreover, every nontrivial C program in practice contains UB.
Nevertheless, it does matter what is and isn't standardized; standardizing such functions would greatly improve the situation for programs like those I mention, because they would be able to rely on the C standard to find out when they're about to run out of stack, an event they already have code for handling. It's true that a pathological C library implementation could still totally break their stack-exhaustion-handling code, and in fact that's already true, but fortunately there are lots of C library implementations that behave well enough in practice.
2. There are very few knobs you can tweak to control how your call stack space is allocated or represented in memory. With a stack container, you have much finer control over that - you can do reallocations or define custom error handling when you overflow the container, you can deallocate/shrink the container when you no longer need the space, you can serialize the container to disk to make your processing resumable, etc.
3. A call stack has a very limited set of operations. You can only access data from the current stack frame, and you can only push/pop (call/return) stack frames. But with your own stack-like data structure, you could extend it to do far more, e.g. accessing data from the first or last n traversals, or insert/remove multiple elements at once.
The meta-problem is what to do with the information.
It's the only sane interpretation I can give to a standard that hides any way to manage stack usage - the compiler/OS layer takes responsibility.
Certainly a few MB per thread isn't too crazy though; at 8MB per thread, you can have a million threads and use only a tiny fraction of the address space available.
Effectively, you can make sure the program runs out of available memory long before a stack exhaustion issue, unless the stack usage pattern is severely unbalanced for some thread.
2. The actual size of the page-table will depend on the page sizes supported by the architecture. On modern x86_64 for programs with less than 32k threads, It might make sense to use 1GB stacks with 1GB guard pages, for just 2 PTEs per thread, less than the 5 needed for 8MB using 2MB pages.
Ancillary data should only be accessed by the macros defined in cmsg(3).
The problem being, the code I was writing was in C# which can't consume C headers or evaluate these macros.
Embedded applications such as automotive, or industrial control come to mind (or even consumer electronics, with less life-or-death consequences possible)
OTOH, there seem to be a lot of "we don't control anything" IoT devices which are random overly powerful processors with random uboot+linux kernels+rootfs's running my_first_c_program.c level applications. All of it tossed together with very little engineering effort, and frequently less low level knowledge, and left to rot because its basically unmaintainable, even if it happens to be the 1% of devices that can get over the air updates.
You can fill the stack region with sentinel data and periodically scan it to warn of impending doom.
I've heard that huge pages might help, unfortunately they're somewhat harder to use.
Of course you can imagine a language where your function signature includes how much stack space it takes up. The stack size being recursively calculated by looking at all the functions your function calls etc.
If recursion is included, then you're correct that this would require solving the halting problem. The obvious solution would be to just not allow recursion. Although, if you also consider languages like Idris (which are technically sometimes not turing complete) you can setup a termination checker and then certain rules can give you an upper bound on how long a recursive function will run (and therefore how much stack it takes up).
For all of these problems there are more complex analyses that can be done to determine e.g. maximum recursion depth or possible functions that an indirect call could resolve to. And failing all of these there is usually a way for an engineer to override or provide this information.