Fast Shadow Stacks for Go
blog.felixge.de
blog.felixge.de
It will not work with Go as-is because the Go scheduler will have to be taught to switch the shadow stack along with the regular stack, panic/recover needs to walk the shadow stack. But for binaries that do not use CGO, it would be possible to enable this fairly quickly. Hardware support is already widely available. The SHSTK-specific code paths are well-isolated. You would not lose compatibility with older CPUs or kernels.
What does the API for accessing the shadow stack from user space look like? I didn't see anything for it in the kernel docs [1].
I agree about the need for switching the shadow stacks in the Go scheduler. But this would probably require an API that is a bit at odds with the security goals of the kernel feature.
I'm not sure I follow your thoughts on CGO and how this would work on older CPUs and kernels.
Regarding older CPUs, what I wanted to point out is that the code to enable and maintain shadow stacks will not be smeared across the instruction stream (unlike using APX instructions, or direct use of LSE atomics on AArch64). It's possible to execute the shadow stack code only conditionally.
Note that it's actually GLIBC_TUNABLES=glibc.cpu.hwcaps=SHSTK to enable it with Fedora 40 glibc (and the program needs to be built with -fcf-protection).
I suppose hardware support (e.g. dedicated call/ret instructions accessing a second stack pointer) is also desirable if you want to do this sort of thing. I recall that Itanium had something like this.
push RBP
mov RBP, RDSP ; DSP stands for "data-stack pointer"
sub RDSP, <frame-size>
... ; locals are [RBP-n], args are [RBP+n]
add RDSP, <frame-size>
pop RBP
ret ; pops address from [RCSP], where CSP stands for "call-stack pointer"
or even, if you don't want a specific base pointer, like this: sub RDSP, <frame-size>
... ; locals are [RBP+n], args are [RBP+frame_size+n]
add RDSP, <frame-size>
ret
In fact, you can do this even now on x86-64: make a second stack segment, point RBP there, and use RBP as if it were that hypothetical RDSP register. The return addresses naturally go into the RSP-pointed area, as they always did. The only tricky part is when you need to call the code that uses SysV x64 ABI, you'll need to conditionally re-align that RSP-pointed area to be 16-bytes aligned which is a bit annoying (in your own code, having RBP always aligned at 16 bytes is trivial since function calls don't break that alignment).There are two main benefits of this layout, as I see it: first, smashing the return address is now way harder that it normally is; second, "the act of traversing all call frames on the stack and collecting the return addresses that are stored inside of them" is now trivial, since the call stack at RSP is now just a dense array of just the return addresses, no pointer chasing needed.
Btw, I think Forth has a separate data stack and return address stack?
On top of that if you have work-stealing task scheduler, that if your current thread waits on I/O reassigns it to some other task to work it may bulb even more.
Scheduling tasks on a work-stealing queue tends to flatten the stack first, eg. a halted task has a nearly empty stack if it's an async coroutine. The scheduler itself may have 3-4 calls deep before resuming the task itself, so that doesn't seem too bad. Unless you're talking about full blown threads.
But then, if you capture the stack (using sampling), and catch for example ThreadContext switch while CreateFileA is happening, you can expect lots of callstack items going all the way down to the kernel
To answer your first question: For most Go applications, the average stack trace depth for profiling/execution tracing is below 32 frames. But some applications use heavy middleware layers that can push the average above this limit.
That being said, I think this technique will amortize much earlier when the fixed cost per frame walk is higher, e.g. when using DWARF or gopclntab unwinding. For Go that doesn't really matter while the compiler emits frame pointers. But it's always good to have options when it comes to evolving the compiler and runtime ...
Do you see this speeding up real world Go programs?
The low hanging fruits to speed up stack unwinding in the Go runtime is to switch to frame pointer unwinding in more places. In go1.21 we contributed patches to do this for the execution tracer. For the upcoming go1.23 release, my colleague Nick contributed patches to upgrade the block and mutex profiler. Once the go1.24 tree opens, we're hoping to tackle the memory profiler as well as copystack. The latter would benefit all Go programs, even those not using profiling. But it's likely going to be relative small win (<= 1%).
Once all of this is done, shadow stacks have the potential to make things even faster. But the problem is that we'll be deeply in diminishing returns territory at that point. Speeding up stack capturing is great when it makes up 80-90% of your overhead (this was the case for the execution tracer before frame pointers). But once we're down to 1-2% (the current situation for the execution tracer), another 8x speedup is not going to buy us much, especially when it has downsides.
The only future in which shadow stacks could speed up real Go programs is one where we decide to drop frame pointer support in the compiler, which could provide 1-2% speedup for all Go programs. Once hardware shadow stacks become widely available and accessible, I think that would be worth considering. But that's likely to be a few years down the road from now.
That being said, I'm sure there are a lot of remaining incremental optimization opportunities that could add up to 10% over time. For example a faster map implementation [1]. I'm sure there is more.
Another recent perf opportunity is using pgo [2] which can get you 10% in some cases. Shameless plug: We recently GA'ed our support for it at Datadog [3].
[1] https://github.com/golang/go/issues/54766 [2] https://go.dev/doc/pgo [3] https://www.datadoghq.com/blog/datadog-pgo-go/
As a developer I like that approach as it keeps a great developer experience and helps me stayed focus and gives me great productivity.
Though I find it unfortunate that the industry considers Go as a choice for performance-sensitive scenarios when C# exists which went the above route and does not sacrifice performance and ability to offer performance-specific APIs (like crossplat SIMD) by paying the price of higher effort/complexity compiler implementation. It also does in-runtime PGO (DynamicPGO) given long-running server workloads are usually using JIT where it's available, so you don't need to carefully craft a sample workload hoping it would match production behavior - JIT does it for you and it yields anything from 10% to 35% depending on how abstraction-heavy the codebase is.
I'll probably bring it up in the by-weekly Go runtime diagnostics sync [1] next Thursday, but my guess is that they'll have the same conclusion as me: Neat trick, but not a good idea for the runtime until hardware shadow stacks become widely available and accessible.