Implementation of GCC's Nested Functions (vs. C++ Lambdas)
uecker.codeberg.page
uecker.codeberg.page
Without all this nested functions are as useful as the "rewritten" examples in the article, one can easily do that by hand without any compiler or language support.
This problem doesn't arise with C++ lambdas because you pass them around as special objects, not as bare function pointers.
For my part, I'll point out that there is one rather important difference between nested functions and C++ lambdas that the author completely ignores, as exhibited by this godbolt example: https://godbolt.org/z/35beWrrTe (note the differences in the generated assembly, especially that which cannot be explained merely by -O0 code generation).
Do you mind elaborating for those of us who aren't familiar with what to expect from the compiler?
If your argument is that ABI doesn't matter, then by all means, propose a patch to GCC to change the ABI and see if it gets accepted.
I do not want to change the ABI for nested functions as it is a useful ABI and a cross-language standard. ABI obviously matters, but it is not a fundamental difference in implementation that makes nested functions fundamentally different to C++ lambdas. If your point is merely that a compiler translating nested functions to lambdas would need to adapt the ABI in this case, I agree. This is not difficult though. And this question is relevant only if one allows taking the address of a nested function, because as long as it is called only locally, the compiler can use whatever ABI it wants.
> And this question is relevant only if one allows taking the address of a nested function
And that question is very relevant since taking the address of such functions (to pass to other functions, e.g., qsort) is one of the main use cases for their existence.
If you take the address of a lambda function you get a pointer to an object of anonymous type, so you can not pass it to qsort, and qsort would also not know how to call this.
But if we added a feature similar to std::function_ref to C (i.e. a wide function pointer type), then such a type could be used to call both, nested functions and C++'s lambdas, and - in fact - many callable entities from other languages too. But for C++'s lambdas this would always involve a compiler generated thunk that adapts the ABI. This is also exactly what happens in C++ if you use std::function, because even in C++ you can not pass the address of lambda to a function without first erasing the type and creating the thunk.
So there is no compatibility problem.
https://godbolt.org/z/vEP5G9Pfr
Similar to how std::function_ref creates a thunk in C++ that calls the lambda so that it has a generic type-erased API that can be passed to non-templates, a conversion to a wide pointer would create a thunk that adapts the call from the nested function pointer ABI (that already exits for other languages also in LLVM whether we standardize the C feature or not) to whatever the lambda needs.
Edit: slightly updated example.
Showing things that are optimized by fully inlining into one function don't actually demonstrate equivalency, because you end up omitting anything that might actually evidence a difference in the semantics. And I get that, for your use cases, those differences might not matter. But as a compiler engineer, I can't say that only those use cases matter and therefore they're equivalent for all practical purposes.
It is also very unhelpful when you insist that this is "compatible" with other languages, where "compatible" actually means "compatible, if you put in a bunch of work in both languages to make something that makes them compatible, none of which I'm actually describing." Especially when there are competing proposals that do have compatibility in the sense of "I don't have to modify the C++ compiler to let it use this thing."
My point is that the implementation is structurally very similar: You synthesize a structure and put it on the stack and then pass a pointer to it around. It is so similar that you can certainly reuse your implementation of the lambda feature to implement this.
The wide pointer ABI question is also entirely orthogonal to other aspects, so we could decouple this discussion. The advantage of being compatible to other languages is because if we would use a common ABI that many other languages also use: Ada, Go, D, etc. In this case, no additional work has to be done for any of these languages. LLVM also supports this already.
The issue with C++ is that it does not use this common ABI and the closet thing it has even as a suitable API is std::function_ref. Where lambdas are not called locally, C++ already needs to create thunks anyway by going to some kind of these adaptors, so there is also no additional burden on the C++ side. One would simply have to implement this thunk in a slightly different way to adapt the calling convention.
I am not even sure that you need to do less adaption for other proposals, as many things the C++ semantics rely on do not exist in C (callable objects, templates), so you also need to adapt anyway at least in how you expose them in the language and in what other features you may need to make it work. In particular, JeanHeyds proposal does not even include the wide pointer part yet, which will also then be required at some point. The proposal also exposes far more features (different ways to capture), which makes it more work.
But I fully realize that the opposition for everything that looks different to C++ from the clang side comes from the perception that it is more work on your side. I can sympathize, but note that in GCC or other compiler that do not have a shared FE, we would essentially have to implement everything from scratch. So let's discuss this more if you want.
I have looked it up and I can already tell you that Go is not using the ABI you would be proposing. I cannot speak for the other languages.
> But I fully realize that the opposition for everything that looks different to C++ from the clang side comes from the perception that it is more work on your side.
That is not where the opposition comes from, and for as long as you continue to believe that, you will fail to understand the opposition at all.
"Closure calls follow the same conventions as static function and method calls, with one addition. Each architecture specifies a closure context pointer register and calls to closures store the address of the closure object in the closure context pointer register prior to the call."
In any case, the documented use of __builtin_call_with_static_chain in GCC and Clang is to be able to call closures of other languages, and for GCC Go is explicitly mentioned.
GCC: https://gcc.gnu.org/onlinedocs/gcc/Constructing-Calls.html
"This built-in can be used to call Go closures from C, .."
https://clang.llvm.org/docs/LanguageExtensions.html
"... as used by some language to implement closures or nested functions."
Yes, it is true that I completely fail to understand the opposition to this.
BTW: I also very carefully analyzed all the semantic differences, which led me to the conclusion that putting C++ lambda semantics into C would be a bad idea. one can read this here: https://www.open-std.org/jtc1/sc22/wg14/www/docs/n3654.pdf
I also do not understand why you think the generated assembly needs to be identical, or why this is a "rather important difference".
If they don't capture anything, I believe you _can_ pass them as bare function pointers.
This one explains how you can avoid the use of trampolines in GCC 17: https://uecker.codeberg.page/2026-07-14.html
The documentation is here: https://gcc.gnu.org/onlinedocs/gcc/Constructing-Calls.html
There are many language features can be rewritten by hand into other simpler forms, you could rewrite loops into gotos, C++ objects into structures, etc. This does not show that those things are not useful. But the point of the article was not to show why nested functions are useful, for which I would certainly have used more interesting examples.
Also misses that the way C++ lambdas work is that one design requirement was that they should be syntax sugar for the functor[0] pattern from C++98.
[0] - Not to mix with ML functors, rather classes with call operator overloaded.
Thus, to access stack variables two enclosing functions up, the static link is walked twice.
A reference to a nested function in D is represented by a pair - a pointer to the function, and the static link. (Called a "delegate" in D parlance.) Interestingly, this is the same layout as taking a reference to a member function, where the "this" pointer takes the place of the static link.
This means that references to nested functions are ABI compatible with references to member functions.
Lambdas in D are just a more compact syntax for nested functions.
[1] The lambda function probably has the same linkage as the function the lambda is contained in, which likely isn't "the only copy of this function is in this TU" but rather "this function may appear in several TUs, but all of these copies are equivalent and you can pick whichever one you like as the actual body." Very different opportunities there!
Generally this is a bit nicer than having explicit lambdas, but I thought the 'best-case' scenario would be if GCC saw into the stack layout of the calling function and could manipulate the calling functions stored stack variables (and saved registers). After all, a debugger can track what variable goes where at every line of code, so this can be done.
Not sure if this would be useful or practical, but would be a nice bit of nerd cred.
A modern compiler IR is probably going to be an SSA-based infinite virtual register set model. In such a model, any variable without its address taken ends up being a register (which may happen to be spilled to the stack). Referencing the variable via a nested function means its a local variable whose address escapes, which kills a lot of optimization potential. It's probably possible to adjust SSA to handle this, but it's a lot of work for little benefit, especially since closure models (closures being regular objects with an unnameable type and an overloaded call operator) have taken over nested functions in language design and thus it isn't really applicable for modern languages.
> After all, a debugger can track what variable goes where at every line of code, so this can be done.
Variable value tracking breaks down pretty much the moment any optimization happens.
> Variable value tracking breaks down pretty much the moment any optimization happens.
I think for nested functions it's not [only] a question of performance/optimizations but of correctness. Even if you properly unwind the stack like a debugger trying to find the parent frame, you may actually find multiple due to recursion. It's impossible to know which one is "yours" unless at least some information about the parent frame was passed to the nested function at invocation time. It can't be a truly static function.
The idea that "closures ... have taken over nested functions in language design and thus it isn't really applicable for modern languages." is true only if you think C++ as basically the only modern language and everything else with nested functions is not modern, which is ... an interesting take.
I'm shocked that you didn't even pick up Rust, given how frequently I mention it on the WG14 reflector.
I have test cases in TXR where qsort is used to sort an array of strings:
https://www.kylheku.com/cgit/txr/tree/tests/017/qsort.tl
https://www.kylheku.com/cgit/txr/tree/tests/017/qsort.expect...