ABI Mistakes
elronnd.net
elronnd.net
If the concern is the overhead of copying large structures around, the obvious solution is to not do that and simply pass structure pointers (like most APIs do). This also gives the API implementer (i.e. callee) full control over whether copies get made or not.
In fact, I think I’d argue that if you’re passing large structures by value, Your API Is Probably Wrong and this is absolutely not the ABI’s problem in the first place.
Yes, and what's wrong with that, if it's transparently done by the compiler?
In this scheme, the caller would pass non-register-sized data "by value" by actually passing a pointer to it. (Importantly, this pointer is register-sized, and so wouldn't necessarily need to be spilled to the stack—unlike caller memory copy, which is always to the stack.)
A callee that can be statically determined to never modify the value, would, under this calling convention, be compiled to code that simply works with the data "through" the pointer that's been placed in its register / on the stack.
A callee that can't be statically determined to never modify the value, would, under this calling convention, generate a memcpy from the pointer onto the stack; where the pointer then goes dead at that point (and so, if the pointer was spilled to the stack by the caller, then the memcpy could be targeted so as to overwrite the pointer on the stack.)
This would be much more efficient under most conditions. Its only inefficiency would come from the situation where the pointer being passed would necessarily be spilled to the stack; and then the callee would be statically guaranteed to make a copy. In that case, you're doing an extra push (for the pointer) that current ABIs avoid by just having the caller do an eager copy.
But this would be quite rare in practice, since this ABI (and the ABIs it replaces) are only for external symbols — the kind whose linkage is fundamentally dynamic (i.e. where the linker-loader could in theory substitute anything it likes for the external symbol with an LD_PRELOAD shim.)
Internal symbols — private static functions — get bespoke compiler-specific ABIs that skip all this enforced caller/callee predetermined role business and just codegen whatever's most efficient on a case-by-case bases, with different callsites getting different monomorphizations or partial inlinings of the callee that distribute the responsibilities differently.
And since these ABIs only matter for external linkage, you have to remember that symbols with external linkage get no "type attribute enforcement" at link-load time, and so your symbol for a function that was originally guaranteed to be a callee-that-always-writes, might be substituted at link-load time for a callee-that-never-writes. In which case, having the pointer on the callee's stack once again becomes handy, rather than useless.
> In fact, I think I’d argue that if you’re passing large structures by value
Not necessarily large structures. More often things like UUIDs—values just large-enough to not fit into a (64-bit) machine register. It's these small memory spills, done in hot loops, that add up. There are huge wins for the cleanliness of e.g. RDBMS engine code, if their various 128-bit to 512-bit column types can be passed by value without necessitating an eager copy.
If the ABI specified that objects passed this way were passed by immutable pointer, then just passing the object using this new method doesn't generate a mutable pointer, so the compiler is free to assume nobody else can modify it.
So not really an optimization anymore.
The caller is required to do a copy as the ABI is right now so if you can't do this optimisation it's only as bad as it was before.
Yes, there is a case where this creates one extra copy. The case where you a) give away a mutable pointer to a variable before calling another function, and b) the called function modifies it's own argument and then c) that function then doesn't use that object as the argument for another function it calls.
That's the only case where it generates an extra copy, and I'm going to go out on a limb and say that I'm almost cirtain that this will reduce sigificantly more copies that it creates.
And it’s not limited to mutable pointers. Sure, if the variable itself is marked const, then it’s UB to modify it. But const local variables are rare in most codebases, in my experience. If on the other hand you pass the address of a non-const variable as a function argument, even if the type of the argument is a const pointer, the function can legally cast away the const and write to the pointer.
Still, I agree that this would probably eliminate more copies than it creates. But I don't think it would eliminate or create very many copies. In my experience, approximately nobody writes code that passes structs by value in the first place.
"const", as a qualifier, isn't novel information about a variable for the compiler — compilers already know whether a variable is used mutably or not, just like they know when and where a variable has its last access and goes dead.
"const" is an instruction to the compiler that if it finds, by its own independent analysis, that the variable cannot be const in that context, then the program is invalid and compilation should error out.
As such, "const" only even has a meaning, if the compiler already independently knows which declared variables are or are not const "in practice."
That information is what would be used to decide whether to do the copy.
> B could end up having to perform another copy for that call
At worst, that's no worse than what happens under the current ABI, where A passes data by value to B, which passes it by value to C.
(Do also remember, as I said in another comment, that ABIs/calling conventions are only "about" functions with external linkage. Functions with "static" linkage do whatever the compiler likes; as do their callers.)
I think, in the specific situation you're describing, there's also prelink-time cross-unit optimizations that can be done here to eliminate redundant copies even further — but that's getting close to just expecting C to do Rust-like ownership analysis, so I wouldn't claim that to be practical in any production compiler in the near future.
Also, you are masquerading an access trough a pointer as a direct access to some data. An access trough a pointer is less efficient than a direct access, and thus one could assume that is doing something efficient while in fact it isn't.
Finally, who guarantees that the programmer doesn't abuse this thing and starts to modify the data without doing the copy on write? On purpose or by mistake, the effect is that passing by value something will result in that something being mutated.
In the end, what is the problem of passing a pointer? It's more explicit and you know what is going on.
> There are huge wins for the cleanliness
What cleanliness? Because you think that
void f(T *x);
f(&x);
is not clean? Well, then use C++: void f(const T &x);
f(x);
You cannot design an ABI that is supposed to be used by all languages on the fact that doing something in a specific language is more ugly than doing some other thing. If it's a problem of C, we should fix C, not change the ABI to change the semantic of C.But the alternative is not "a direct access." It's a copy and then direct access. A copy, mind you, that occurs through a pointer to grab the original data from the stack position it was sitting in. That's the very same indirect read that the callee does in the proposed ABI!
Thus, the indirect memory read will happen either way; either at caller copy time (existing ABIs), or at read time (proposed ABI, copy optimized away), or at callee copy time (proposed ABI, copy not optimized away.)
The only difference in the number of indirect memory accesses, with the proposed ABI, is that you might be able to avoid an indirect memory write (i.e. of the memory the caller is writing to by doing the copy.)
> Finally, who guarantees that the programmer doesn't abuse this thing and starts to modify the data without doing the copy on write?
Uhh... we're not talking about features exposed to the developer. We're talking about calling conventions—instructions to the compiler on how to distribute the glue code it generates between the call-sites it builds ("caller-[verb]ed" things) and function prologues/epilogues it builds ("callee-[verb]ed" things.)
The "copy on write" part here [if you want to call it that—it's not "on" write] would be part of such a generated function prologue ("callee-copied")—the same code that gets generated at the call-site in existing calling conventions ("caller-copied"), just moved to the other side of the linkage. With the important difference that, if the compiler could prove that the callee (or anything it calls, passing the variable, transitively) wasn't going to do anything write-like with the data behind the pointer, then the copy in the function prologue could be optimized out.
This compiler optimization can only occur if the copy is the callee's responsibility. Given external linkage (think: separate compilation units, e.g. calling a function defined in a static library that's already compiled and only available in binary form), the caller can't guarantee the properties of the callee. But a callee can always guarantee its own properties during function-prologue/epilogue generation—since it's generating the function prologue/epilogue in the very same compiler session that has access to the AST that describes the function's static-analysis-proven properties in full.
(Yes, this means that a compiler could be mistaken about whether a callee writes through the implicit pointer underneath the passed value, and therefore generate the optimized function prologue when it shouldn't. This is known as a compiler bug, and wouldn't be something J Random Hacker would ever run into if using a stable, non-alpha release of their compiler, with a stable, non-alpha release of support for this calling convention.)
> What cleanliness?
The functionality being proposed is that you just pass-by-value (which makes obvious that the semantics your code is enacting is "the function can do whatever it likes with the value, and it won't affect the caller, because if it writes, it's writing to a copy of the value"); but sometimes it'll be automagically faster than it could have otherwise been, because sometimes, when the compiler can prove certain properties, no copy — only an indirect memory read — would be occurring.
This allows you to always use pass-by-value when you want those semantics, and to get code that's "as optimal as it can be"; rather than having to pass pointers around and explicitly copy-on-write and all that mess in order to get that same level of optimization.
(And yes, that's the alternative in heavily performance-oriented code like a database's tuple materialization logic. You don't get to write clean code at the expense of performance. Your "options" are to write performant code using dirty hacks; or fix the conventions you're coding on top of so that performant code doesn't require dirty hacks.)
Depends. If the structure was constructed in the caller and not needed afterwards the compiler could choose to directly construct it in the stack space used for the function call.
To be honest, the first option is much better for readability. Visual inspection of the caller lets the reader know whether or not f can modify x.
That second option is horrible for the maintenance programmer who is reading the caller; whether or not you are interested in what the function f() does, you need to actually look it up to see if it is allowed to modify x.
I remember how I had to solve a pile of thread bugs in a C++ project by changing string assignments into iterator constructors which would bypass the GNU CoW reference counter.
They tried, and tried, but never quite succeeded in fixing every single possible race condition with the reference count CoW implementation.
Copy-on-write is "the called function receives the original, but if it needs to write, then there's explicit code that makes a copy and replaces the reference to the original with a reference to the copy."
What we're talking about here is:
• A function is being compiled. It receives one of its formal parameters "by value." That parameter's data is larger than a machine-register in size.
• The compiler uniformly generates the call-sites of this function to provide the address of the original data, passing it in a register. The call-sites of the function look the same whether or not the function does the copy.
• The compiler, by static analysis of the function, determines whether the function really needs a copy—i.e. whether it would make a difference to the code whether the copying ever happens.
• If it needs a copy, then the compiler generates the function to have a prologue that does a copy from the address in the caller-passed register to the address in the stack pointer; and then the compiler generates code that interacts with the copy as a local stack variable. (Because, after the copy, it is!)
• But if the function doesn't actually "need" a copy, then the compiler generates no copy in the function prologue, and generates code for the function body that interacts with the original indirectly through the passed address in the register."
This has nothing-at-all to do with copy "on" write. This is "copy now, if we can't guarantee there won't be a write later."
Which is already what compilers do. That's what "pass by value" means. It's just that current calling conventions restrict compilers from being able to ever make that guarantee; whereas under the proposed calling convention, they would sometimes be able to make that guarantee.
Sane ABIs already pass by-value parameters via up to two registers (or via vector registers). One register is just to small for many common things like 3D/4D vectors or start+end pointer pairs.
> Internal symbols — private static functions — get bespoke compiler-specific ABIs that skip all this enforced caller/callee predetermined role business and just codegen whatever's most efficient on a case-by-case bases, with different callsites getting different monomorphizations or partial inlinings of the callee that distribute the responsibilities differently.
Unfortunately this is still an area where compilers could do a lot more.
Yes that’s the point the author is making. His scheme incurs strictly less overhead than the scheme currently used.
Some people may want to program in a “value oriented” way without pointers at the C level for various reasons but for efficiency reasons they cannot.
But there is a larger assumption that underlies your post that I think is important to correct.
You seem to be caught up in the idea that the code you write, and the code the compiler generates must have some kind of one-to-one corrospondance, but this isn't the case at all.
The code the compiler generates must have the same final result, but it doesn't have to actually work the same way.
So if you pass a structure by value, and the compiler can avoid a copy somehow because it's more efficient, then of course it should do so.
Making an ABI that is ammenible to allowing the compiler to make better optimisations is of course also a good idea. And the proposed ABI here achieves that.
You're obviously correct, but I think this ignores downstream costs of sufficiently magical compilers. If my compiler tunnels into a parallel universe, executes the code there, and returns the result, that's fine, but it's going to be enormously painful to debug.
Obviously very few developers (including myself) have an appropriate mental model of highly complex, speculative processor architectures that end up executing their code, but I do think there is at least some benefit to having a correspondence between the code you write and the code the compiler generates.
Maybe it's not a problem with the compiler, but a problem with languages. Maybe we just need better languages?
> Obviously very few developers (including myself) have an appropriate mental model of highly complex, speculative processor architectures that end up executing their code, but I do think there is at least some benefit to having a correspondence between the code you write and the code the compiler generates.
It’s not really clear to me why you think this? Or at least, it’s not clear to me how you can write this without levelling the same complaint against essentially any C compiler written after the mid 1980s?
Because for all the hand—wringing in the comments here about compilers generating code that doesn’t do what your C said it would do in the way you wrote it, that ship sailed decades ago.
That quaint little integer addition loop you write in C will, as like as not, be exploded by the compiler into an unrolled stream of SIMD instructions and packed registers that doesn’t resemble what you wrote in the least, which the CPU may then further scramble and interleave, resulting in nothing remotely like what a naive person might imagine is a reasonable translation of what you wrote – and yet, debuggers remain broadly useful tools.
And people have for decades now shown a strong desire to enthusiastically and with great speed drop compilers in favour of newer ones that performed better optimizations, which is to say, people have shown a strong preference for compilers that are good at generating code that looks nothing like what you wrote, provided the results are the same and the performance is as fast as possible.
Turning a copy into passing a pointer that may later copy if and only if the compiler sees the possibility of a write coming is, in terms of modern optimizations, an absolutely conservative optimization that doesn’t routinely happen only because platform ABIs unnecessarily rule it out. Far from being a bridge too far, in terms of modern compiler behaviour it’s simple and low-hanging fruit which isn’t exploited even though far more exotic tricks are already the rule of the day.
It doesn't, but when it comes to debugging, it would help to have them similar.
[expr.call]: "The lvalue-to-rvalue, array-to-pointer, and function-to-pointer standard conversions are performed on the argument expression."
[conv.lval]: "... if T has a class type, the conversion copy-initializes the result object from the glvalue."
The way a program can tell if a compiler is compliant to the standard is like so:
struct S { int large[100]; };
int compliant(struct S a, const struct S *b);
int escape(const void *x);
int bad() {
struct S s;
escape(&s);
return compliant(s, &s);
}
int compliant(struct S a, const struct S *b) {
int r = &a != b;
escape(&a);
escape(b);
return r;
}
There are three calls to 'escape'. A programmer may assume that the first and third call to escape observes a different object than the second call to escape and they may assume that 'compliant' returns '1'.It's still a win because 1) you can avoid making copies in many places, and 2) code size decreases because the copy happens one time in the callee rather than many times for every caller.
In theory when they sort that out Rust should be able to turn on restrict where it applies globally "for free".
[0]: https://github.com/rust-lang/rust/pull/82834
LLVM is a large project that's mostly written in pre-modern C++, and "noalias" is a highly non-trivial feature that affects many parts of the compiler in 'cross-cutting' ways. It would be surprising if it did not turn up some initial bugs.
This means optimisation phases would need to explicitly opt-in and aliasing optimisations would commonly be missed by default, but it would avoid miscompilations.
Also, even without LTO, LLVM can perform the same ABI optimizations on functions that are local to a single translation unit and not exported to other ones. In Rust that means a function not exported across crate boundaries. In C that means a `static` function.
Since only the JIT sees the native code, alongside JIT caches[0] you have basically LTO and PGO data across runs and some call graphs might eventually be collapse into simple basic block, or structs might be reduced to the actual set of fields that are used.
[0] - This includes all modern JVM implementations including the Android cousin, and .NET to certain extent already with usability improvements coming in version 6.
Every major ABI is listed here as containing the same mistakes. I'm inclined to think the people who designed these ABIs were smart enough to understand the consequences of their design decisions.
I don't know whether this author is correct or not, but my gut is there is something missing here with respect to non local control flow (like exception handling, setjmp/longjmp, and fibers).
I don't really know enough to weigh in on this, but I can say that having pursued a lot of WTFish things in my career so far, 90% of the times I've encountered bad decisions, the explanation for it was either "it was done that way because legacy reasons" (i.e., it had to be done that way then, the reason it had to be has changed, and now it would break things to do it 'correctly') or "it was easier" (i.e., at the time the badness wasn't really going to affect anyone, or not measurably, or was very intentional tech debt, and it's only 'now' that anyone is noticing/caring).
The rules for this proposed ABI are exactly the same as the existing amd64-SystemV C ABI, with one difference: the stack-to-stack copies aren't generated at the call-site; instead, the generated code at the call-site passes the address (in a register, or spilled to stack) for what it would have copied. The compiler generates the stack-to-stack copy in the generated function's prologue, using the address it was passed. Nothing more, nothing less. It's just moving the required location for certain generated code across the linkage, and keeping a temporary alive a little bit longer to make that work. (And in exchange, the temporary that the local stack variable gets put in isn't created at the call-site, so the register-file "pressure" of the change is net neutral.)
This is no more or less complex than the current ABI. It doesn't create more exceptions or edge-cases than the current ABI. It doesn't make the ABI harder to implement. The only thing it does, is choose differently in the matter of a basically-arbitrary choice of where to put some generated glue code (the stack-to-stack copy).
The only practical upshot of this change, is that this enables compilers to sometimes do an optimization that they can't currently do, because doing said optimization would go against the rules of the amd64-SysV ABI (i.e. a caller that pushed a register instead of copying the value wouldn't be an amd64-SysV caller any more, and wouldn't be compatible with precompiled amd64-SysV callees any more; and vice-versa for the callee.)
But if-and-when a compiler does do that optimization, it's internal to the generated function. It doesn't mean that there are two potential callee "signatures" under the proposed ABI. There's only one.
Here's what the proposed ABI would probably say about stack copies:
> "The caller always passes large values by reference; the callee always receives them by reference. If the callee is taking a parameter pass-by-value, then it's up to the compiler of the callee to insert code into the callee's function prologue to turn the passed reference into a stack-local copy of the referenced data."
With that particular legalese, the callee's generated copy is still "required" by the spec, but its effects are now also "hidden" from the caller — i.e. its observable results are no longer leaking across the linkage. Therefore, the compiler is now empowered to optimize out the callee copy, as long as it can ensure the resulting code has observably equivalent results from the caller's perspective.
Note that this isn't anything the person implementing the ABI targeting code in the compiler has to worry about. They just write the code to generate a callee function prologue that does a stack-to-stack copy. It's the person writing the optimization pass that comes after that codegen step, who can now can take that stack-to-stack copy and — static proof of read-only access by the callee in hand — drop it out.
The optimization opportunity being enabled by the change, isn't part of the ABI's spec. The proposed ABI is just about moving the stack-to-stack copy into the callee. What the compiler chooses to do when targeting an ABI where the callee does stack-to-stack copies, is up to the compiler. Presumably, it will do "whatever fiendish things it can" at -O3, and "nothing much different" at -O0. Like usual.
And either way, the linkage itself looks the same. The optimization doesn't change the linkage. Any and all tooling that examines the linkage — debuggers, disassemblers, tracers, etc. — would see the same thing, whether the optimization has occurred or not. Because the optimization isn't part of the linkage; it's internal to the codegen of the callee, enabled by the (uniformly!) modified structure of the linkage.
I've also seen "bad" decisions made due to outside constraints. These decisions look like bad decisions, except that if you try to "fix" those decisions, it becomes a lot harder than it looks.
is just not possible. CPUs don't know about `const`. So you have to work with the assumption that functions that you call can do anything to their arguments. Thus copies cannot be avoided.
Besides, that's irrelevant. There's nothing stopping my function from following every pointer on the stack and smashing up its contents; are you going to defend against that, too? If not, how is this any different?
Instead what you'll do is specify the constrained inputs and expected output behaviour. From there you can out anything that violates those constraints as non-conformant. As long as you maintain those constraints between versions, there's no ABI breakage.
Also you can absolutely have constant references in an ABI. There may be ways of ignoring the const depending on how you design the ABI but they will be obvious abuse.
It didn't matter before, as compilers were not optimizing as much, code had a much closer 1:1 correspondence to assembly (if you are passing it by pointer and not register, you would want to make that clear in code).
It's much easier to implement in simple compilers. On the side of the callee you don't have to check if you manipulate your arguments, which is generally hard. Being able to manipulate your arguments is another shortcut for keeping the compiler simple. On the side of the caller you don't have to check if you hand out a mutable pointer.
Also finally and most importantly: memory access was much cheaper in terms of cpu cycles. Just look at cdecl: all parameters are passed on the stack instead of registers. Our current calling conventions stem from performance hacks like fastcall that were only optimizing for existing code (you pass big structs by pointer by convention).
(Post author.)
Mechanically, what happens is essentially the same as what ms/arm/riscv do: the caller creates a reference and passes it to the callee. The only difference is that the callee is more restricted than it would otherwise have been in what it can do with the memory pointed to by that reference. So I don't think that there can possibly be any implications for non-local control flow.
You can get around the immutability of the reference if your compiler implements the ABI with copy on write semantics, which I think is a reasonable compromise. But I'm still not certain how you would handle arbitrary control flow that the compiler may not be able to reason about.
If for example your arguments may be behind const references, how would you implement getcontext/swapcontext for your ABI? If everything is an integral value in registers or on the stack then it's really easy, but i would think it would have to be a compiler intrinsic if it depends on the function signature of the calling context, in order to perform the required copies.
Then if you care about the possibility of signal handlers modifying the original... you pretty much have to make a copy every time anyway.
Using rust and propagating the single writer xor multiple readers requirement in an ABI, this might be interesting. But with C/C++, I'm afraid copies would be forced "all" the time.
void print_foo(FILE *outf, struct foo foo) {
fprintf(outf, "foo '%s': %i, %i\n", foo.name, foo.x, foo.y);
}
That one would gain a speed-up and code-bloat reduction from the proposed ABI, and there are many like it.But even if every single function had to fall back to making a copy, the argument is that there's still a significant code bloat saving by putting the copy in the callee rather than in the caller. After all, the instructions necessary to make a copy takes some space, and with the proposed ABI, those instructions are put in the called function, rather than in every function call. Most functions are called more than once, and all functions are called at least once (hopefully), so anything which can be changed from O(number of function calls) to O(number of functions) is an improvement.
1. In many cases, no copy would have to be made. There are lots of small non-complex functions out there where the compiler can prove that it's safe to not make a copy.
2. In many other cases, a copy has to be made. But the copy is made by the callee, not by the caller. That means that all the instructions necessary to copy the argument ends up in the binary once in the callee, rather than once for every function call, leading to less code bloat (which has its own performance advantages).
In fact, a stupid compiler could just always make a copy without analyzing the function body. This would result in a compiler which generates code that's about as fast as it would be with current ABIs, but with a smaller size.
Could you provide an example of a situation where there would be more copies made using the proposed ABI than in traditional ABIs?
struct A { int m; } global = {5};
int f(struct A a) {
global.m = 7;
return a.m;
}
int main() {
f(global);
// need to make a copy of 'global' here
// otherwise f will return 5 instead of 7
}In the first version, the worst case situation is that only one copy is made, and it's always made by the caller. However, the caller has to make a copy if the object is referenced after any function is called, because that function might otherwise modify the parameter if a pointer to the caller's version of the object has leaked out somewhere.
In the second version, the worst case situation is that two copies are made where old ABIs would make just one copy (if the caller has to make a copy and the callee has to make a copy). However, the callee would only have to make a copy if it actually does something which might modify the object through the pointer passed as an argument, so the optimization would apply for more functions.
I think it's fairly clear from the article that your intended ABI is the first version, due to the sentence "In the event that a copy is needed, it will happen only once, in the callee, rather than needing to be repeated by every caller" . But in this comment, you're implying that the caller makes a copy if it can't guarantee that nothing else has a pointer to the object?
Your first interpretation is essentially what the ms/arm/riscv abis do. The reason I don't think that works as well is—
In general, it's rare for functions to mutate their parameters by value. We can effectively treat this as an edge case, and 'compensate' by making copies in the callee when necessary. But, when does the caller need to make a copy?
Version 1: whenever the object is aliased before the call, or read from after it
Version 2: whenever the object is aliased before the call
I think using the same struct multiple times is something that happens relatively frequently, so compared with v1, v2 elides a lot of caller-side copies. In exchange, it adds a relatively small number of callee-side copies. Which, despite the few pathological cases, seems likely to be overwhelmingly worth it most of the time.
Anyways, your intention is clear now at least. I'd be a bit worried about an ABI which might produce two copies for one parameter. It would be interesting to analyze a bunch of real-world code and see A) how often would my version create a copy, B) how often does the MS/ARM/RISC-5 version have to make a copy, C) how often would your version make a copy, and D) how often would your version require two copies.
Would also be interesting to see an analysis of code bloat due to copying parameters.
So the callee has to know what every caller of it will ever do? That's ... an ABI. The whole point is that functions can exist in a vacuum without knowledge of who they will be called by.
To be clear, I think it would be really cool if compilers could generate ad-hoc calling conventions using lto to optimize spillage, but that's not really useful as an ABI.
> would be interesting to analyze a bunch of real-world code and see A) how often would my version create a copy, B) how often does the MS/ARM/RISC-5 version have to make a copy, C) how often would your version make a copy, and D) how often would your version require two copies.
> Would also be interesting to see an analysis of code bloat due to copying parameters
I agree!
The ABI I had in mind was similar to the AArch64 ABI:
>If the argument type is a Composite Type that is larger than 16 bytes, then the argument is copied to memory allocated by the caller and the argument is replaced by a pointer to the copy.
But with a slight modification to put the copy in the callee:
>If the argument type is a Composite Type that is larger than 16 bytes, then argument is replaced by a pointer to the copy. The callee copies the pointed-to object into memory allocated by the callee.
This immediately has the advantage of less binary bloats, because the amount of parameter copying instructions in the binary will become O(number of functions) rather than O(number of function calls). (As an aside: That can probably be a huge advantage for C++ with its large, inlined copy constructors.)
When the copy is made in the callee, we can start identifying cases where a copy isn't necessary, or cases where only certain parts of the struct has to be copied. It would have to be fairly conservative though, since unlike with your ABI, there would be no guarantee made by the caller that there are no other references to the parameter.
I think my version is a clear and obvious improvement over the status quo, with decreased binary sizes and as good or better performance. Your version is more risky where the worst case is two copies per large parameter but, your version will probably achieve zero copies in way more cases than my version. "Low risk / medium reward" versus "medium risk / probably high reward".
---
Anyways, I might end up writing a blog post on this stuff. If I do, it will refer to your blog post. How should I refer to you? Moonchild or elronnd or something else?
So I think this idea is virtually impossible. Your ABI is probably not wrong, but your blog post probably is :P
The compiler would have to produce a copy in the case of this function:
void foo1(struct large l) {
bar();
baz(l.x + l.y);
}
Because the call to `bar()` might change the caller's copy, so the callee `foo1` has to make a local copy before calling `bar()`.But it would not have to produce a copy in this example:
void foo2(struct large l) {
baz(l.x + l.y);
}
Because nothing can change the caller's copy before the callee `foo2` is done using it.The compiler would not have to produce a copy in either case.
The call to bar() can not change the value of l because the person who called foo1 gave you an immutable pointer to it.
How did foo1's caller know it was immutable?
If the value passed to foo1 was a local variable, or one of it's arguments passed to it by immutable pointer, then all the compiler has to do is check if, in that function, it created a mutable pointer to the value and put it somewhere accessable to someone else. If it did not, then it can just pass the pointer right through.
If it did create a mutable pointer, or if the value came from some other place the compiler has no knowledge about, then it would make a copy of the value before passing a pointer to foo1, just like it would do for the old ABI.
When generating code for `foo1`, the compiler has to know how the parameters are laid out. The code generated for `foo1` can't depend on the code which calls `foo1`. It is my understanding that the proposed ABI says that `foo1` will get `l` passed with a pointer, and that the term "immutable pointer" just means that `foo1` isn't allowed to modify the object through the pointer.
The context (invisible to foo1) could look like this:
struct large *lptr;
void bar() {
lptr->x += 10;
}
void do_thing() {
struct large l = { ... };
lptr = &l;
foo1(l);
}
So the compiler must assume that the call to `bar()` can modify the object pointed to by the "immutable pointer".---
Or are you suggesting that the callee's decision on whether to make a copy or not should be a runtime decision, and that the caller passes in information about whether the callee has to copy or not?
Thefore, before the call to foo1, the compiler simply creates a copy of l, and then passes an immutable pointer to that to foo1. Effectively falling back to how the old ABI worked.
The code generated for foo1 does depend on the code that calls foo1. That is what an ABI is. An ABI is a specification of the contract between the caller and the callee.
The proposed new ABI says "If you are passing a struct by value, then you, the caller, will give me a pointer that you guarantee can't be modified until I return. In turn, I guarantee that I will not modify the pointer you give me."
My understanding of the ABI was: "The callee will always receive a pointer to the caller's copy. If the callee can't guarantee that nothing modifies the caller's copy, the callee makes a copy."
Your understanding is subtly different, because the caller guarantees that nothing else has a pointer to the parameter (and the caller guarantees that by making a copy when that's the only option).
To be honest, I think the ABI according to my understanding is better. It does reduce the amount of cases a copy can be avoided, but it means that the worst-case is that one copy is made (by the callee). With your version, you would have a bunch of situations where the caller has to make a copy (because a pointer has escaped) and the callee has to make a copy (because it modifies the object). With my version, the worst-case is that a copy is made once (by the callee), but the worst-case would be hit more often.
I think my understanding is closer to the intent of the article, due to this sentence: "In the event that a copy is needed, it will happen only once, in the callee, rather than needing to be repeated by every caller." In your version, the copy is sometimes made in the caller and sometimes in the callee.
> The callee also has more flexibility, and can copy only those portions of the structure that are actually modified.
It's clear that the author understands that the callee needs to copy any value it wants to modify.
But also, your understanding doesn't actually work, due to the issue you mentioned. So I think it's pretty unreasonable to assume that is what the author meant.
In my version the callee can absolutely make a copy of only the modified fields if the function A) modifies some fields but not others and B) doesn't do anything which demands a full copy.
Anyways, discussion on the intent of the article isn't super productive, I wrote a reply to https://news.ycombinator.com/item?id=27091726 which asks the author what their intention was.
Consider:
void other_thread(struct large* l) {
struct large new_l = { ... };
*l = new_l;
}
void do_other_thing() {
struct large l = { ... };
thread_t t = start_thread(other_thread,&l);
foo1(l);
wait_for_thread(t);
}
Not only can `foo1` never avoid making a copy (because it can't assume it's unknown caller is not this perverse), it can't even make a copy safely unless its caller allocates a new `struct large` for it to copy from. (And the caller does not, in general, know that `start_thread` (or `process_thingy_maybe_concurrently^W^W`) starts a thread with the aliased `&l`, so it would have to make such copies on any substantially escaping alias just to be sure, at which point you're back to the original proposal.)The pass-by-non-mutatably-aliased-reference ABI seems terrible IMO, but it at least allows the callee to decide whether to copy using only information that's actually present in the callee.
There is no such thing at the level that the ABI works, especially at a kernel-userland boundary (where the choice is to fully marshal the arguments or accept that you have a TOCTOU issue).
You ABI could say a pointer may only be accessed on Wednesdays if it really wanted to.
The problem I see with that proposal is that it introduces a new type of pointer, the immutable pointer. That seems like equal to a const pointer, but it's not. With a const pointer the callee can not mutate the pointee. But it can still be mutated from the outside. That means that any time you handed out a mutable pointer to something you have to make a copy for handing out an immutable pointer to it. An ABI like this would probably be much more complex to implement for a small gain.
You would end up with hardly predictable behaviour wether the struct gets copied or not. C# structs suffer a lot from this, because methods are mutating by default (https://codeblog.jonskeet.uk/2014/07/16/micro-optimization-t...). The biggest problem is simply that this is not explicit.
Also there is a case where a calling convention like this can make things worse, as you will have to make two copies:
void fn1(A*);
void fn2(A a) {
// Have to make copy here because mutating
fn1(&a);
}
void fn3()
{
A a;
fn1(&a);
// Have to make copy here because reference to a might have escaped.
fn2(a);
}That's incorrect. In c semantics, at least, it is legal to take a pointer to non-const, cast it to a pointer to const, pass it back, and mutate the object pointed to. It is only illegal to mutate if the pointer actually points to const memory.
Which means that, in a general sense, if you're given a const pointer, since you can't know if it actually points at const memory, you shouldn't mutate through it. But if you're handing out const pointers to non-const memory, you shouldn't count on the memory not being changed through those pointers.
IOW const is all but useless in c (except as a declaration of intent).
> An ABI like this would probably be much more complex to implement for a small gain
Why? Mechanically, it's almost the same as the ms/arm/riscv abi, except that a copy is made by the callee rather than the caller.
It is not about enabling compiler optimisations but about preserving invariants.
void mutate(const int *x) {
*(int*)x = 5;
}
int main() {
int a;
mutate(&a); // ok
const int b;
mutate(&b); // undefined behaviour
}So then actually no - it's not guaranteed to be portable. That non-const location qualification of course means portable code cannot just cast const pointers to non-const.
Let's be specific: compilers for languages that support aliasing. For example, FORTRAN does not permit aliasing and therefore has all sorts of optimizations that languages that do permit aliasing cannot have.
It's a tradeoff like any other, and isn't specific to a compiler beyond the fact that a given compiler X can compile language Y.
This is specially noticeable in OSes outside the UNIX culture, where two C compilers might not share the same ABI for all its features, while having the same ABI subset for OS libraries and syscalls.
See also https://en.wikipedia.org/wiki/Aliasing_(computing)#Conflicts...
Still, definitely an interesting optimization. I can definitely see a place to use this optimization myself in the near future (custom new ABI) either way.
Not that there's anything intrinsically wrong with having a differentish calling convention for C and C++ - it'd just add a smidgen more complexity is all.
Having said that I think the suggestion would be observably break both the current C and C++ standards, so something like this shouldn't be done without an explicit attribute.
struct X { int i; };
X* p;
void bar(X x){
P->i=1;
assert(x.i==0);// fails with transparent pass by ref
}
void foo() {
X x{0};
P=&x;
bar(x);
}This is something the compiler already keeps track of anyway in order to be able to know if it needs to re-load variables from memory when they are used again after a function call.
Basically the optimsiation degrades to exactly what happened before in the case where you do this.
In practice this optimization would be very fragile (like most optimizations relying on escape analysis) and people would keep passing by reference in fear that a small change like taking the address of an object might force a copy of a large struct.
a much better solution of course is to only apply it when strictly necessary, which seems not to be the case de facto in the wild
this is one very easy way to be 'faster than c' in practise
Could theoretically be for Rust, but it sounds like Rust doesn't have the aliasing optimizations working yet. Based off [1], it looks like Swift's ABI might be able to take advantage of these things. (I don't know much about Swift) Based off [2], it looks like Zig definitely can do these sorts of optimizations. Ada also has aliasing constraints, (you need to specify to the compiler that you want a thing to be aliasable) but I don't know if any compilers use them for optimizations. I know that Fortran can do these sorts of optimizations. Both Ada nor Fortran seem to be undergoing a PR renaissance, so they might be "trendy new languages" for that purpose.
However, given that the other two articles on the author's website [3] are about Delphi and complaining about AT&T assembly syntax, I'd suspect that the author is just a system's programmer annoyed about ABIs.
[1] https://gankra.github.io/blah/swift-abi/#ownership [2] https://ziglang.org/documentation/0.2.0/#Type-Based-Alias-An... [3] https://elronnd.net/writ/