> An access trough a pointer is less efficient than a direct access
But the alternative is not "a direct access." It's a copy and then direct access. A copy, mind you, that occurs through a pointer to grab the original data from the stack position it was sitting in. That's the very same indirect read that the callee does in the proposed ABI!
Thus, the indirect memory read will happen either way; either at caller copy time (existing ABIs), or at read time (proposed ABI, copy optimized away), or at callee copy time (proposed ABI, copy not optimized away.)
The only difference in the number of indirect memory accesses, with the proposed ABI, is that you might be able to avoid an indirect memory write (i.e. of the memory the caller is writing to by doing the copy.)
> Finally, who guarantees that the programmer doesn't abuse this thing and starts to modify the data without doing the copy on write?
Uhh... we're not talking about features exposed to the developer. We're talking about calling conventions—instructions to the compiler on how to distribute the glue code it generates between the call-sites it builds ("caller-[verb]ed" things) and function prologues/epilogues it builds ("callee-[verb]ed" things.)
The "copy on write" part here [if you want to call it that—it's not "on" write] would be part of such a generated function prologue ("callee-copied")—the same code that gets generated at the call-site in existing calling conventions ("caller-copied"), just moved to the other side of the linkage. With the important difference that, if the compiler could prove that the callee (or anything it calls, passing the variable, transitively) wasn't going to do anything write-like with the data behind the pointer, then the copy in the function prologue could be optimized out.
This compiler optimization can only occur if the copy is the callee's responsibility. Given external linkage (think: separate compilation units, e.g. calling a function defined in a static library that's already compiled and only available in binary form), the caller can't guarantee the properties of the callee. But a callee can always guarantee its own properties during function-prologue/epilogue generation—since it's generating the function prologue/epilogue in the very same compiler session that has access to the AST that describes the function's static-analysis-proven properties in full.
(Yes, this means that a compiler could be mistaken about whether a callee writes through the implicit pointer underneath the passed value, and therefore generate the optimized function prologue when it shouldn't. This is known as a compiler bug, and wouldn't be something J Random Hacker would ever run into if using a stable, non-alpha release of their compiler, with a stable, non-alpha release of support for this calling convention.)
> What cleanliness?
The functionality being proposed is that you just pass-by-value (which makes obvious that the semantics your code is enacting is "the function can do whatever it likes with the value, and it won't affect the caller, because if it writes, it's writing to a copy of the value"); but sometimes it'll be automagically faster than it could have otherwise been, because sometimes, when the compiler can prove certain properties, no copy — only an indirect memory read — would be occurring.
This allows you to always use pass-by-value when you want those semantics, and to get code that's "as optimal as it can be"; rather than having to pass pointers around and explicitly copy-on-write and all that mess in order to get that same level of optimization.
(And yes, that's the alternative in heavily performance-oriented code like a database's tuple materialization logic. You don't get to write clean code at the expense of performance. Your "options" are to write performant code using dirty hacks; or fix the conventions you're coding on top of so that performant code doesn't require dirty hacks.)