I'd expect scalar replacement of aggregates (SROA) would eliminate the use of a struct like this in some cases, but there are definitely limits, especially if not compiling with LTO enabled, since optimizations across compilation units will be limited. Honestly doesn't seem worth the cost.
Using registers to pass structs is an ABI thing and doesn't need LTO, it happens during compilation (see godbolt.org links on the rest of this thread).
Only if the function is inlined, at least it seems that way from experimenting on godbolt.org.
I guess that's because the struct size is below 16 bytes (see my sister comment). Adding two more ints to the struct places the entire content on the stack.
Sure, so it doesn't have to do with inlining.
Note that this is not an optimization (as in something the xompiler can do when optimizing). This behaviour is mandated by the ABI.