Smaller size is good for cache, but other factors matter
From the above
Single-Threaded Environment
When a single thread on a single core is accessing the data in a struct, we can improve caching performance by using as little cache lines as possible.
By optimizing for memory footprint, as discussed in the previous section, the struct uses less memory and hence occupies less space in the cache.
By placing heavily used members close together, we hope (based on the locality of reference principle) that they will end up closer together in the cache, preferably even on the same cache line, and hence use less cache space.
By separating hot fields form cold ones, we reduce the amount of cache lines filled with unused data.
So for HPC code, you very much do want the ability to re-order them manually.
C, C++, Pascal all define the order explicitly, the platform ABI generally defines the padding rules.
If you include toy languages, you might get there.
Even languages like Go and Rust recognize that at API boundaries you need a stable and defined ABI.
These languages still allow you to consume and provide C compatible ABIs explicitly but this does not interfere with data optimizations for native data.
Go and rust can’t be used for system libraries: you have to create C interface. This means that if you have two libraries, both written in rust then they have to communicate through a C layer.
The alternative (what rust does) is to have every application contain a complete copy of every library it uses, which is horrific for performance.
Yep, not shipping (native) compiled code is usually how you end up with no ABIs. But these these languages and runtimes still don't really support optimizations of data structures very well, I think it's largely because they weren't specified and implementerd to do it from the start and now there are all kinds of ingrained things about the semantics and estsabilished implementations and user expectations that get in the way of doing big things like feedback based rewriting of data layouts.
Language specs rarely support it well, so much of the blame is on language designers. But it's a somewhat chicken and egg situation. Language users don't demand it either because they never had it in other languages.
There is more to it than lazy language designers.
I submit it's mostly solvable in language design. They're not all lazy of course, but claiming to be very performance-oriented is a half truth as long as you don't have a strong story here.
Neither is an issue in Go, since everything’s statically compiled and it doesn’t make any ABI guarantee.
Making this opt-out (or making precise layout opt-in) actually improves the situation there, because then you have clear, explicit guarantees.
As for the documentation,
Check -dynlink and -shared on "go compile"
https://pkg.go.dev/cmd/compile@go1.17.6
Also -buildmode and -shared on "go link"
https://pkg.go.dev/cmd/link#hdr-Command_Line
You can either create a dynamic linked package, that will dynamically link with other Go compiled code (from same toolchain), or expose a C ABI from a Go compiled .so (which may or may not include the runtime as well).
As for one possible example,
https://www.ardanlabs.com/blog/2020/07/extending-python-with...