Thoroughly Understanding C++ ABI (2024)
ykiko.me
ykiko.me
Also C ABI also does not exist, people keep mistaking the ABI of their favourite C compiler with the OS ABI, which only overlap if the OS was written in C to start with.
For example on mainframes and micros, naturally not written in C, it is either a bytecode based ABI like TIMI on IBM i, or language environments like on z/OS, ClearPath MCP, OS 2200 and so forth.
And as a reminder, from a famous WG14 and WG21 member, and former Rust contributor,
"To Save C, We Must Save ABI"
https://thephd.dev/to-save-c-we-must-save-abi-fixing-c-funct...
All you need to do is clarify the first post on the thread with documentation.
Sometimes I do miss these Usenet like rounds.
I believe GNU follows the SysV ABI on amd64: https://refspecs.linuxfoundation.org/elf/x86_64-abi-0.95.pdf
That ABI is known for UNIX systems yes, although it isn't uniformly implemented, and again comes back to the whole OS written in C, like UNIX naturally is given it is the reason C exists in first place.
Which as mentioned, is not followed by every UNIX variant anyway.
Now GNU OS, still don't know what that is, Linux kernel with glibc?
Then again, using other libc doesn't require the same ABI.
COM said, "objects are cool and modules are cool so we need an ABI for objects" and specified a vtable. It isn't the only programming model which continued from the same observation: Python did too with its absolutely abysmal leaky PyObject ABI
On OS/2, Smalltalk enjoyed a role similar to what would be .NET and VB on Windows.
Sure you can use delegation with aggregation, and type libraries (nowadays .NET metadata), but still isn't as ergonomic.
Here is an earlier comment of mine with some details - https://news.ycombinator.com/item?id=49142987
(Though you cannot call just any method asynchronously; the interface's IDL needs specific annotations and the underlying object needs to implement ICallFactory.)
You can also use IRpcChannelBuffer3 to make the call asynchronously no matter what the IDL says. You'll have to write the proxy plumbing yourself, but it'll work, and you don't even need to rerun MIDL.
COM was one of the best realizations of "Component-based Software Engineering" and Brad Cox's "Software-IC" model. It was a binary standard and so components could be written in any language and yet be assured of perfect interoperability provided you followed all the rules/conventions. There was a bunch of boiler-plate for the framework itself but once you understood it, everything was smooth sailing in your language of choice.
One of the things i always advise people is not to focus only on the current way of doing things but to study older well-known libraries/frameworks/architectures/kernels/etc. to really understand "Software Engineering" from many perspectives. That is where insight comes from and real understanding happens.
For people interested in understanding COM, see the classic Essential COM by Don Box.
Regardless, WinRT is COM.
It only adds IInspectable as additional interface alongside IUnknown, and .NET metadata files instead of type libraries.
If you were to interview any "standard" windows programmer today i can almost guarantee they know nothing about COM even though WinRT itself is an evolution of COM and .NET also provides COM interop. Only the older senior programmers who have used it directly know of it.
Also no one uses WinUI, if that is what you are implying with WinRT, only Microsoft employees on Windows team forced to deal with that clusterfuck framework after Project Reunion pivoted into WinAppSDK.
What do you think Qt, VCL, FireMonkey, ImGui, wxWindgets, JUCE make use of?
Also you do not need COM to program in C/C++ directly on Windows API since they are all C interfaces anyway. COM is just an architecture you choose to use or not depending upon your needs.
Ever since managed languages/virtual machines became "standard" programming platforms most Windows programmers only write to these. Everything is at such a high-level now that only curious programmers delve deeper into the rabbit-hole.
It is like using ReactOS instead.
All Windows APIs since Vista are delivered via COM.
How much Windows development are you doing actually?
Not that much apparently.
Also someone has to surface Windows APIs to managed languages, they don't appear by magic.
Nope, This right here tells me you do not have much actual programming experience in Win32/Win64 apis nor of Windows Internals. The C-style windows apis are from kernel32/gdi32/user32/etc. user-mode dlls which call into intermediate ntdll/win32u dlls which then calls into kernel mode ntoskernel.exe/win32k.sys. With modern Windows there are another layer of abstractions with "Windows API Sets" (https://learn.microsoft.com/en-us/windows/win32/apiindex/win...) which decouple those user api from their actual implementation dlls.
COM is at user-mode and so interfaces to the first set of dlls only. Since the windows api breadth is vast not all of them are exposed via COM. WinRT uses/enhances classic COM but also calls Win32 api as needed. So Win32 api and WinRT api coexist with the latter providing the "modern" way to api access - https://en.wikipedia.org/wiki/Windows_Runtime But because WinRT is oriented towards secure sandboxed apps many low-level Win32 apis dealing with memory management, thread/process manipulations, system hooking etc. are limited/removed entirely from WinRT api.
For your edification see this detailed older article; Turning to the past to power Windows’ future: An in-depth look at WinRT - https://arstechnica.com/features/2012/10/windows-8-and-winrt...
Your claims must be backed up with references before asking others about their experience. I only see you name dropping and quoting historical data in all your comments which is not relevant to discussing/understanding anything.
That is for new APIs, not extensions of existing APIs.
IIRC, most non-COM methods you'll see for those are wrappers around COM calls. Back when WinRT was promoted as the next big thing, once you got outside of the small circle of WinRT propaganda there was explicit acknowledgement that all of the APIs are actually COM based enabling continued development of non-.NET compilers.
Of course underneath it all, especially once you get to userland-kernespace interactions you go back to combinations of IOCTL, memory mapped IO, some basic object calls and undocumented fun of OpenVMS-derived IPC, but a considerable chunk of that is not exposed or documented for non-blessed programmers.
WinRT is made up of component dlls and therefore exposes functional interface bundles. But the implementation of these components mostly took the path WinRT -> Win32 api -> ntdll (aka "Windows Native api" which is undocumented). But since MS wanted to push WinRT they allowed some paths to be WinRT -> ntdll. AFAIK this is what the mandate/waiver was for.
So you have the current situation that some Win32 apis have no analogues in WinRT and vice-versa. The DirectX path uses "Nano-COM" (a stripped down version of COM) which calls into user-mode runtime dlls which then takes another path to the kernel.
Lets not even get into how .NET projections/interfaces/wrappers do their job ;-)
The key point to know is that COM is only a packaging architecture/framework/middleware for structuring at a user-mode binary level.
It is all quite elaborate and complicated and so making a one-line statement like "everything is COM" is sheer cluelessness.
Fun fact: Microsoft developed COM based on how Zortech C++ virtual functions worked. (It predated Microsoft C++ by years.)
Pdf at https://www.google.com/goto?url=CAESawHuR6pN4SDa5tLkqrSuoaot... and reissue at https://www.researchgate.net/publication/2396782_Multiple_In...
I read this in one of the documentation books (still have it somewhere) which came with SCO Unix (don't recall the date). This book contained articles on C++ only and was an introduction to the language and compiler (since the language was very new). I also remember that it had an article on Template mechanism.
Next was Jan Gray from Microsoft's excellent 1994 article in MSDN titled C++: Under the Hood which explained class layouts/vtables/etc. It is still a excellent read today. Pdf at https://www.google.com/goto?url=CAESewHuR6pNy2XLp9fzgwckhWpg...
Finally of course you have Inside the C++ Object Model by Stanley Lippman explaining everything in detail.
In a tragic way, given the prevalence of COM in Windows, especially since Vista, someone at Microsoft should offer a few Delphi and C++ Builder licenses to the teams responsible for doing COM tooling.
Because MFC/OLE, ATL, WTL, WRL, C++/WinRT all have their sharp edges and could be so much better, if someone actually cared about productivity and framework ergonomics.
In other words, the ABI we rely on today isn't really part of the C+ standard—it's more like the Itnaium C++ ABI or the MSVC C++ ABI.
In the end, I think the ABI stays stable because of community conventions established by compiler vendors.
Afaik Carbon is at this point the only attempt at a successor language that's still going?
Anyone that can reach out to Rust, Go, Java, C#, whatever, should do that preferably.
There are also some efforts to tame existing C++ with profiles, and replacing UB with erroneous behaviour, and that's about it.
ABI is more of OS-level thing. Most systems these days follow the SysV ABI, which is largely defined by the hardware manufacturers via the processor-specific supplements (the x86-64 one is here: https://gitlab.com/x86-psABIs/x86-64-ABI). These largely delegate C++-specific conventions to the Itanium C++ ABI (which they likely directly reference), although the ARM ecosystem uses a somewhat different layout for the exception handling tables. Microsoft uses a different ABI for both the underlying C ABI and for the C++ compatibility layer built on top of the C ABI.
One of the issues that crops up is that vendors end up needing to add extensions to the ABI for various reasons, and these extensions tend to end up being incompatible, since they're added before they've had a time to be standardized. 16-bit floats is a particular historical bugbear, as is the C23 _BitInt stuff.
There is often a distinction made between "language ABI" and "library ABI" in documentation which is highly confusing.
The answer is: library API + compiler ABI = library ABI
GNU libstdc++ ABI Policy and Guidelines - https://gcc.gnu.org/onlinedocs/gcc-9.2.0/libstdc++/manual/ma...
With exception of Swift, D, and bytecode based languages, the ABI is left to the vendors.
For those that think ISO/IEC 9899:2024 PDF has anything related to ABI, the actual ABI used by C compilers, is the OS ABI, if the OS was written in C to start with, and naturally this overlaps quite nicely with UNIX like OSes, and Windows.
It isn't like that on other platforms that decided to either use other systems languages, or expose their OS APIs in a different way, e.g. mainframes, micros, Android, ChromeOS, WebOS,...
If you are on e.g. z/OS you would be using ILE, Integrated Language Environment, on ChromeOS JS/WASM, on Android either DEX or JNI,...
80% of Android APIs are exposed via Java, and even native ones have to use JNI.
Binder APIs very much do not follow C ABI other than the part where you deal with wrapper around ioctl()
Binder is just C++ with a thin IPC compiler, more or less similar to any IPC idl. Could be sunrpc, but Google wanted a c++ wrapper.
But in any case this is completely irrelevant, because there is no such thing as a "userspace driver" (unless you're speaking about hurd or plan9). Drivers are by definition kernel-space code (which on Linux is either built into the kernel, or is a loadable kernel module). Those binder wrappers are just a convenience layer, underneath it's all still Linux, /dev/video0, v4l2, and so on.
As for Binder being C++, no it isn't. Historically it has connection with C++ due to how it was used, but like other IPC it's not bound to C++, and arguably given how it is structured compared to some other options there's even less C++ specific stuff in it.
Also, IIRC the use of Binder predates acquisition of Danger (?) by Google
All smartphone drivers would have to be rewritten for a different kernel.
Same applies to C and UNIX/POSIX, OpenGL/Vulkan/OpenCL/NVN/LibGNM(X),...
Thus analysers and language subsetting tools, enforcing language style guides and safer coding practices.
It is complex/complicated/baroque but for the power, flexibility, real-world-usage that it provides it is worth the effort. The mistake that people make is to try and learn it all at the same time which overwhelms them.
Use a good book like Discovering Modern C++: An Intensive Course for Scientists, Engineers, and Programmers by Peter Gottschling and you should have no problem at all.
I often see on HN these sorts of comments whenever C++ is brought up and it really needs to stop. It is wrong, adds no value to the discussion and does a grave injustice to the programming community.
We should be encouraging programmers to learn C/C++ since just knowing those two languages enables one to program MCU/embedded/desktop/server machines across all layers of software from apps/scientific software/OS/system software/bare-metal in the real-world. One can imagine the job opportunities for a programmer with such a skillset as market needs change. With both Hardware and Software evolving so rapidly nowadays, C/C++ are the one constant interface language in the industry.
It is just another language and not rocket science.
I don’t see exceptions there. Apple defines the ABI of Apple’s Swift’s implementation, Walter Bright (or his team) defines that of D, Python defines its ABI, etc.
If you were to write a Swift/D/etc compiler, you’re free to define your own ABI. Disadvantage is that you would give up linking with code compiled by the other compiler (workarounds such as pragmas, C++’s extern "C", are possible)
You're free to do whatever you feel like on your implementation, including not being compliant with the official language standard.
The C++ standard did force gcc to break the ABI of std::string (copy on write was banned - for good reason). They then watched the gcc community work through 10 years of pain to make the transition. They are also well aware that python 3 broke compatibility with Python 2 - and again it resulted in 10 years of pain for the python community to deal with that. With this history there are a lot of experts strongly against any breaking change. It might happen anyway, but only with strong justification and likely an attempt at a migration of some form (what? there are a lot of examples of migration plans that don't work that they are aware of)
Again, it is not because of convention, it is because of painful experience from those who break it.
During the same time the MSVC compiler broke ABI multiple times and the world didn't end.
The fact that breaking ABI was such a shitshow for libstc++ is more related to the way C++/ABI/SOs is/are handled in Linux...
Microsoft never said it outright, but I’ve always assumed the major con that made them switch was that people didn’t buy new Visual Studio versions because upgrading VS required also upgrading all of your binary blob closed-source dependencies.
For VS 2015 they adopted a policy of having a stable ABI and have preserved ABI compatibility for over 10 years now [1].
[1] https://devblogs.microsoft.com/cppblog/binary-compatibility-...
GCC changing C++ ABI usually meant horrific time for every C++ application on Linux unless you use something like Nix so you don't chance loading incompatible binaries.
Program Obfuscation via ABI Debiasing (pdf) - https://www.google.com/goto?url=CAESYgHuR6pNsi8_4x_I2WBy7lmf...
PS: James Coplien in his excellent Advanced C++ Programming Styles and Idioms shows many techniques one of which is actually replacing vtable entries at runtime to mimic features from more dynamical runtime languages (eg. Smalltalk).
[0]: https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2008/n26...
A nice succinct post by Pavel Khaipov; Copy-on-Write in Modern C++: Why the Standard Dropped It - and Why Qt Still Uses It - https://www.linkedin.com/posts/khaipov_copy-on-write-in-mode...
A detailed post by Andrii Nikishaiev; Why Copy-On-Write (COW) removed in C++? - https://www.linkedin.com/pulse/why-copy-on-write-cow-removed...
The std document on Concurrency Modifications to Basic String - https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2008/n25...
Finally, C++11 tightened the specification for std::string viz. contiguous memory, some methods must have O(1) access time, when references/iterators/pointers can/cannot be invalidated etc. which doomed COW.
One thing i try to do is point people to lesser-known but very good and advanced books. I generally find that most HN book suggestions refer only to popular books most of which are just so-so while many great books remain unknown which is a shame. For example how many people know of Operating Systems in Depth by Thomas Doeppner? It has lots of helpful illustrations, compares concepts in both Linux/Windows implementations and is an all-round great book. I really understood the nuances of signal handling between kernel/user modes when i read it here.