The glibc s390 ABI break (2014)
lwn.net
lwn.net
The difference between the Linux and the OpenBSD mentality in one example.
The ABI must be changed.
Linux: That was really hard, we are never going to do that again.
OpenBSD: That was really hard, we better get good at it.
My opinion: When your deliverable is made up of source code(like these open source projects are), the ABI and ABI stability is not that important, it is the API(the source interface) that is critical.
MS: "Let's design it so either it never breaks or if we can't avoid that we'll just add a FooEx or Foo2 method call"
This is why so many Win32 methods take structs that have a size parameter that must be filled in. That acts as a version parameter. That's not to say they haven't had ABI fun moments... glares at COM.
I get this feeling every time I read a post on The Old New Thing (Raymond Chen). He'll patiently explain why some weird wart is the way it is, or how misusing some API is bad, and it's always extremely interesting. But at the same time I'm sitting there just thanking my lucky stars I've managed to avoid ever having to write software for that environment, because it's always absolutely bonkers.
I much prefer the EVENT based system Win32 goes with. You can just suspend on the event and not worry about it. The kernel will wake you up when it matters. Also MS seems to have very serious API design rules that help keep usage patterns consistent, which is important in making sure things aren't accidentally misused.
…And that’s without addressing any other parts of the system beyond the console.
I’d still take Linux over Windows every day if the week though.
Whatever version of windows you make your software work on it just works on all future versions.
If you don’t like the old APIs don’t use them.
Being able to put software on a website and have people download and use it immediately is kind of amazing compared to Linux.
COM is very annoying tho. I’ll give you that.
That said, I couldn't name you a Win32 that actually has more than one valid value for said cbSize member. Are there any?
One example: https://source.winehq.org/git/wine.git/blob/62df608d3ed84aac...
https://devblogs.microsoft.com/oldnewthing/20031212-00/?p=41...
um/webauthn.h has a bunch of structures that explicitly document fields added in various dwVersion s. However, no individual structure has changed since being published in the Windows SDK last I checked. e.g.:
//
// The following fields have been added in WEBAUTHN_AUTHENTICATOR_MAKE_CREDENTIAL_OPTIONS_VERSION_2
//
// Cancellation Id - Optional - See WebAuthNGetCancellationId
GUID *pCancellationId;
um/wincrypt.h also has a few structures chopped up with #ifdef ..._HAS_EXTRA_FIELDS s that are presumably differentaited between via cbSize.um/ShlObj_core.h has COMPONENT, which extends and is differentiated from IE4COMPONENT - presumably by dwSize.
It also wouldn't suprise me if cbSize is also used to differentiate between 32-bit and 64-bit versions of the same structure as well, with WoW64 blindly forwarding said structures.
I find it interesting upgrading the SDK could make your code invalid, due to uninitialized struct members. Compiler warnings can catch that sort of thing, thankfully—at least these days.
There are other subtle edge cases as well - like needing to make sure all code (including third party libs) is built with the same defines and settings if they might share said structs without painstaking checking cbSize & never passing by value. (Even that might not be enough when statically linking, as it's an ODR violation - and thus undefined behavior - to mix and match different definitions of the same struct...)
In the real world,
- people use operating systems to get work done,
- such work is often done with proprietary software that people buy, and that lot of people get paid to develop
- breaking such software is not acceptable, because it means that suddenly dozens of thousand of people can't work
- telling all those people to "stop working" until their software is fixed is not acceptable either.
People use Linux for work in the real world, and that's why Linux tries very hard not to break ABIs.
Microsoft has had a stable ABI for 30 years, and people still run software compiled 30 years ago on the latest Windows version today.
What this tells you about OpenBSD users, is another story.
Ultimately, it is about the APIs when you're compiling from source, a package that isn't using a new API is going to need the old library, and both might need to be compiled with and linked against the same toolchain. I think Linux and BSDs are closer than one might think here. Packaging and upstreams have both gotten better here over time I think, at least over the last couple of decades I've used Linux. I've only played around with the BSDs in VMs.
This is also the reason why containers are so popular these days, ship a software with all its dependencies to avoid having to recompile stuff each time you upgrade or change the operating system.
I think ISV's disagree.
It is absolutely important. ABI instability is the number one reason why packaging software on Linux is difficult. Breaking ABI causes pain even for maintainers of free software projects and packages. The Linux kernel is the only project in a Linux distribution that seems to take it seriously.
commit 2f438e20ab591641760e97458d5d1569942eced5
Author: Stefan Liebler <___.ibm.com>
Date: Thu Jul 31 20:04:54 2014 +0200
S/390: Revert the jmp_buf/ucontext_t ABI change.Both this move to glibc and the prior move from a.out to ELF taught me an incredible amount about Linux (and unixoid systems in general) because I had to rebuild pretty much everything in the system by hand, reading everything's READMEs and learning a bit more about them. (My system at the time originally came from a distribution installed by CD, but was then never upgraded the official way but only move forward with source tarballs, i.e. no package management. I got that luxury much later.)
Fun times.
libpng using setjmp.h for error handling was always my least favorite part of libpng, especially since the libpng documentation indicates it was just done for convenience (tell that to the authors of the OP!):
> The motivation behind using setjmp() and longjmp() is the C++ throw and catch exception handling methods. This makes the code much easier to write, as there is no need to check every return code of every function call.
It is true that changing the size of jmp_buf breaks binary compatibility, but that's hardly a unique fatal flaw in setjmp/longjmp; it's also true of fd_set, struct timeval, struct tm, struct sockaddr, struct stat, and all the other exposed memory-layout interfaces mentioned in the article: __pthread_unwind_buf_t, PerlInterpreter, png_struct_def, etc. Indeed, you'd think it would be much less of a problem for jmp_buf, because normally the things stored in a jmp_buf are precisely the callee-saved registers in your ABI; changing that involves comprehensively breaking your ABI anyway.
There is a big advantage in C to exposing the memory layout of a struct in this way: it doesn't need to be heap-allocated, so it doesn't introduce a dependency on heap allocation (which wasn't part of the standard library at all when setjmp was defined, and is still forbidden in many contexts), and you can statically bound your program's memory use, so you can be sure it won't fail. Heap allocation can always fail, so you can only ever use it in programs where failure is an option. You don't want your antilock braking system to raise an exception and reboot because its heap has become fragmented.
My brain hurts-- if libc.so.6.1 is a nightmare, then what is the utility of having the libc.so.6 soname numbering at all?
I'll put it in a more effective wrong-thing-on-the-internet style for receiving responses: There's no point in using NixOS. Just use lib.so.versionNumber. (Sorry) :)
[1] Or, perhaps more accurately, an incompatible version of libc requires the creation of an entire distinct target triple. You can have multiple incompatible versions if you've got a multiarch setup of some kind (like 32-bit and 64-bit x86 code on the same system), but an application can't simultaneously use both in the same process in such systems still.
Just open up some process in procexp and see how many msvcrt##.dll you find loaded.
It means every library has to be careful about providing both malloc and free calls in their API. Eg if a library has a `Foo *create_foo(void)` it has to provide a `void delete_foo(Foo*)`, instead of expecting the user to call `free(the_foo)` because that `free` might be from the wrong CRT. With Linux libraries I often find that you have a `create_foo` but no `delete_foo`, and you're expected to just `free` it.
I assume Linux libraries never took this much care historically because software was generally open-source, so everyone was getting their software from distro repos that rebuilt the world to link to a single libc.
The actual functionality that glibc broke in this specific ABI break is really the sort of functionality that is provided by kernel32.dll (i.e.., that's where the SEH functions live) on Windows. And you can't provide multiple copies of kernel32.dll on Windows, just like you can't have multiple copies of glibc.
Nobody is trying to run multiple incompatible versions of glibc in a single application. The point is to run multiple different versions of the same application on a single OS, where each version uses a different version of glibc.
NixOS is wildly over-engineered for this purpose. Just use lib.so.versionNumber, full stop.
And, scene. :)
on windows multiple incompatible versions of the C runtime in the same application are mostly fine as far as I know and fairly useful, one can load a DLL built in 2003 in a program built in 2022.
I guess Csound would be the grandaddy to check. But I have a feeling this would essentially be limited to running old proprietary plugins for which no source is available?
Yes, I have some songs from circa 2007-2009 which depend on windows 32-bit versions of freeware plug-ins whose developers have been MIA for 15 years now. Well, now I know better and try to only use software that I can recompile or write for my music (including pd :-)). But then I also spent 10 years not making much music because of that.
I also work with artists who keep old 10.6 Macs around in case they'd have to perform one of their past songs.
Because of the songs being locked away, or because of having to learn an entirely new software stack?
> I also work with artists who keep old 10.6 Macs around in case they'd have to perform one of their past songs.
Thanks to the various guild orgs from the early 2000s, many commercial filters from that time period were GPL time-bombed. Regardless, it was likely registered and is still held in the Trust's archive. Contact your local maker rep-- you may be able to make a commons release request and obtain the abandonware key from them.
Oops, wrong timeline...
glorp
because of writing my own software stack to do exactly what I want aha (https://ossia.io)
> Oops, wrong timeline...
you really had my hopes up for a few seconds
The issue is then when porting stuff from unix world which expects that you can pass things like FILE* between different modules or that you can return malloc()'d thing that can be free()'d by the caller.
Edit: another thing is that on windows there are CRT implementations that do not implement C++ exceptions and setjmp/longjmp in terms of SEH, which creates additional dimension of hard to debug ABI incompatibilities (and sidestepping this issue involves being sure that SEH unwind will not happen through SEH unaware code, which is essentially impossible to do unless you handle that as an fatal error).
IIRC I tried kylix in 1999. It still requiring libc5 is strange. Well... those were not best inprise/borland days indeed.
What libc implementation was that, then? I thought glibc was the first to support Linux.
Despite being version 2.0, they set the so version to 6 because the Linux specific fork had reached version 5. That is why libc6 is synonymous with glibc2.
Here is what I found on Google trying to confirm these memories: https://lwn.net/Articles/417848/
There were other examples from the gnu world where an experimental fork ends up with rapid development and then gets superseded by an upstream release. Egcs vs gcc is another I remember from that time period.