C Portability Lessons from Weird Machines
begriffs.com
begriffs.com
It’s amazing that, by carefully writing portable ANSI C code and sticking to standard library functions, you can create a program that will compile and work without modification on almost any of these weird systems.
The big question, which I ask whenever someone harps on about portability, is does it even make sense? Is your average PC/smartphone application realistically ever going to be ported to a 6502 or a Unisys mainframe? Keep in mind that I/O in particular is going to differ significantly, and something like standard input might not even exist (think of a fixed-function device like an oven controller with only a few buttons and LEDs for I/O.) I don't think it's particularly amazing, because the "core" of C essentially represents all the functionality of a stored-program digital computer; so if you completely ignore things like I/O it's not hard to see how the same piece of code can express the same concepts on any of the latter type of computer.
It should also be noted that these "weird" environments are often not "strictly conforming" either, because it's either impossible or doesn't really make sense to. Besides the "omissions" they will also have "extensions" that help the programmer to more fully utilise the platform's capabilities.
For entire apps it doesn't always make sense but being able to port code fragments and small libraries is definitely nice. You can write a (non-optimized) portable memset or memcpy very easily (although one should be careful with how they're used if the arch has "weird" CHAR_BIT). Ditto for printf, a simple compression library, a function to strip whitespace and non printable characters from a string, code to handle the FAT filesystem etc... That's stuff that's useful potentially everywhere, from a 32core Xeon to an 8bit controller.
Look at kernels like Linux that are ported to a wide range of devices, there's still a significant amount of shared code.
These days it's also fairly common to have code that runs on both on amd64 on the desktop and ARM32/AARCH64 on smartphones and tablets. It's not as large a gap as going from some DSPs to general purpose CPUs but there are still enough pitfalls that being able to write portable C without too much difficulty is a good thing.
It's not that the ability to write code that is so portable isn't useful. For something like a library of algorithms (e.g. compression - think zlib), there's actual value to be derived from having a single ultra-portable implementation that can run everywhere. But does something like Evolution or LibreOffice really need to never assume that CHAR_BIT is not necessary 8, or that int might be less than 32, or that int64_t might not even be defined? I would say that of all the C and C++ code that's running on modern devices today, the vast majority could safely assume flat memory addressing, 8-bit chars, 32-bit ints, two's complement, IEEE floating point etc - and nobody would even notice. In fact, a lot of it likely already does assume some or all of that, just not explicitly.
It would be nice to have an ISO-standard superset of C that catalogs such assumptions. Basically, a "non-DSP, non-mainframe" version of C, that's portable across all modern non-exotic platforms, and provides definitions for as many things that are UB or implementation-defined in standard C as possible.
IME I've rarely, if ever, needed to make any of those assumptions, though early in my career I unfortunately did. Unless you're writing a kernel or a compiler there's rarely a good reason to assume a flat memory model. You would assume a flat memory model if you're manipulating objects in a way that grossly subverts the typing system; I say grossly because assuming a flat memory model usually isn't necessary even for most illegal type punning hacks.
Assuming 8-bit chars is useful, yes, but really only so you don't need to insert masks everywhere when doing bitwise operations on chars, such as when marshaling integer types (preferably without type punning). AFAIU an 8-bit ASCII string won't be packed 2 characters to a char on implementations using a 16-bit char, so pointer arithmetic and string manipulation will always look the same assuming the same string encoding. As with integers more generally, most of the time all that matters is that an integer type has at least N bits, whether or not you'll use them.
If you're depending on fixed-width types then usually there's a leak in your abstraction somewhere or you're doing bitwise manipulations where it's probably wise to use explicit masking, if only to make the code more clear, if not provide the ability to parameterize the value range. Typical exceptions would be cryptographic code, hashes, etc, particularly code that needs to perform rotations. But thankfully C has provided fixed-width types for nearly 20 years now, and almost all (if not all) compilers, even niche compilers, support these.
Assuming two's complement is only useful when you're depending on overflow characteristics. But signed overflow is undefined, anyhow (and for still-good reasons), and unsigned integer types are already guaranteed to be using two's complement representation (which requires emulation on some architectures, similar to how compilers transparently synthesize 64-bit long long integers on 32-bit platforms).
I've never done heavy floating point arithmetic so can't speak to the usefulness of assuming IEEE floating point for, e.g., managing error accumulation. But IME general purpose software likes to assume IEEE floating point for its indirect benefit, like how JavaScript provides the ability to perform accurate 32-bit scalar arithmetic by dint of providing a single 64-bit IEEE floating point integer type, or how VMs (including JavaScript JIT VMs) abuse floating point to implement tagged types; but arguably the reliance on 64-bit IEEE floating point has been on balance detrimental by, e.g., making the adoption of 64-bit scalar arithmetic in languages like JavaScript more difficult.
Also, it's noteworthy that WebAssembly makes use of some of the flexibility (or, rather, stricture) of the C standard. Appreciating the subtleties of C semantics helps make your code portable not only to archaic environments, but to bleeding-edge and future environments.
And yes, you can do that. The question is, why have all this complexity when in practice int32_t is there pretty much all the time - as you've said yourself - and when lots of existing code relies on it?
Same thing on other counts. Yes, you can work without all these assumptions. But why waste time on making the effort, when it simply doesn't matter?
For C2x and C++2x it looks likely that they will drop support for sign-magnitude and ones' complement representations of signed integers, essentially mandating two's complement.
https://gustedt.wordpress.com/2018/11/12/c2x/
https://herbsutter.com/2018/11/13/trip-report-fall-iso-c-sta...
http://www.open-std.org/jtc1/sc22/wg21/docs/papers/2018/p090...
No and why should they? I assume all these things routinely when I write code that I don't expect to run on "weird" systems. And I don't have to worry about int being less than 32 since I when I need this guarantee use "int32_t" like a civilized human being.
The overhead is only here if you care to really write entirely portable C code but in practice very few people bother to do that outside of comp.lang.c.
>It would be nice to have an ISO-standard superset of C that catalogs such assumptions. Basically, a "non-DSP, non-mainframe" version of C, that's portable across all modern non-exotic platforms, and provides definitions for as many things that are UB or implementation-defined in standard C as possible.
POSIX C then? It assumes that chars are 8bits, that data and function pointers are castable and a bunch of other assumptions that are not possible in strict C. It also offers a richer API than C's stdlib. When I write C for a "non-DSP, non-mainframe" environment (that is, 99.9% of the time) I almost always target POSIX C, not strict ANSI.
And yes, POSIX C is a very good example (and I completely forgot about the function/data pointer thing... that one is probably one of the most common assumptions people rely in, and I bet most of them don't even know that it's not standard ISO C!). It's a great starting point, but POSIX being as old as it is, it doesn't include some things - e.g. it doesn't guarantee any particular size for int above and beyond ISO. At the same time, the richer POSIX API is not actually available on some platforms - notably, Win32 - so taken as a whole, it's not really a pragmatic portable superset, if you want to target all mainstream platforms.
What I was thinking about is more along the lines of, look at existing platforms that "matter", find the set of common assumptions that ISO C doesn't guarantee, but which all the platforms nevertheless provide in practice, and give an official blessing to that set.
I feel like if you start going down that road, towards “commonly portable”, you naturally end up with one of the not-c variants (zig, nim, etc). ...and ifc, if you start writing a new language, you might as well toss in features from the last thirty years of language design.
C is C largely because its so stupidly (and impressively) portable; if you aren’t drawing that benefit, its likely a case of the wrong tool for the job.
I'm working in medium sized embedded systems (not embedded Linux, but also not the smallest embedded systems). I've only rarerly worked on 8 or even 16bit Microcontrollers. Most of what I work on is 8bit chars, 32bit, Little Endian normal Microcontrollers.
I really don't need the portability, the targets are all the same, however they only have C/C++ compilers. And we can't use GC. Rust would be a very nice choice at least for the application code. I'm not sure how to write low level drivers in Rust, but with FFI I would just push that to the boundaries, and make at least 90% of the applications Rust. But there's no LLVM backend for most embedded processors I work with, and even if there is, the maturity might not be there, and no company would switch to that considering that the next target might not have a backend.
We have to use C, because it's the only thing that we are offered. But for desktop applications I see no real reason to start something in C nowadays.
https://www.mikroe.com/compilers
But as you say, the company needs to give the option to developers.
But you can already make those assumptions, and safely ensure that they are satisfied by using compile-time asserts (i.e. static_assert), which are very powerful in C++.
Need to know that int is 32 bit? Not a problem. Need to know that long is strictly larger than int? Not a problem. And if you do break someone's build with these static_assert's, that's fine -- it wasn't intended to be a supported platform anyway.
The other problem with this approach is that it effectively defines lots of distinct subtly different pseudo-platforms, each a set of assumptions. And your average developer might not know which of those assumptions actually are "portable enough", and which are not. The point of defining such a profile would be for experts to provide a blessed set of assumptions that is, in fact, portable across all general-purpose platforms that are in active use.
For example, the fact that CHAR_BIT is 8, or that int32_t is defined, or that ints are two's complement, are all reasonable assumptions. But two's complement overflow for signed arithmetic is not, even today - even though in practice it works just fine on some platforms (more specifically, on some implementations).
I'll continue under the assumption that C is capable of equally powerful compile-time asserts.
> it effectively defines lots of distinct subtly different pseudo-platforms, each a set of assumptions
Indeed, but using this approach, all that's needed to define a new 'ordinary platform C' is a clearly defined set of requirements. No need for a whole new standard in the usual sense.
> experts to provide a blessed set of assumptions
Sure, and that's compatible with this approach. It could take the form of a header file.
Good question, I'm not certain that's possible. A quick google turned up surprisingly little.
The ghastly hack alternative would be to detect which compiler/target we're dealing with and compare against a whitelist.
> I think this sort of thing would need some official blessing to be widely adopted, regardless of whether it's a written spec or a header.
It would certainly help to have it properly branded and given real credibility, yes. I'm reminded of MISRA C, which has a good deal of tooling and documentation.
I have 15 years of programming on 8but micros where that assumption/declaration is crazy. :)
Never use int. Always uint32_t. There! Problem of portability solved.
Right, but that's the point here - we're discussing how to describe a stricter variant of C, which make exactly the kinds of guarantees that preclude exotic architecture targets. Ultra-portability isn't a goal for many codebases.
By 'standard' do you just mean more widespread use? Some APIs already do something like this, such as Windows with its 'DWORD'.
* It's big-endian
* unaligned reads will fail (if you have to read packed structures, you need to declare the structure as packed and the compiler will generate code to do unaligned reads)
Given you mention "memory layout" it seems you are expecting possibly a little-endian system, and maybe unaligned reads.
Even if you stick to recent platforms, depending totally on FILE* can be a mistake. For example Windows has all its idiosyncrasies with sharing modes and whatnot, not all of them captured by fopen. Or you may wish to implement streaming or some alternate data source not on the filesystem. The author of a library can't predict these things, so should allow the caller to bring their own.
Just one example, because I'm writing one right now. Protocol encoding/decoding. I just write it once, can share headers with struct definitions, and can compile it for both ends, be it x86_64 on one end and PIC on the other.
A related question when talking about portability to odd or historic machines, exactly how is this supposed to be my problem and not the problem of the guy that decided to go that route?
Yes. I probably won't port my program to a PDP-11 or a Motorola 68000. But by relying on undefined behavior or implementation-defined behavior, I am making it hard to port my program, say, from Windows/x86-64 to Debian/ARM (e.g. Raspian).
Unaligned accesses don't trap with a SIGBUS or anything, they just round the pointer value down to an aligned address and read whatever that is.
Reading from and writing to NULL will generally succeed (just as on SGI).
Function pointers are just indices into tables of functions, one table per "function type" (number of arguments, whether the arguments are int or float, whether it returns an argument or not). Thus, two different function pointers may have the same bit pattern.
Is this likely/guaranteed to happen by the wasm standard itself, or, could runtimes/emcc fill in stub table entries to avoid this?
CHAR_BIT == 32
sizeof(char) == 1
sizeof(int) == 1
sizeof(short) == 1
sizeof(float) == 1
sizeof(double) == 2
I suppose the alternative would be to use an addressing scheme encoding bit offset in the pointer, like some of the other machines in this story. But that's also much more expensive, and this is a DSP architecture, so they went with something more straightforward. Curiously, this set-up is still fully accommodated by ISO C standard.This allegedly means that the standard doesn't require fgetc() (which returns a char converted to an int on success, or EOF on failure) to have sensible behavior on such platforms [1].
That said, I'm at a loss to explain why i = i++; is undefined and not merely unspecified.
you can compile in a way that unsigned overflow traps with gcc / clang's -fsanitize=undefined.
It has solved countless bugs of mine
Given the standard, it's obvious why i=i++ is undefined: it modifies i twice between sequence points. The question is why it was "undefined" (in which the implementation is allowed to make demons fly out your nose) and not something like, say, the "unpredictable" that occurs in many architecture specs (in which the program behaves as if i had been set to something in particular, but with no requirement on which value it is: unchanged is ok, incremented is ok, 12345 is ok if i is wide enough to hold it, but silently failing to generate any code for statements that occur after the i=i++ isn't).
Maybe it's UB only because implementation-specified is too onerous and it didn't seem to be worth defining a bounded "unpredictable result" behavior.
Later, the alpha was FreeBSD's first 64-bit platform, and when working on the port to alpha, we hit a lot of the same issues in the FreeBSD kernel. As alpha was also FreeBSD's first RISC machine with strict alignment constraints, we also hit a lot of unaligned access issues in the kernel.
Ah, those were the days. I now find myself grumbling about having to check to ensure my code is portable to 32-bit platforms.
Meanwhile, ARM does care about alignment, and its now the most popular architecture for anything that's "not a PC".
My first experience with this was writing some smartphone code that died with a SIGBUS when trying to make a function call, where the reason was totally non-obvious from simply looking at the code.
A couple decades ago, sure, ARM was different. Had it stayed that way, ARM would not be so popular today.
Almost zero software out there actually needs unaligned reads.
Most of the troubles related to -fstrict-aliasing involve unaligned reads. All sorts of file formats, TIFF for example, are most easily handled with unaligned reads.
A typical issue would have code like this:
foo = * (bar * )baz; // baz is a char pointer into a binary blob, and bar is a type that needs alignment
If the data were all properly aligned, then most likely there would be no desire to write such code. The correct types would be used.
Aside from gcc abusively optimizing, the above works fine on x86, PowerPC, and modern ARM. It does not work on older ARM.
That sort of code is everywhere.
If the instruction was something like "ADD @R1, @R2, 5", it would fetch the word, register, or bitfield pointed at by R2, add immediate 5, then save it to the word, register, or bitfield pointed to by R1.
The machine didn't have shift/rotate instructions, but it could be effected by saving to a bitfield then reading from a bitfield offset by n bits.
They had a working (but not polished) C compiler but that project got shut down when they realized the system was not going to take off.
I never thought of it that way, but that's true. However, he didn't mention the biggest issue with C on the 6502, i.e., the extremely constrained 256-byte hardware stack. To do anything practical requires some sort of software-maintained stack to have stack frames of any decent size or quantity (in parallel, or replacing the use of the hardware stack completely). "Downright hostile to C compilers," indeed.
I'm going to take a look at yet another alternative 6502 language called C02 now: https://github.com/RevCurtisP/C02
That sounds like any SGI, or Sun, and although they're mostly gone there's still the Power series from IBM (runs AIX), and the only reason to use the expression ".. unlike most computers nowadays" is by counting the sheer number of Intel, AMD and ARM chips in use. Of course those numbers are overwhelming - ARM alone sells billions - but it's not like big endian is some obscure old concept in a dusty corner. (The irony is that ARM can be used in both BE and LE modes, by setup). Anyway, at work I have to write all the code so that it runs on BE as well as LE architectures. BE is alive and well.
According to wikipedia, there's also IBM's z/Architecture and the AVR32 µc. PPC looks to be switchable like ARM.
> The irony is that ARM can be used in both BE and LE modes, by setup
I've always wondered how common it is for ARM CPUs to run in BE mode. Does anyone have info?
> BE is alive and well.
If only because network protocols are generally BE.
Fun fact: the optimized string library for the original ARM1176 raspberry pi had a bonkers implementation of memcmp() which used the SETEND insn to temporarily switch to bigendian, because on that core setend is only 1 cycle and it saved an insn or two later. (On newer cores setend is a lot more expensive and the trick doesn't work.)
https://en.wikipedia.org/wiki/Endianness#Bi-endianness lists ARM >= v3, PowerPC, Alpha, SPARC V9, MIPS, PA-RISC, SuperH SH-4 and IA-64 as bi-endian (can be configured as wither). MIPS is still used a lot in routers. I don't know if the MIPS-compatible PIC32 is also bi-endian - there's maybe no reason it should be.
About ARM, I don't think I have run into any BE-mode ARM systems in a while, but I do recall using big-endian gcc versions for something in the past. And apparently some folks are running Cubieboards in big endian mode (if that's common or not for Cubieboards/Cubietrucks I don't know).
Yeah, network protocols will keep BE alive if nothing else - and yet, the Power 7 systems I'm coding for at the moment are big endian and IBM won't change that in the reasonably long term.
>If only because network protocols are generally BE.
I think a big part of this is that, until x86 conquered all, big-endian was far more common than little-endian. So when all those "foundational" protocols were originally designed, it was little-endian that would have been the bizarre choice.
But both POWER and ARM have unsigned chars by default :)
The lesson of weird machines is that weird machines can lurk anywhere; that's why they're 'weird'. Something to do with CPP or bit endianness, perhaps, or exploiting undefined behavior is what I thought clicking on it.
And the reference in the manual to the Unisys machine not having byte pointers: the PDP-6 (the original 36-bit machine as far as I know) had byte pointers that allowed you to specify bytes in the width 1 to 63 bits wide. It was common to have six-bit characters (you could pack six to a word) as well as 7-bit ASCII characters (you could pack 5 to a word).
Then misaligned access will trap on x86. Can be useful for emulating architectures that don't support misaligned access.
Http://pdp10.nocrew.org/docs/instruction-set/Byte.html
It is hilarious to think about the possibility of having to make your code portable to a 9 bit big endian system too.
Here is the issue though, portability is a joke even between chips of the same series let alone chips of different mfgs and code for PC.
In a PC I’m not stuffing every possible bit in a structure because I’m going to push this “object” over a 10kbps bus at some point or because on system still measures it’s ram with a K not a G.
And typically it’s moot. Even if I have an RTOS, the majority of what a micro is doing is loading data in and out of peripherals. When it’s CPU heavy it’s typically (for me anyhow) just crunching some data in order to load a result into a peripheral. Well - those peripherals are never the same between mfgs. So portability it’s really as important to me as understanding how the mfg intended the peripheral to be used and doing so efficiency.
I’ve been doing this for some time and while it’s a nice goal to write once use many, reality gets in the way.
Before the great serialization options we have now (MessageBuffers, ProtoBuf, etc) there was a bit more ambiguity with what your data stream was. TLV (type-length-value) packing being pretty common... but I guess I’ve written plenty of domain specific ‘protocols’.
If you mean a constant bitstream of never ending data, if you can use an established format like I2C or SPI the hardware on both sides really takes care of most the gritty stuff giving you a nice interrupt on both sides. Doesn’t really matter one side has the “preference” to look at in in 32bit chunks and the other in 8bit. Besides, even on ARM the physical transfer is typically in 8bit units anyhow. SPI can send 16s or 32s, but can also stop at 8s usually. It’s the ASIC/drivers/devices that are more rigid in their streaming requirements (this device MUST accept 8bit transfers, etc).
I think mips r3k has a better behaving add ("addu"), or is that only on some? if your compiler outputs these you don't have to worry about special behavior.
I'd say a bigger concern, for vax - the lack of IEEE754 is noticeable when people pick unsuitable float constants, or it traps by default. or the many GCC bugs now.
For mips r3k, the complete lack of atomics. And the load delays.
And load delays are annoying when writing asm manually, but not for the C compiler.
And yeah, addu was in MIPS-I.
Course for 98% of use cases IEEE754 is broken.
http://bitsavers.informatik.uni-stuttgart.de/pdf/dec/pdp8/pd...
I started with this design for my relay computer, but evolved it to be even simpler: no accumulator.
(There was a C compiler on it, but it cheated and used 9-bit chars).
Worth to read it.