The x86 architecture is the weirdo, part 2
devblogs.microsoft.com
devblogs.microsoft.com
+ HPPA's upwards-growing stack
+ SPARC register windows
+ PowerPC's slightly oddball varargs ABI
+ CHERI capabilities are massively spending weirdness budget
+ these days, being big-endian is spending weirdness budget
Doing a few odd things is fine, but if you blow your weirdness budget entirely then your new architecture is in danger of failing because you've cumulatively made adoption of it too expensive for the existing software ecosystem. So you need to be careful that you're spending it on things that really matter; if something's not a big deal, do it the same way everybody else does. (Certainly some of the things in the list above are deliberate choices for things that are, or at least seemed at the time, important. But nobody tried to do all of them at once!)
Of course if you're a pre-existing dominant architecture like x86 then you get a lot more freedom to do weird things :-)
edit: oh, and softfloat vs softfp vs hardfp vs vfp
edit: oh, and how they have two incompatible assembly language dialects that are mostly the same, but in non-trivial code, incompatible
Heh, having dealt with x86 for years, this is comparatively such a nothing burger. It's always a simple known fixed offset.
> softfloat vs softfp vs hardfp vs vfp
That's something not really unique to ARM per se. Any architecture with options for hardware FPU are going to practically need ABI specs for the soft and hard cases (you don't absolutely need anything more than hardfp since you can always emulate but it will be slow as shit) - and ARM is certainly not unique in having multiple hardware floating point implementations either.
> two incompatible assembly language dialects
Curious what you're referring to here - but I personally wouldn't consider assembly language dialects to be part of a CPU architecture.
I assume they're talking about ARM vs. Thumb?
The poor ELF specification ends up quite tortured by this, IMO.
Arm64 is even worse! There is almost exactly the same instruction, but it also zeroes the low bits of the target address, so as you relocate code you also have to change the offset in the 2nd instruction even if the distance between the reference and the target stays the same.
RISC-V was originally going to implement a hypervisor mode which would only have worked with Xen-like hypervisors. Luckily we were able to head that off early and the actual hypervisor extension we got can run KVM efficiently.
The base is too base and their bitfield extension is weird.
I'm sure those that only need the base disagree. For the rest of us, there's G (IMAFD).
ARM endianness is switchable at runtime.
They are? Care to point to a ratified RISC-V standard that lists said instruction sequences?
slli rd, rs1, {1,2,3}
add rd, rd, rs2 Fused into a load effective address
...this is so insane. Whoever thought that this is okay and good, has, in my opinion, severe psychological and psychiatric problems and would do well to seek professional help. If this gets "fused" into a lea, why just not implement a hex code for lea? I'm just completely at a loss as to how messed up that is.
You know what, I'd like to know what a person who thinks that this is okay looks and behaves like.
Thus they are replacing not two instructions but three. This addition was indicated because of the number of critical loops in legacy software that, against C recommendations, tries to "optimise" code by using "unsigned" for loop counters and array indexes instead of the natural types int, long, size_t, ptrdiff_t. This can indeed be an optimisation on amd64 and arm64, but it is a pessimisation on RISC-V, MIPS, Alpha, PowerPC.
One codebase that uses "unsigned" in this way is CoreMark and they explicitly prohibit fixing the variable type. But it's also common in SPEC and in much code optimised for x86 and ARM in general, where using "int" pessimises the code. If they used long, unsigned long, or the size_t or ptrdiff_t typedefs the code would run well everywhere.
While the .uw instructions were being added, it was very low cost to add the versions using all the bits of rs1 at the same time.
So, in the context of this discussion, having 32 bit operations sign extend the results to 64 bits is a weirdness. More ISAs do is than zero-extension, but the most common in the market zero-extend. Note that at the time RISC-V was designed arm64 was not yet announced, so only amd64 did zero-extension.
An unsourced table on Wikichip (which was added in 2019 and not updated since) barely counts as a proposal (in terms of it standing a good chance of becoming part of a ratified RISC-V standard).
No doubt many high performance implementations will choose to use fusion and no doubt they'll all go for different combinations, with different edges cases. Yes there likely will be significant overlap but it could become a bit of a nightmare for a compiler writers. A thorough standardized list of instruction pairs to fuse would definitely help here, but we don't have one.
That's not in the spec. It's just something high end implementations do, and low end ones don't.
So if fusion is a weirdness it's a nearly universal one.
The good thing about fusion is the program works fine if you don't do it, so low end minimal area CPUs such as microcontrollers can just not bother.
+ 'char' being unsigned
+ handling of unaligned accesses (in early architecture versions a value is read from the aligned address and rotated, which is useless behaviour that falls out of the original implementation because of how it dealt with byte loads; subsequently it was at least made to fault, but it wasn't until I think v6 that unaligned accesses were made to Just Work)
+ the weak memory model
[0] See "Expanded Instruction-Length Encoding" in the user spec.
No one has done it, no one seems to be keen to be the first to do it, and even how the instruction length encodings work is not a ratified part of the spec -- it's just a proposal at the moment, even for the next step of 48 bit instructions.
There has been discussion of encodings better than the one proposed in the current spec, especially around instruction length encoding schemes that would make more opcode bits available in 80 bit instructions than in the scheme in the spec, so as to have a possibility of encoding 64 bit literals in an 80 bit instruction.
I've looked at wide RISC-V decoder design and the variable length is no problem at all out to at least decoding 32 bytes of code per cycle i.e. eight 32 bit opcodes or sixteen 16 bit opcodes, or somewhere between for a mix (average would usually be about 11-12 or so).
You just need 8 decoders that can decode any instruction, plus 8 decoders that only have to understand C instructions. The 16/32 decoders each need a 2:1 mux in front of them selecting either bytes 0..3 or 2..5 from a six byte window. They always output a real instruction. The C-only decoders will sometimes be told just to output a NOP instead [1]. Each decoder type needs a 1 bit input to tell it which option to take. Those inputs can be chained like carries in a simple adder, or they can be calculated in parallel like in a carry-lookahead adder. For an 8-16 wide decode you need a 2-deep network of LUT6 to do this (in FPGA terms .. also not very deep in SoC terms).
Note that this is a VERY wide machine. Possibly well beyond the point of usefulness given typical basic block lengths and what you can sensibly do in the OoO back end. x86 is currently doing 3-4 wide decode, and Apple M1 is doing 8 wide.
In short: no, it's not a problem.
[1] or not output an instruction at all. Outputting a NOP makes it easier to insert the decoded instructions into an output buffer. Then you need to filter out NOPs later -- which is needed anyway, as programs contain explicit NOPs, OoO machinery turns register MOVE instructions into NOPs by just updating the rename tables, etc.
- Every instruction being conditional (all instructions have a four-bit condition field, with one of the 16 possible modes being "always");
- The barrel shifter, which can be used on nearly every data processing instruction;
- The program counter being one of the general-purpose registers (and on the original ARM, the same register also containing the flags), so that any register move can alter the program flow;
- The load-multiple/store-multiple instructions, which can load or store up to 16 registers, plus incrementing or decrementing the base register; and since the program counter is one of these registers, it can restore several registers from the stack, update the stack pointer, change the program counter, and switch to Thumb mode (stored on the least significant bit of the program counter), all in a single instruction.
Being able to manipulate the program counter in the ARM processors directly is just being honest, simple and straightforward, rather then having it always done implicitly. Seems very intuitive to me now that I think about it.
There are four kinds of instructions that play hell with pipeline and OoO design:
- instructions that might cause traps, dependent on the values processed
- instructions that you don't know whether they will change the control flow
- instructions that you don't know where the control flow is going to go to
- instructions where you don't know how long they will take to execute
RISC-V, for example, bans the first category entirely other than load/store, and carefully separates the other three so any one instruction only had at most one of those problems.
ARM load multiple has all of those problems. At least you can examine the register mask at instruction decode time and know whether it will change the PC or not and tag the instruction in the pipeline as being a Jump or not. Imagine if there was a version that took the bitmap from a register instead of being hard-coded...
Load/store multiple don't increase performance much if at all on a CPU with an instruction cache and/or an instruction prefetch buffer. On an original 68000 or ARM without any cache, sure, a series of load or store instructions requires interleaving reading the opcodes with reading or writing the data, while load/store multiple eliminates the opcode reads. An instruction cache also eliminates them, leaving only the code size benefits. But load/store multiple is a perfect candidate for using a simple runtime function instead, at least if you have lightweight function call/return as RISC designs usually do.
The performance of movem.l in the MC68000 comes not from multiple load, but from multiple store, because the main memory access incurred a tremendous, extremely punitive penalty. This has not changed, even decades later, in systems with the fastest memory chips available: any writes to random access memory incur tremendous penalties.
PC being a general-purpose register was historically not uncommon. PDP-11 and VAX both did it and they were kinda popular at one time.
Load/store multiple was also fairly common with, for example, both 68000 and VAX having it. IBM 360 also, though using a register range rather than a bitmap -- a less general solution, but good enough, and much easier to make go fast.
Worse, iirc what looks like it should be the slot for the "never" condition actually does exactly the same thing as "always".
The A64 manual says 1111 on a Bcc etc disassembles as NV but does the same as AL.
The ARM7TDMI manual says 1111 is reserved and don't use it. I don't know the actual behaviour.
Aha. The Welsh&Knaggs "ARM Book" says before ARMv3 NV meant NV. In ARMv3 and ARMv4 NV is unpredictable. And in ARMv5 NV is used to encode "various additional instructions that can only be executed unconditionally".
Sorry, its been a while, but i found the idea, to generate those instruction isles into the code quite weird.
What happened to the Mill "unlimited weirdness budget" CPU guys anyway?
Ivan was saying that they weren't sure about how to get money, and then someone quite rightfully pointed out that this is probably the most liquid time in years for funding startups.
Mill team if you're reading this: Get your IP sorted out then open source everything, then start talking to VCs aggressively. Not the other way round because everyone thinks you're either a joke or a group of fantasists.
Cool ISA, hopeless execution. Where's the FPGA model?
I feel like the Russian "Elbrus" architecture did the same exact thing.
At this point, we can safely say that Mill was just an odd engineer's weird fetish.
Being VLIW was probably the least weird part of Itanium. The NaT ("not a thing") bit on every register was the weirdest part and probably the one which tripped people the most. And like SPARC, it had register windows. It also had a pair of stacks instead of a single one, a set of predicate registers (which also had their own register windows), and plenty of other assorted weirdness.
And on the frontend, sure they mark instruction boundaries within an instruction bundle, but you still need N barrel shifters (where N is the number of max instructions per cycle) to get at them, and that's the expensive part of wide decode. Even 16bit aligned is so much nicer for the critical path than 8-bit, bit aligned is absolutely crazy town.
The biggest thing that they seem to have is that they can compute an enormous number of instructions per cycle without blowing up the size of the bypass network.
Granted, this is a top of the line superscalar processor with AVX-512 and megabytes of cache, but my point still stands. The vast majority of the gates are in everything else: cache, execution units, branch prediction, load/store, die interfaces, etc.
[0]: https://twitter.com/GPUsAreMagic/status/1256866465577394181
Instead, launching a new architecture is an herculean task. Trying that inside anything that isn't a huge company or consensus on the academic world has a history of nearly 100% failure since huge companies started building computers. (And even the consensus on the academic world is iffy.)
The point of designing a new architecture is you have a point to make (generally, "doing X will lead to faster execution"). So by definition you are adding unfamiliar architectural weirdness, else why get involved.
The big problem is that the pervasiveness of the C model has fossilized design decisions of the PDP-11 that still haunt us to this day. For example a lot of transistors are devoted to hiding the effects of the micromachine (superscalar implementation, OOO execution etc). Newer paradigms (even trivial older ones like vector instructions) don't get used enough because of higher level language constraints (e.g. C's memory aliasing model). There's a lot of innovation going on but it's all stick in one small bay of a large lake.
When you're in a C straightjacket it's hard to take advantages of other architectures like transputer, connection machine, or cell. Ever run bsd unix on a Cray? It was dog slow because the CPU assumed very deep pipelining and always had to continually resynchronize because of the frequent branches in the C code.
GPGPU is the sole recent exception, and even then it's still at the margins. And we have CUDA which is an attempt to, yes, get back to the C model.
These modern CPUs benchmark pretty well, but in practice little code can really take advantage. For example the M1 is a multicore screamer, yet most apps are still restricted to a single one.
CUDA allows different programming models, it's the point of having the PTX virtual machine as the exposed programming model.
CUDA C++ is popular, but is far from the only option people have.
(and it can go very deep, https://www.ibm.com/docs/en/sdk-java-technology/8?topic=egpu... for example)
FORTRAN has tighter restrictions around what you can do in terms of aliasing/etc, which allows compilers to be more aggressive about optimizations without as much undefined-behavior "this looks correct and will run fine until we add another optimization and it doesn't" shenanigans. So in practice it can be faster than C, and it's actually quite popular among the HPC community and other performance-sensitive segments.
It's not like C has to be this way, with all the "undefined behavior" nonsense, it just is an artifact of a programming language written in the early 70s. C makes a lot of sense when a compiler looks a lot more like an assembler with a bit of optimization sprinkled in, and isn't performing super deep introspection and deciding that this entire function can be optimized away because of some arcane "undefined behavior" rule.
I'd say that the most popular option at the end-user level today is a whole level of abstraction above: using Python w/ PyTorch (or less frequently, TensorFlow).
Combined with CuPy and cuNumeric, which is available at https://developer.nvidia.com/cunumeric and handles scaling to multiple machines with Python-written code, through the Legate runtime (https://nv-legate.github.io/legate.core/README.html). This covers a huge amount of what the CUDA userbase uses.
That's why you have things like memory order constraints on atomic operations in C, and why you need to be extremely careful with how you place things like memory fences when doing anything fine-grained between several threads.
Generally, most programmers shouldn't need to be reasoning about OOO or memory order and be using fences. Rather, they should use libraries that are written by experts that use extra-lingual mechanisms (like compare-swap or fences) to guarantee a higher contract.
I mean, you're basically saying that it's only observable if you care to observe it.
I seem to recall that there was a mainframe back in the 1970s that actually did that. Univac 1108 or 1100, maybe? Maybe only in some particular mode?
does this mean that because you can’t be sure if two vectors don’t overlap in C you therefore can’t auto-vectorize loops?
Ha! As if C compilers really cared about correctness. This is why type-punning and aliasing memory through different typed pointers is fraught with peril, often UB. It's so C compilers can cheat and use type-based alias analysis, breaking programs that have "UB" but would work absolutely fine if the compiler wasn't so aggressive.
In general, almost all arguments between language design and optimization for C/C++ are settled with "let's not think too hard about how to help users; go ahead and optimize this and make it their fault (UB) when they observe the hard cases". It's so rude.
C is especially awkward here, because compilers are expected to precisely conform the C specification, but at the same time C users don't actually want to program for the abstract machine from the C specification (the one that can't overflow signed integers or shift by a full int width). C users expect their programs to actually behave like programs for a particular hardware they target, without any overhead or discrepancy from "emulating" the imaginary C-spec machine. In face of such conflicting requirements, C compilers will always be forced to balance performance, predictability, conformance, and compatibility.
This is tangential and somewhat pedantic, but I want to point out an oft-repeated mischaracterization.
C is essentially a refinement of B with some additions [1], and B was written for the PDP-7, which was a very different machine.
For example, the increment and decrement operators in C were inherited from B; they could not have been modelled after the PDP-11. From Dennis Ritchie himself [2]:
> Thompson went a step further by inventing the ++ and -- operators, which increment or decrement; their prefix or postfix position determines whether the alteration occurs before or after noting the value of the operand. They were not in the earliest versions of B, but appeared along the way. People often guess that they were created to use the auto-increment and auto-decrement address modes provided by the DEC PDP-11 on which C and Unix first became popular. This is historically impossible, since there was no PDP-11 when B was developed. The PDP-7, however, did have a few "auto-increment" memory cells, with the property that an indirect memory reference through them incremented the cell. This feature probably suggested such operators to Thompson; the generalization to make them both prefix and postfix was his own.
Another person puts it this way [3]:
> It's a myth to suggest C's design is based on the PDP-11. People often quote, for example, the increment and decrement operators because they have an analogue in the PDP-11 instruction set. This is, however, a coincidence. Those operators were invented before the language [i.e. B] was ported to the PDP-11.
[1] Notably, the char type, since the PDP-11 was byte addressable but the PDP-7 was not.
[2] https://www.bell-labs.com/usr/dmr/www/chist.html
[3] https://retrocomputing.stackexchange.com/questions/8869/have...
"char" is a useful type for the program's problem domain regardless of whether the machine is byte-addressed or not. Pascal was developed on machines that were not byte-addressable but always had a char type. This is easily visible in the PACK and UNPACK standard functions and the PACKED keyword for arrays and records. It is slow to do access to a random char from an array packed several to a word, so there are standard functions to copy everything from a packed array to an array that uses one word per item for easier processing, and back again.
If anything, you need the compiler to support the char type more if you don't have a byte-addressed machine because it's such a pain for the programmer to deal with themselves.
Standard Algol 60 didn't have "char" but some manufacturers added it non-standard. Algol 68 had char. These ran on primarily word-addressed machines too, as was the fashion in those days.
Though a general purpose machine it had instructions specifically designed for low cost Lisp implementation, arbitrary numbers of stacks etc…all inaccessible to C
Algol68 allowed char to be as small as 6 bits, the size on the ICL 1900 it was first developed on (with a 24 bit word size).
Pascal was first implemented on the CDC 6600, with 60 bit words and characters again 6 bits. The first other machine Pascal was ported to was, again, the ICL 1900.
Examples of non-pdp-11 things are discussed in various comments in this thread.
Once instruction sets got features such as index (or base) registers and register indirect addressing (including stack addressing) instead of absolute addressing (including for indirect jumps/subroutine calls) there was no longer any need to support self-modifying code during execution of the program itself, and Von Neumann architecture is not required except for initial loading of the program code into memory by the OS.
I'd think most modern instruction sets are easily capable of running with code and data in truly different address spaces.
The only real difficulty is in loading constant data and especially constant tables from program space. Microcontroller ISAs such as AVR have special instructions for program space loads. Any ISA where the PC is a GPR (PDP-11, VAX, arm32) or that has an explicit PC-relative addressing mode (preferably indexed) such as 68000 make it easy to detect that a memory reference is PC-relative and do it in the program space instead of the data space.
Arm64 with ADR and ADRP and RISC-V with AUIPC make it tricky because hardware would have to track that a GPR contains a pointer derived from the PC. x86 and PowerPC make it even more difficult, because the only way to get the PC value is to do a fake function call to the next instruction (or to keep return stack prediction happy, to a real function that just saves the return address then returns). On these ISAs you're probably better off assembling all constants using load immediate and shifts, and arrays or tables of constants using computed jumps to load immediate instructions.
In short, using a conventional modern ISA with Harvard architecture is doable if you want to.
To be honest, I never really understood why more architectures don't have upwards-growing stacks. Visually, one adds things to the top of a stack, after all, so it makes sense to increment the stack pointer. I suppose it's a relic of the old days when folks just set the start of the stack to some remote part of memory and hoped it never grew down into the code or the heap?
> + these days, being big-endian is spending weirdness budget
Hard disagree, big-endian is right and little-endian is wrong. At least, unless you think this year is 2202. I don't know how Intel got this so badly incorrect.
And remember that network byte order is big-endian. Any protocol which emits little-endian bytes on the wire is broken by design.
But then again, HPPA makes it more obvious that stack should be limited size because the above only works for single-threaded processes
Worse things about HPPA is delayed jumps... no wonder they were the ones that partnered with Intel on Itanium.
Network has no natural order (do you start with the last bit?)... for new stuff I'd do the little-endian any way if I cared to bother.
Thanks!
I agree home/consumer stuff is just basicly coasting on 3-4 generation tech though. I ran cat-6 cable around the house hoping to get 10G sometime in my lifetime for less then the cost of a kidney...
If I wanted to show where ordering truly does not matter, I would rather look at encoding schemes. Some versions of Ethernet use https://en.wikipedia.org/wiki/64b/66b_encoding , meaning that the smallest unit that can be read at once is 8 bytes. That removes the ordering from IPv4 addresses, but IPv6 can still get handled earlier by taking ordering into account.
"When the bytes which make up such multi-byte numbers are ordered from most significant byte to least significant byte, that is called "network byte order" or "big endian." (1)
Seems pretty straightforward MSB to LSB? which is what I learned way back in 'C' times...Did something change ?
1) https://datatracker.ietf.org/doc/html/draft-newman-network-b...
A downward-growing stack is more intuitive. To access items on a downward-growing stack, you use positive offsets; for instance, sp+0x10 could be the second stack-passed argument to a function. For an upward-growing stack, you have to use negative offsets.
> Hard disagree, big-endian is right and little-endian is wrong.
Little-endian is more natural. With little-endian numbers, the byte b at offset n has value b*256**n; with big-endian numbers, the byte b at offset n has value b*256**(l-n-1).
> And remember that network byte order is big-endian.
That's mostly by historical accident; the architectures originally used to implement the network protocols we use nowadays were big-endian. It makes sense in some situations to have on-wire numbers be big-endian, especially network addresses, since it allows the hardware to decide before the whole number arrives; that is less relevant nowadays because not only networks are much faster, but also it's more common to read the whole header to a buffer before making these decisions.
I would expect you to use negative offsets to access the previous elements.
Why would one use negative offsets with the stack pointer instead of positive offsets with the base pointer?
Using the stack pointer, I can never dynamically push something onto the stack without changing the offsets.
If you think a bit about it, it would have been much more natural to write number in the opposite order compared to what we have now, since we read and manipulate numbers from least-significant digit to most-significant digit and read text left-to-right.
When you read a number like 346624596, you need to parse it twice to know its value, but also to know its magnitude or do anything significant with it. In the opposite order, everything becomes simpler: adding two numbers, knowing the magnitude or ignoring the least-significant digits ("its about 350 millions"), ...
Our numbers originate from people using a right-to-left script...
[1] https://mathshistory.st-andrews.ac.uk/HistTopics/Bakhshali_m...
I don't know of any new architecture that is choosing big-endian. Almost all of the ones that I can think of that are/were big-endian have added a little-endian mode.
Thank god for that! In Wasm we just decreed little-endian, and I think this was one of our better calls.
For example, endian-ness looks like weirdness to a programmer, but not to a chip designer that needs to read in lower-order bytes first to decode an instruction quickly. Why not just change the instruction decoder? Because it needs to be backwards compatible to what was once a byte-sized ISA.
But for programmers who don't understand the history, it seems insane.
Second, processors like the Motorola 68000 were 32-bit, so there was no need to read just the first byte. Refer to my point about instruction fetch. Like I said, if you have an ISA that came from byte-sized opcodes, and then expanded later, you need to fetch bytes first. However, if your opcodes are 32-bit, then the code is bigger. (In the old days, CISC code size was a selling point compared to RISC, this opcode compression was one of the reasons.)
This is what you learn in first-year computer architecture. However the world has changed a lot since the 80's. Single bytes aren't fetched in x86, 256-byte cachelines are, and the decode happens on that.
The moment you help perpetuate a monoculture rather than subvert it and fight it at every turn and opportunity, you become part of the problem.
+ MIPS branch delay slot
+ Everything about VLIW
+ Everything about Mill
+ iAPX 432 bit variable length instructions
SPARC register windows: once one gr0ks them, they are a feature of wonder! The register windows in an UltraSPARC processor effectively provide 256 virtual registers in a processor with only 32 physical ones, giving one increased performance.
Michael and I talked about exception design a lot in g++ (I certainly had experience of what to do and not do from CLOS) and there was never a consideration of doing anything that might impact runtime performance.
At least now I finally understand why some people, especially in the game industry, want to disable exceptions.
As a side point, I'm disappointed by Raymond's (18 yo) claim that "Zero cost exceptions" aren't really zero cost. Everything called "zero cost" in C++ means "doesn't add any runtime cost if you don't use it". He argued against a straw man.
https://devblogs.microsoft.com/oldnewthing/20220418-00/?p=10...
Table-based exception-handling metadata requires function prologue/epilogue sequences to have constrained forms that can be described by that metadata. When NT was first designed for x86-32, there was a lot of existing asm code that people wanted to easily port over to NT, and that ported code included pretty much every weird function entry/exit sequence you could think of (and lots more that you wouldn’t imagine anyone would ever think of). Switching to table-based metadata would have required modifying all that code, making porting more difficult. At least, I assume that’s the reasoning – I was deeply involved in the SEH runtime code in the ’90s and ’00s when I was on the VC++ compiler team.
Note that in Microsoft compiler land there are two kinds of exception, regular C++ ones and "structured exception handling". https://docs.microsoft.com/en-us/cpp/cpp/try-except-statemen...
The latter does exception unwinding in the kernel on runtime faults (segv, division by zero, etc). It does _not_ unwind destructors or deallocate storage. And yes, I know about this because I had to debug it, on WinCE ARM, where we eventually discovered that there were two different MS compilers, one of which generated code that could be unwound and one of which didn't.
The code worked literally everywhere else, but it turned out that on Windows longjmp doesn't just load the registers from the buffer but also does some SEH bollocks. If I recall (and this was in 2006), it didn't matter for the exception handling, but it was a hard crash for switching threads.
So I had to write my own setjmp/longjmp for Windows in assembly language (i.e copy theirs and cut out the SEH bollocks).
IANACW, but I think that in principle is almost always possible to implement truly zero cost exception path in the non-taken path, by moving all necessary compensation code to undo optimizations into the exceptional path, but in practice it might be too hard for compilers to do it.
But OK, it could potentially be nonzero. Let's say epsilon cost instead.
On modern CPUs, probably the most efficient exception model is one that simply uses error codes. Especially when paired with an optimised calling convention like what Swift does. Checking for an error flag is essentially free on a superscalar CPU anyway and all this stuff is transparent to the compiler, simplifying the translation process and enabling more optimisation opportunities.
P.S. I ran a bunch of tests with C++ a while ago and a Result-like type error handling (implemented sanely) was always as fast as C++ exceptions for the good path and faster for the bad path — unless you are going a hundred or so nested functions deep (but then you have a massive code smell problem anyway). With an optimised calling convention it is likely to be even better.
Yeah. At its root this is a Conway's Law bug. The people responsible for "exceptions" had a senseless wall between them and the people responsible for "stack frames".
Can someone explain how is this x86's fault?
The compiler knows f1 and f2 live in the abstract C machine and cannot access it.
I think its like the difference between errno, a global inside the abstract C machine, and __foo, which a compiler might create, and, if it did so, could assume to be completely under its control.
So either: exception handling in 32 bit mode is sort of an hack and the exception handling code generation doesn't actually expose the store the store through fs to the rest of the optimizer or, more likely, things are much much more complicated than described in the article.
Probably stores through fs are pragmatically not considered globals stores as they would hinder other optimizations.
Um, it's always been on MS's domain. It has moved from e.g. msdn.com (I think) to microsoft.com, but that's more about MS periodically moving around where its documentation is hosted because shrug, but it was always in same way officially hosted by MS.
As far as I know, The Old New Thing is the only blog from that wave that's still being written - and it's been written every weekday for essentially that entire time. To give an idea of the scope of the work, he did a compilation of posts in book form sixteen years ago, and it was over 500 pages long.
https://www.amazon.com/gp/product/B004YWCL5W/ref=dbs_a_def_r...
I don't read it daily anymore (hasn't been relevant to my work or professional interests for ten years or more), but I do read it occasionally and continue to be both impressed and grateful for the effort he's put into the blog.