A secret Apple Silicon extension to accommodate an Intel 8080 artifact
bytecellar.com
bytecellar.com
The Datapoint 2200 was a programmable computer terminal announced in 1970 with an 8-bit serial processor implemented in TTL chips. Because it was used as a terminal, they included parity for ASCII communication. Because it was a serial processor, it was little-endian, starting with the lowest bit. The makers talked to Intel and Texas Instruments to see if the board of TTL chips could be replaced with a single-chip processor. Both manufacturers cloned the existing Datapoint architecture. Texas Instruments produced the TMX 1795 microprocessor chip and slightly later, Intel produced the 8008 chip. Datapoint rejected both chips and stayed with TTL, which was considerably faster. (A good decision in the short term but very bad in the long term.) Texas Instruments couldn't find another buyer for the TMX 1795 so it vanished into obscurity. Intel, however, decided to sell the 8008 as a general-purpose processor, changing computing forever.
Intel improved the 8008 to create the 8080. Intel planned to change the world with the 32-bit iAPX 432 processor which implemented object-oriented programming and garbage collection in hardware. However, the 432 was delayed, so they introduced the 8086 as a temporary stop-gap, a 16-bit chip that supported translated 8080 assembly code. Necessarily, the 8086 included the parity flag and little endian order for compatibility. Of course, the 8086 was hugely popular and the iAPX 432 was a failure. The 8086 led to the x86 architecture that is so popular today.
So that's the history of why x86 has a parity bit and little-endian order, features that don't make a lot of sense now but completely made sense for the Datapoint 2200. Essentially, Apple is putting features into their processor for compatibility with a terminal from 1971.
Parity is expensive to compute with 1970s hardware. If you look at the die of the TMX 1795, the parity flag computation is a substantial part of the die, about half the size of the register file or the ALU. The first problem is that XOR is an inconvenient gate to implement, especially if you use standard MOS logic. Processors usually have special pass-transistor tricks to make XOR more compact. The second problem is that parity needs to XOR all the bits together, so you don't get parallelism.
Computing any of them is equal amounts of delay if you want to mux/select any of them to output, and then on top of that you have to sequentially NOR-chain (and/or some hierarchical manner) the result for zero, and waiting on the adder carry chain for the carry flag. If you were sequentially (and/or some hierarchical manner) XOR-chaining for parity in parallel, is it really that much more delay? Especially if XOR uses domino logic etc.
I see a parity in the same realm as the adder carry chain as far as prop. delay goes (for computing sign/carry at the end.) Of course carry chains can be made more efficiently than a direct sequential chain, but you could be hierarchical about xor-chains for parity too..
Technically parity could become a stable value before the zero flag could! ;)
Oh but I did forget that a single multiple input NOR could be large but without exponential amounts of gates
I'm still displeased with macOS Catalina and iOS 11 dropping support for 32-bit apps (including many games) but Apple has a long history of old apps not running on newer OS versions and newer hardware.
Over time macOS dropped support for 68K, PowerPC, Classic, Carbon, and 32-bit apps. We'll undoubtedly get a macOS that won't run on x86, and I will be unhappy, but not surprised, if/when Apple removes Rosetta 2 and x86 app support. I can even imagine an unpleasant time when they remove the Objective-C runtime and frameworks.
I'm sure there were some Apple 1 users who were unhappy that their software didn't work well on an Apple 2, even if the Apple 2 had much better hardware and software.
Intel tried to force a split in desktop/server design with a significant price premium and shipped a very complicated CPU that moved significant complexity/work into the compiler. It managed to scare Alpha, Mips, and PA-risc out of the market, and generally shipped years late and hit performance milestones even later.
AMD on the other hand doubled the x86's registers, added 64 bit support, improved performance and did so with basically zero increase in cost. So suddenly a random x86-64, even a desktop, could gasp make full use of over 4GB ram at zero cost.
So Itanium retreated from servers, to HPC, to enterprise, and HA enterprise then of course died.
They must have performed deep analysis of the requirements of Rosetta down to the level of individual instruction emulation before they finalised the design of the M1.
Also the Rosetta team had enough influence to get the hardware team to build it. Not necessarily a given in all organisations.
What would be interesting - and perhaps not surprising - is if Arm added this to the standard Arm ISA in order to facilitate x86 to Arm translation elsewhere. After all they added the infamous FJCVTZS instruction which speeds up x86 derived floating point conversions in JS.
Bloomberg estimates that Apple intends to ship less than a half-million headsets it's debut year. By comparison, the Meta Quest 2 sold 10 million units in it's first year at sale.
> Also, surely anything at that price point is going to be positioned against the Quest Pro, right?
At what, $2,000? At that point, you're competing against everything. You're competing against the price of a new-in-box Valve Index with a VR-capable computer to-boot. You're competing against the price of buying everyone in your family a $300-400 headset. You're competing against the price of a Playstation 5 with PSVR2.
This product is a suicide mission for Apple. It will come out, and it might even be great, but it's value proposition is non-existent in a market that already lived through Beat Saber, VR Chat and Half Life Alyx.
(a) for apps and (b) for hype, while Apple works on bringing an actually viable consumer headset to market in 2024.
This is a pretty bold departure from Apple's normal strategy of "wait until several others have rushed to market and them blow them out of the water with a 'premium' product you claim is better than all the competition, at a higher price point than all the competition."
In part it's because companies like Pimax are already sort of doing that, and because Meta got Apple scared that Zuckerberg was going to beat them to monopolizing the VR market if they waited too long.
(We'll have to wait and see the reception to the Quest 3 to see how concerned they really should have been, because the reception for the Quest Pro has been... not great. Turns out, making a quality headset is just expensive)
JS libraries break all the time. APIs and experimental language features break sometimes.
But the low-level core fundamentals of the language, such as the way math works, have never changed in it's entire existence. Even the fairly minor changes in strict mode require explicitly opting-in.
Instead, they really just wanted a deeper selection of binaries to analyze for other various reasons.
I would bet real money that this goes away with the generation that no longer supports rosetta2
Of course credit goes to Intel here, because they're the ones that retained that 8080 behavior even in their latest x86 chips.
Apple has made Rosetta 2, one of the most impressive software/hardware mixture ever, and they can’t wait to get rid of it.
Microsoft will make a crappy conversion shim and keep it in their code forever. For. Ever.
This is such a BS take, the nice thing with opensource is that you can go look at the commits that say remove windows 7 support from python. And pretty much anyone who isn't a child can see its less a technical move, and more political. Because out of a million line codebase those half dozen lines were causing so much grief.
Frankly, all this "we have to remove legacy ports" and "legacy code" is some kind of OCD levels of mental illness.
That you have to call this approach a mental illness is poor criticism.
Technical debt is when people make a mess, its perfectly possible to clean that mess without breaking ABIs. Particularly in OS's where the ABIs tend to be decoupled from the underlying code by abstraction layers. For example you can swap the filesystem in use and still maintain the behavior. Linux's syscall ABI has been basically static for decades, its only the userspace layers that don't try and adhear to those levels of compatibility.
PS: We like to give apple all this credit for having a "fast" machine, but the real question should be, if it can't solve the problem I have, does it matter how fast it is? Mac's for the vast majority of the people I see using them are basically trendy chromebooks, where the users are spending 99% of their time in safari. So its probably a good choice for apple if that is the user base they are interested in.
And it only works when your application doesn't have HW requirements that can no longer be met. Or the company in question just doesn't provide drivers anymore.
Frankly, the driver thing is somewhat understandable, but as you point out its easy to bolt software compatibility layers on, why the OS vendor can't manage to make it work well/transparently is a mystery.
macOS doesn’t ship a 32-bit version, because the 64-bit version on every platform is much faster and more secure, and shipping both would be twice the disk space.
(More secure because you can do so many tricks like PAC with those spare bits in every pointer.)
Apple developed the 68k emulator used in PowerPC Macs in-house, and never removed it.
Apple licensed Rosetta from Transitive at a time when they didn't have tens of Billions of dollars in cash lying around.
Rosetta 2 was developed in-house, but even if they were still paying a licensing fee every time they shipped a new OS version, it wouldn't even be a rounding error to their bottom line today.
Why would they be in any rush to remove Rosetta 2 at all?
Due to the belief that any backwards compatibility comes at the cost of a higher difficulty moving forwards, I suppose.
There will probably come a time in which a new feature will collide with this compatibility layer and I have little doubt that the compatibility layer will lose.
“This almost entirely prevents inter-instruction optimisations. There are two known exceptions. The first is an “unused-flags” optimisation, which avoids calculating x86 flags value if they are not used before being overwritten on every path from a flag-setting instruction.”
Does this mean the emulator is buggy (if I write a loop that sets and clears flags, but doesn’t do anything with them, an interrupt handler that samples the flag values would need to see both values), but in a way that no sane code would notice?
It does have to emulate signals but its moderately easy to handle those by deferring them to "safe" points where all architectural state is known
The article says: "While almost no modern applications read these AF and PF bits"
It can't really be legacy software, as most AMD64 software on the Mac is fairly recent (2006) and I suspect that we're talking software older than that for the AF and PF to be in regular use. So it also have to be something fairly important like virtualization or something used in image processing or compression. It seems like overkill for something that would be a problem for an end-user application that would transition to Apple Silicon anyway.
Since flags can be load/stored or push/pop-ed I'd speculate avoiding the hit in correctly handling every flags access was deemed to be worth it. Based on experience from a long time ago (8-bit micros, z80 in particular, and implementing a 6800 emulator, DAA was more code than any other): unusual opcodes and edge case behaviours are common in optimisations, and anti-debug/disassembly. Games perhaps?
In any case, I find the idea of a correctly behaving, deterministic binary translation more appealing than a bunch of beartr^W on-the-fly software fixups.
It's very hard to conclude Apple care much about breaking a bunch of games on the Mac when they cut off 80% of the Macs gaming library a year earlier by removing 32bit execution, presumably mostly to ease the use of Rosetta 2.
And, y'know, literally every other action Apple has taken for twenty years.
People here complain about software subscription models, but then also like to complain about OS's that maintain backwards compatibility. I mean someone has to pay for engineers to spend their days keeping up with all that frequently pointless churn. Why would a game that works perfectly well with a 32-bit address space need 64 bits?
There isn't any reason for old programs to suddenly stop working because someone in the OS/library chain was to lazy to provide backwards compatible ABIs.
There's a reason that BCD is in the middle of EBCDIC, IBM used BCD math so heavily in its decades of history that even their text encoding and punch cards were built around it. I've seen Enterprise software that still needs to read/write and sometimes even do math in BCD for compatibility with old Mainframe apps and software still written in COBOL.
I don't know of any applications the use AF, or the PF result in question, but I haven't looked into it. I'd maybe check some of the big professional things with a lot of legacy history like Photoshop or Excel. As I understand it Rosetta 2 also supports 32-bit x86, not for native macOS applications (where 32-bit isn't supported), but to allow Wine/Crossovers to work, so a whole bunch of ancient 32-bit Windows games might be in scope too.
Microsoft claims they use IEEE 754 floats (https://learn.microsoft.com/en-us/office/troubleshoot/excel/...), but they don’t completely do that.
For example, do
A1: 0.3
A2: 0.2
A3: 0.1
A4: =A1-A2-A3
A5: =(A1-A2-A3)
B4: =A4=0
B5: =A5=0
A4 will show as “0”, A5 as “-2.77556E-17”, and B4 and B5 show these aren’t purely display issues. B4 shows “TRUE”, B5 shows “FALSE”.My hunch is that they sometimes use BCD arithmetic.
More info at http://people.eecs.berkeley.edu/~wkahan/ARITH_17.pdf.
Bill - j'accuse!
It should be false/false for the IEEE floats...
https://en.wikipedia.org/wiki/Decimal64_floating-point_forma...
"formally introduced in the 2008 version of IEEE 754 as well as with ISO/IEC/IEEE 60559:2011."
https://web.archive.org/web/20190629181619/https://support.a...
cmp al, 10
sbb al, 69h
dasOr are there some common patterns where making such analysis is difficult? Something like copying the whole state of status register for later use (don't remember if x86 even allows that) instead of checking flags directly. Although in that case you would need an extra code anyway to reorganize value so that it matches the strucutre of x86 flag register. I assume the order on x86 and arm is not identical.
>In a VM, it isn’t able to configured the host CPU, so it can’t use this functionality. There are two other options. Either, you can skip computing the flags, because they’re mostly useless and most software won’t care. Or you can compute them the long way shown above. Rosetta 2 chooses the second option, and this mostly works out fine, because they have an “unused flags” optimisation that avoids the computation a lot of the time.
So there's an optimization pass that does what you mentioned, it's just not needed when you have the right instruction available, and the thing is probably a wee bit faster if you don't do it.
(I haven't looked at Linux Rosetta 2, other than to note the parity-flag computation coincidentally showing up in a screenshot posted to Twitter, so I'm not sure exactly how often they do the manual computation, but I'm guessing it's anywhere flags are used, in which case the unused-flags optimisation could be extended further by tracking AF and PF separately to the usual flags, which is roughly your suggestion.)
You could still remove the vast majority of the remaining computations with some heuristics that work well, but then you've gone from 100% correct to 99.9% correct, which is nice to avoid when you have the option.
At least Intel could make some more revenue for a few years off of the Apple arm switch with some licensing.
An actual hardware core would have the advantage of better compatibility. And you only spin it up when needed.
But it would probably slow adoption too....
Throw on fully functioning x86 cores and you've defeated the entire point.
Remember, running x86 code in emulation on ARM uses less power than running the native code on Intel.
In some senses it's even more important for Snapdragon - Google doesn't yet ship a native Windows ARM build of Chrome. So for MacBook competitors using Windows on Snapdragon, you either have to use Edge or the emulated x86_64 Chrome.
> A generic Snapdragon ARM SoC, say, would deliver notably less performance in this specific scenario that is critically important to Mac users.
Sorry, I'm not following.
If almost no modern applications read these AF and PF bits, why is it critically important to Mac users?
UPD: and ARM seems to be infringing as well:
> There’s a standard ARM alternate floating-point behaviour extension (FEAT_AFP) from ARMv8.7, but the M1 design predates the v8.7 standard, so Rosetta 2 uses a non-standard implementation.
> (What a coincidence – the “alternative” happens to exactly match x86.
Good luck persuading a judge that this was "a coincidence".
In this case there would be no copying of microcode since the underlying hardware is completely different.
AMD has the rights to x86, and some like transmeta did x86 translation in the chip.
It's very unlikely but there could conceivably be a patent, but it would have long since expired.
No case to answer.
Edit: Just to clarify - referring to treatment of one instruction here. As peer comment has said Rosetta translates ISA rather than implements so it's even further removed from being a copyright issue.
For those who weren't following the 32-bit to 64-bit transition on the x86 world back then, the x86-64 ISA is from the year 2000 (https://web.archive.org/web/20000817014037/http://www.x86-64...), so any patent which applies to that ISA (without the ISA being prior art for the patent) is now over 20 years old.
Thinking aloud, I wonder if AVX 512 is translated?
Interestingly, https://en.wikipedia.org/wiki/X86-64 says this:
> x86-64/AMD64 was solely developed by AMD. AMD holds patents on techniques used in AMD64; those patents must be licensed from AMD in order to implement AMD64
The usage of these flags is NOT common.
And as observed above, if they're not used, then you can "have an “unused flags” optimisation that avoids the computation a lot of the time".
So I don't see how it follows that this is a "critically important" scenario. Reality is that adding the logic to compute these flags is almost free, so it's a cheap (in terms of hardware) optimization to do to squeeze some marginal extra performance. But to say it's a deal breaker? If it is, I don't think you can conclude that based on the evidence presented.
I suppose it would be much harder for JITs and other dynamically generated code.
In other words, it's probably not necessary but was included early on out of an abundance of caution. They knew that if it turned out to be used in some binary somewhere, it would be a major performance killer.
I'd be interested in seeing a benchmark with the hardware flag turned off and the translation/optimization setup used for linux enabled. I bet the difference would be negligible.
Computing them while doing operations is nearly free in hardware, so it makes sense to add them in hardware if you can. It's not nearly free in software, but it's important to do it where not doing it might be obervable in normal flow (tricky things with interrupts are out of luck, even if you always do the software calculation, you could interrupt in the middle of it). In many cases, it's easy to determine the flags aren't observable and you can skip software computation.
Which is... just about anything these days, I suppose? Half of the desktop apps are Electron, half of new CLI apps are in Node.JS, half of old CLI apps, including near-ubiquitous ones like git, are a random assortment of half a dozen scripting languages... I haven't actually counted it properly, but ad-hoc random sampling gives me an impression that a third of typical Linux distro userspace is in Python, and most of it not even compiled AOT.
The overlap of "is JIT" and "doesn't have a runtime for ARM" is overwhelmingly small. That's probably mostly because runtimes and JITs are opensourced, which mean you can just recompile them for the target. Rosetta 2 is more focused on the proprietary space where they don't often develop proprietary JIT runtimes.
The more common instances of JITs are, like you said, Electron/Node.js programs, Java programs, and the odd Ruby program. Anecdotally, I think Rust and Go are very common languages for new CLI apps, and I definitely have significantly more Rust and Go CLI apps than Node.js ones installed at the moment.