ARM or x86? ISA Doesn't Matter (2021)
chipsandcheese.com
chipsandcheese.com
[1] https://www.tomshardware.com/news/tenstorrent-shares-roadmap...
[2] https://www.anandtech.com/show/21281/tenstorrent-licenses-ri...
It still kinda makes me sad that the discussion always stops there. It will still break some of whole classes of programs, it's not magic and can either be finicky or just too limited.
From the top of my head USB and peripheral access in general is a PITA every time a translation layer is added. Raw device access in emulation never has worked "pretty well".
For Mac it wasn't that much of an issue because anyone with deeper hardware needs was probably not there in the first place. But on Windows I'd predict it will be a significant transition pain. A bigger share of people will stick on x86 basically forever to cover their long taily needs.
That said this is a solved problem - because plenty of people have arm based machines, and run x86 binaries on them (it obviously comes with a perf cost, but there’s only so much that can be done when dealing with poor engineering)
I wonder which ones he considers those to be -- the interview doesn't say, sadly...
[1] https://www.theregister.com/2024/08/27/tenstorrent_ai_blackh...
Lunar Lake now appears more than competitive in that regard with even the higher end M3 SKUs, though of course M4 Pro and Max are around the corner. Apple, more than the ARM ISA, seems responisble for the prevailing impression that ARM is inherently more efficient because Apple were simply the first to put such major investment into targeting specifically mobile SKUs, rather than scaling down from server focused products. Snapdragon X showcases that quite impressively, ARM based but efficiency wise Apple, Intel and AMD SKUs appear more competitive.
This [variable-length coding] is a disadvantage for x86, yet it doesn’t really matter for high performance CPUs ...
Why are they leaving out difficilty of building a compiler backend when the ISA has variable-length codes? I would assume an ISA needs to consider its burden on compiler authors. (Itanium is an extreme example of an ISA too tedious for compiler authors)I once extended a Common Lisp compiler to emit machine code for SSE4.2 instructions (specifically minss and maxss). The experience was a bit bad due to subtle differences in prefixes and specific fields needing to be set to activate some mode for SSE4.2 instructions.
Now suppose you want to debug a compiler backend targetting x86. Good luck, x86 disassembly is an undeciable problem because you don't know where the instructions start. Meanwhile ARM has two kinds of instructions (full-size and half-size, called "thumb"). Thanks to fixed-length instructions and alignment rules, you always know whether you're in thumb or not. Emitting machine code (and disassembly) is much more straightforward
You're either using an established backend that does all this for you, or realistically you can throw some money at it and make it go away.
Neither x86 nor arm are easier to compile for, because they both have different quirks that cause headaches. X86 has two operant forms, which require more work. Arm has restrictions on conditional jump distance and immediate encodings, which requires more work. So, an instruction set that had no two operant forms like x86 and that allows any instruction to have any sized immediate would probably be easiest to compile form. Allowing any instruction to have a 32-bit or even 64-bit immediate would lead to either all instructions being huge or to variable length instructions. X86’s use of variable length instructions allows 32-bit immediates everywhere, so in that sense, it makes compiling easier.
> I once extended a Common Lisp compiler to emit machine code for SSE4.2 instructions (specifically minss and maxss). The experience was a bit bad due to subtle differences in prefixes and specific fields needing to be set to activate some mode for SSE4.2 instructions.
I assume this was some toy compiler or a non-optimizing compiler. LLVM or GCC (or any other industrial strength optimizing compilers) have no trouble whatsoever dealing with any of those. The difficulty with more complex instructions like vector instructions is in optimization / being able to find the code pattern that can take advantage of the complex instructions, and that has nothing whatsoever to do with them being variable length encoding or prefixes or knowledge about instruction set themselves. If the program is already written for it - e.g. using intrinsics - emitting and mapping to the machine code is trivial, regardless of how complex the instruction encoding rule is.
I dont know the definition of "toy compiler", but compare the following (x86 backend vs arm64 backend)
https://github.com/Clozure/ccl/blob/d960a0e/compiler/X86/x86...
https://github.com/Clozure/ccl/blob/d960a0e/compiler/ARM64/a...
I would argue the former is a lot more complex compared to the functional equivalent in arm64
The specific extension to allow e.g. minss would look something like this
(def-x86-opcode minss ((:regxmm :insert-xmm-rm) (:regxmm :insert-xmm-reg))
#x0f5d #o300 #x0 #xf3)
Now try doing the same for e.g. movsxd, and you will have to be careful with the ModR/M byte, due to the VEX prefix changing semantics.This juxtaposition is (I assume unintentionally) hilarious. ARM is the ISA where you might reasonably see two different ISAs (and thus a disassembler needs to handle both) in the same object file, where x86 only has one, at least since the mid-90's. Yet somehow x86 is undecidable and ARM is straightforward!
In reality, the problem you're talking about for x86 is trivially solvable with what's known as recursive disassembly. You start disassembling from known function locations (since you mention debugging a compiler, that means your binary should be fully symbolized, but finding common function entry points for an unsymbolized binary isn't that difficult of a challenge), and then continue disassembling instructions until you get to various kinds of jump instructions. Then you add branch locations to the list, and rinse and repeat until you're out of new locations to disassemble.
As someone who did this literally yesterday, no you do not. Doing this properly requires tracing control flow, which is an undecidable problem. In fact x86-64 is way easier in practice because real code is somewhat self-synchronizing if you just disassemble linearly and for ARM this doesn’t work nearly as well.
Because compiling into the maching code and the machine code decoding by the CPU have different time constraints. One can even argue that a compiler has an infinite time (however impractical it is) to translate a programming language into the binary code, whereas the CPU does not have the same luxury hence the haggle over the ISA's. So it makes sense to leave the compiler code generator out of the picture.
Very amateur-ish question and counterpoint from me: how is Apple M1 performance and power consumption so good compared to the other laptop cpus then?
Apple gets volume from the iPhone.
This is also why x86 outperformed expensive (and so low volume) RISC workstation CPUs for a long time at single thread perf, and then eventually on multi thread perf too.
Do you have data that says otherwise?
So the comparison is Intel vs TSMC. If you consider that TSMC doesn’t just have Apple as a customer, then holy cow the difference is big.