Fujitsu Switches Horses for Post-K Supercomputer, Will Ride ARM into Exascale
top500.org
top500.org
Honestly, ARM is a very complex architecture these days. There's somewhere in the neighborhood of a thousand instructions. It's very close to x86 in complexity. Tons of cruft built up over the past thirty years too (although it was a cleaner architecture than x86 to start off with so it has that going for it).
We fought that battle 15 years ago, and delay slots lost (for general-purpose CPUs).
Edit: These machines also do a lot of cool stuff under the hood. DAISY's native binary format looks more like a CFG than we pure instruction steam. The Crusoe has some really sweet exposed speculation hardware. I'm sure Denver does some really cool stuff too, there's just not a whole lot of publicly available information on it.
What is the acronym "CFG", I wasn't able to find that anywhere.
[0] https://www.youtube.com/watch?v=oEuXA0_9feM
[1] https://github.com/ssvb/tinymembench/wiki/Nexus-9-(Tegra-TK1...
I'll have to disagree, I think ARM mostly did the right thing with the A64 instruction set. The AArch32 instruction sets were just getting ridiculous (just having multiple instruction sets for one architecture is pretty messed up), with A32, T32 and a bunch of other extensions (ThumbEE, Jazelle) which pretty much no one used. In addition, the A32 and T32 instruction sets have grown organically over time, with new instructions being crammed in any unused / invalid space, which is how we ended up with one encoding doing this one thing, except if one register is the SP or PC, then that is a whole other instruction or with immediate values being split up over 2, 3, 4 fields. Also, a few features had turned out to be pretty bad ideas: for example access to the PC as a general purpose register is overly complicated to implement in OOO cores (verification cost) and comes with performance penalties, but it still had to be supported; the IT instruction is another fun one, especially in the context of dealing with exceptions.
A64 pretty much takes the set of instructions which has emerged over time, removes the bits which are not suitable any more or were bad ideas in retrospect, fits them cleanly in a new encoding space and removes some of the flexibility which was getting in the way (e.g. VFP having either 16 or 32 double precision registers). I expect AArch32 to be gone from most high performance implementations in a few years, with a software DBT option remaining available for legacy software which still hasn't been ported, which is when the implementers will be able to fully benefit from having a simplified AArch64 compared to AArch32.
Apple has been requiring iOS apps to include 64-bit binaries for more than a year, and in iOS 10 they're publicly shaming 32-bit apps [1] on launch. Wouldn't be surprised at all if 32-bit support is gone from iOS and Apple's SoCs by iOS 12 / 2018.
[1] http://appleinsider.com/articles/16/06/15/ios-10-warns-users...
In effect thermodynamics (or some relative) once more have us by the balls...
Ease of decode - You want your instruction stream to be dense to fit easily in I-cache. You want it to be self-synchronizing for easy multiple decode. You want the input registers to be at predictable offsets from the instruction start to speed up fetching them.
Data Dependancies - The design of your OoO execution engine will depend on the maximum number of outputs and especially inputs an instruction can have. Flags are inputs too. If some instructions can only partially update your flags (like x86) then the flags count as multiple inputs now.
Atomic execution - Handling precise exceptions from timer interrupts or page faults or whatever mean you have to go back to some coherent state. This is easier if your instructions only touch memory at most once each. Especially difficult is multiple writes to memory in an instruction but read-modify-write is also bad. Read-modify and modify-write are pretty ok but still add a bit of complexity.
And one can argue about whether it's better to have 16 or 32 registers but with x86-64 everybody these days has one of those and is good enough. Well, actually SPARC with its register windows has way more than 32 and that might be Fujitsu's biggest motivation in moving to ARM.
It seems like the Parallella tried to get close, but did not succeed in breaking into the market. (On Amazon, looks to be marketed more for a high powered raspberry pi alternative...)
Which seems to be more useful? The power to compile same code with alternate flags (Xeon Phi style); cross-compile to other arch's (x86 host -> ARM binary); JIT bytecode, LLVM bitcode, or Intermediate-Representation formats?
What kind of access patterns would be most common for a hobbyist or an enterprise to cater to? For example, one issue the Xeon Phi has is memory controller contention, which makes it less optimal for less structured relational analysis.
Beside big picture, there are such pesky details as "Fortran for ARM" HPC compiler. All these bearded and not so guys who walked Sun hallways for years ... The Intel HPC compilers also well established. What is available for ARM in that department?
This hardly looks "completely new" compared with Sunway, for instance, especially given work with it in Europe.
Fujitsu have their own compiler -- amongst their own essentially everything for K -- but GCC is "well-established" in HPC and you can see the Fortran-based HPC-type software in Fedora and Debian aarch64. (Yes, there's a lot more to it than that.)
I wonder how much better the library is than the free alternatives.