Arch: Remove Itanium (IA-64) architecture
git.kernel.org
git.kernel.org
> I'm a little bias because my company is a re-sellers of the HP Itanium ia64 hardware (RX & ZX boxes), as well as PA-RISC. For that reason, I would hate to see it fade away in any sector. The ia64 platform is still widely used with HP-UX Unix and Open VMS users worldwide. This hardware is embedded in most every data center and large and medium companies that have been around since the 80s/90s, its probably the oldest box they have in there but its the one thats in the corner running for 20 years, long before most people started working there.
Can anyone speak to how true this is? Are there still a bunch of companies that run some critical service using machines/OSs like this? I have a morbid curiosity about "obsolete" tech still being in everyday use.
VMS for x86-64 has just emerged for production deployment in the past year. I don't know if it has the OS-level emulation of previous architectures (macro32 and vest, if I am correct in the VAX components, and their Alpha and Itanium equivalents), but the most modern VMS installations are on Itanium.
Intel was a big user of VMS in the past, although I don't know if that remains true.
TruCluster was definitely interesting, but expensive and fairly esoteric. Most Linux people would just horizontally scale their applications without depending on its features (a filesystem that was shared seamlessly across 8 machines, along with virtual IP address aliasing with fast failover).
More details on the underlying technology: https://www.hpl.hp.com/hpjournal/dtj/vol8num1/vol8num1art1.p...
Still works as of today, AFAIK.
It's funny that companies can ostensibly track their expenses down to the penny but can't recognize the costs of technical debt.
Oracle offers services like hosting services, consulting, training, financing, and compliments these with hardware and software products, manufacturing and design. So.. Sales, sales, more sales, coffee, then sales, some sales, and more sales.
Oracle is a marketing engine. I am going to call that true.
If you found whatever PC that you were running from this era, would it be slower than your laptop? Very likely.
It syphones your data faster. That's all.
?bad grammar?
Plus, value of RDS = not having to upgrade/replicate cluster and backups manually.
Of course there are better managed services...
"Heinz Mauelshagen wrote the original LVM code in 1998, when he was working at Sistina Software, taking its primary design guidelines from the HP-UX's volume manager."
https://en.wikipedia.org/wiki/Logical_Volume_Manager_(Linux)
> None of the original companies behind Itanium still produce or support any hardware or software for the architecture, and it is listed as 'Orphaned' in the MAINTAINERS file, as apparently, none of the engineers that contributed on behalf of those companies (nor anyone else, for that matter) have been willing to support or maintain the architecture upstream or even be responsible for applying the odd fix. The Intel firmware team removed all IA-64 support from the Tianocore/EDK2 reference implementation of EFI in 2018. (Itanium is the original architecture for which EFI was developed, and the way Linux supports it deviates significantly from other architectures.)
...
> There are no emulators widely available, and so boot testing Itanium is generally infeasible for ordinary contributors. GCC still supports IA-64 but its compile farm no longer has any IA-64 machines. GLIBC would like to get rid of IA-64 too because it would permit some overdue code cleanups.
The market evidence suggests it’s not, but there haven’t been that many well-funded, competent attempts (Itanium, Transmeta, ???) so the sample size is small.
As I and others have pointed out, instead of a constant 3 instructions in 128 bits, a variable number of instructions in 64 bits should get you competitive code densities. (Say, have a 4-bit prefix that indicates if 1x60 bits, 2x30 bits, 4x15 bits, or 5x12 bits instructions are packed in the remaining 60 bits. For 12-bit instructions, you probably have the first instruction only able to store into r1 or r2, the second instruction only able to store into r3 or r4, etc. so that you have a 6-bit opcode, 1 bit for destination register and 5 bits for operand register. Plenty of old processors had an accumulator that was the implicit destination for most instructions.) Some of the other 4-bit prefixes would be used to indicate a combination of instruction widths and instruction parallelism to allow lower-powered in-order superscalar implementations.
Fixed alignment of 64-bit bundles of variable-width instructions makes parallel instruction decoding cheaper (and makes static analysis easier/finding ROP exploit targets very slightly harder).
Your high-performance cores are probably still going to have dynamic out-of-order scheduling in hardware. However, your power-efficient in-order cores might use parallelism data embedded in that 4-bit prefix.
Ideally, the hardware would have reservoir sampling for which branches are mispredicted and which instructions stall the pipeline. This information could be aggregated for a background process to re-optimize the binary, similar to what the current Android Runtime does.
Efficient hardware branch tracing (perhaps a second stack where, when the tracing machine state bit is set, the destination of each conditional/indirect branch target is pushed, along with an interrupt when the trace gets full) might allow for efficient runtime re-optimization of binaries, including inlining of dynamic library code into the code hot spots.
On a side note, I know RSIC-V at least at one point had a proposal for an extension to use the integer registers for floating-point. Does anyone have a feel for how costly it would be for register renaming logic to handle separate integer and floating point register files so that low-power implementations could use a single unified register file (perhaps without register renaming) and higher performance implementations (which would presumably have register renaming anyway) could use separate integer and floating-point register files?
In general, I hope we can find ISA designs that leave room for both very-low power implementations that can push some of the work into software, and high-performance implementations that aren't hindered by the features that allow lower-power/simpler implementations.
Speaking of which: so was AMD Terrascale: https://en.wikipedia.org/wiki/TeraScale_(microarchitecture)
-----------
It just needs to be in the right situation. GPUs almost entirely compute inside of register space, while FPGAs have carefully laid out memory exactly perfectly for their compute, so once again VLIW can fly.
SIMD is an important technique to scale vs GPUs, because... well... SIMD is basically free scaling so might as well get it.
So that's the combo IMO. Remove the random memory delays associated with general purpose compute, have a crap-ton of register space to never touch memory under normal circumstances, have huge kernels for DSP / GPU like tasks, and let the machine fly.
----------
I think it so happens that SIMD-parallelism was easier than expected, while VLIW-parallelism is harder than expected. So SIMD takes priority, but both should happen in these use cases. If you consider that the 2000s period of AMD64 vs IA-64 going on, the AMD systems focused on SSE / SIMD parallelism to provide us with the high-performance multimedia DivX decoders and whatnot (and other multimedia problems that consumers needed SIMD compute for back in the day). That parallelism was enough to be competitive vs IA-64.
"General purpose CPUs" are so RAM-latency limited in practice. So out-of-order machines that can find work while waiting for RAM (again: AMD64) get a benefit.
The second iteration, McKinley, was designed by HP and had more reasonable performance, but not enough to stop AMD64.
A big problem was that a quality compiler was hard to find, and without it everything just ran very slowly.
The first issue is the static versus dynamic scheduling problem. Static scheduling first requires a sufficiently smart compiler to output the correct static schedule. But you're also limited in your ability to parallelize based on what you can divine statically. Some operations have fundamentally dynamic execution times (memory operations, branches, and division operations are the most common of these), and if you guess wrong as to how long they're going to take, you're going to force the computer to do nothing where a dynamic scheduler could have slotted other work in instead
Additionally, you can only schedule work to be done in parallel within a relatively small scope, largely a basic block, and definitely on function boundaries. Functions like sin and cos boil down to polynomial evaluation, a chain of ~5 FMAs that have to be executed sequentially. Even if you have two FMA execution units, a static scheduler is forced to execute sin(x) * cos(y) serially, whereas a dynamic scheduler can usually get some overlap between the FMA chains.
A sufficiently smart compiler that can create the optimal static schedule for a VLIW chip can also emit the code so that the superscalar chip will dynamically choose the optimal schedule, so a VLIW chip ends up getting no better throughput than a regular superscalar chip, and if the stars align less well, will likely get worse throughput.
The other fundamental issue with VLIW is that it forces you to expose more of your microarchitectural details at the ISA level, which makes it harder to adjust those details (e.g., add more execution units) should you desire to in later chip iterations. Now Itanium does have a design which ameliorates this to a degree, but the cost of such a design is largely bringing back the transistor- and power-hungry aspects of the superscalar chip into your VLIW design, at which point you start asking how much you're really saving.
Had it not been for them, Intel and HP would have managed to push Itanium no matter what.
I always wonder if there's some kind of strange explicit smt or on-thr-fly instruction building or swapping... dynamically melding coroutines into ones main code.
Simics also used to support Itanium years ago, but unfortunately that support doesn't exist in the version that Intel released to the public a few years back