In terms of a “theoretically even better ISA”, if anyone has any proposals, I would love to see links to them. In particular, the article “RISC-V: They Should Have Known Better” does not propose an ISA (except mentioning that the author feels AArch64 is much better), so I find its criticisms dubious.
[1] I will make one ISA propsal: x86_64 should have had 32 instead of 16 registers, but 16 registers at least is quite a bit better than the 4-8 registers the 386 had.
Is "registers" still a real thing or its just an illusion at this point? I am pretty sure there are "virtual" registers or micro instructions (something like that, I forgot the terminology)
Being able to store, e.g. 10 vs 5, local variables in a larger register file has performance implications[0] strong enough to consider larger register files an advantage.
[0] The reduced number of memory-to-CPU and vice versa data transfers for transient computation results.
https://en.wikipedia.org/wiki/Register_renaming ?
But yes, register count is still a thing. Why? If you run out of registers, operations 'spill out' into main memory. Read: L1 (data) cache. Which is fast, but not as fast as CPU registers. And probably uses a lot more transistors & power.
Of course there's limits to that due to # of opcode bits available (eg. 32 regs, 3-operand instruction -> 15 bits needed to encode source & target registers).
Yes, register renaming. What I mean is that why does anyone care still what instruction I am running on? The CPU have its own microops, instruction decoder, its own pipeline, scheduler, etc etc
Why can't you have a CPU that have 1024 registers and pretend that it only have 32?
"[...] and pretend that it only have [sic] 32" is EXACTLY what register renaming is.
You can still run out of value names, though. We rarely run out of them in straight-line code, we sometimes run out in absurd loop code, and we almost always have to do something for our function calls/returns to avoid running out -- we can't just assign each function its own set of value names. This means we need spill/restore instructions, either before and after the call instruction or at the beginning and end of the function. Sometimes both. We also divide our value names into parts that have to be saved by the caller and parts that have to be saved by the callee.
Maybe smarter call/return instructions will help a bit in the future by doing multiple register renames as part of a single instruction in order to make those spills/reloads faster/less necessary/more asynchronous.
Sparc and Itanic tried to do something like that (register stacks), but in a way that ended up being expensive to implement in hardware (high clock speeds were hard) and inflexible in practice -- and also annoying to support for the OS, the compiler, and the debugger.
APX proposes exactly this along with 3-register syntax and some other things.
There are two big issues IMO.
1. It will take at least 15 years before most software ships with this because unlike something like AVX where you typically just rewrite a small part of your code that needs AVX, APX requires a 100% rewrite to take advantage.
2. APX instructions require an additional byte each time you use them compared to current instructions. I think there are still savings to be found, but they won't be as big as it might seem at first.
Lack of serious incentive combined with decades-long rollout seems a recipe for non-adoption at a time when RISC chips offer these features now.