I've been reading the detailed pdf to see more details. It looks to me like 32-bit mode, specifically in the sense of a 32-bit distribution of modern software, is being retained. Segments would be converted more towards their 64-bit extremely attenuated mode, rather than the more robust functionality they have right now. So most 32-bit software shouldn't be affected, unless you're getting really dirty with the segmentation features or trying to support something like Win16.
Which, ten years down the line, should free up a bit more opcode space when they decide to no longer support CS/DS/ES/SS segment prefixes. The space could be used for instructions that are the same in 32-bit and 64-bit modes.
The primary ("first byte") opcode space is quite busy in 32-bit mode. Intel and AMD figured out a way to "hide" extra instruction prefixes in illegal modes of certain 16/32-bit instructions. Intel did this with the VEX prefix and AMD did it with the XOP prefix. The instructions they chose were LES/LDS and POP. This is the reason why some of the bits in those multi-byte prefixes have to be inverted, btw.
Having 4 clean bytes available with no need for such silly games should prove useful in the future.
The 4 string I/O instructions they suggest removing will be useful immediately.
Removing all ring 3 I/O instructions (also part of the plan) will free up 8 further bytes in the primary opcode map for user space code. This could perhaps be used to provide shorter aliases of certain existing encodings -- so the instructions in question would still be available for kernels but applications would see better encoding density.
And then there is the 67h byte being freed as well due to removing address size overrides.
If they are sensible, they will also remove the BOUND instruction from 32-bit mode, but that is not in the current version of the plan.
So what are the benefits?
Free space in the primary opcode map => better encoding density... and they might not be doing this to give us shorter programs. They might be doing this so the parallel instruction decoders can handle more instructions per cycle.
Getting rid of 16-bit data will make the data flow engine in future CPUs simpler (no more partial register stalls).
They also get to remove lots and lots of special cases that undoubtedly cost engineering time and validation time.
The document explicitly says that segment prefixes are _ignored_.
> Removing all ring 3 I/O instructions (also part of the plan) will free up 8 further bytes in the primary opcode map for user space code.
It does not unless the same codes are removed in ring 0 as well. But if ring 0 IO commands are converted to multibyte, why ring 3 IO commands can't get the same?
> And then there is the 67h byte being freed as well due to removing address size overrides.
Again, not ignored - unsupported styles get GP instead.
> Free space in the primary opcode map => better encoding density...
Won't happen. More so, most encodings are shared with 32-bit mode and still developed this way.
> Getting rid of 16-bit data will make the data flow engine in future CPUs simpler (no more partial register stalls).
This is missed as well. 16-bit data are still handlable. Only addressing is disallowed.
> So what are the benefits?
Seems pretty none among what you listed:(
" 3.3 Removal of 16-Bit and 32-Bit Protected Mode
16-bit and 32-bit protected mode are not supported anymore and cannot be entered. The CPU always operates in long mode. The 32-bit submode of Intel64 (compatibility mode) still exists. ... "
This is somewhat beyond my knowledge of system programming, but my read is that it's just possible to support 32-bit guest kernels in the VM with some more emulation of the de-supported instructions. I don't know how many of those instructions already cause VM exits.
If we're talking firmware, the gigantic mess that used to be called ACPI, still leave users suffering every day for example.
If it's 2% of instruction cost, then kill it with fire. If it's raising the development cost by 1%, kill it with fire. If it's a fixed cost and isn't hurting anything, then meh.
...and eliminating 16-bit operations would be pure ideologically-driven stupidity, given how many classic RISCs with word-width-operations-only suffered greatly and the ones that survived (like ARM) evolved to add them back in.
To be honest, there are two legacy x86 features which I suspect do have a noticeable tax on the support in the execution units: the x87 FPU stuff (due to the 80-bit floating-point types, and a variety of hardware-implemented transcendental instructions that need to be supported), and the legacy addressing modes.
If they're going to still support virtualisation to run legacy code, which is what the article is implying, they'll still need the logic to deal with all the existing modes being used inside a VM.
...and those modes are honestly not hard to support. (Protected mode) segmentation is far simpler than paging, and realmode is even easier given that the original 8086 with its 29k transistors could implement it.
Or, they are adding something new that puts them over that budget, and cleaning up the conditional branching lets them stay under it.
> ...and those modes are honestly not hard to support.
I have to imagine that there's a decent amount of extra logic in the load/store units to use the current processor mode to work out memory semantics, and now I'm thinking about questions like "what would happen if one thread is in virtual 8086 mode while the hyperthread partner isn't". That's a noticeable burden in things like validation.
And even if the implementation of segmentation is simpler than paging, that's not the issue; the issue is that it's a different pathway, and potentially one that cuts deep into the execution unit. If it's the sort of thing where eliminating support for anything other than paging means getting an L1 cache hit back one cycle faster... then sign me up.
If we're talking firmware, the gigantic mess that used to be called ACPI, still leave users suffering every day for example.
And replace it with what? APM? I'm completely unaware of another power management abstraction that has even a fraction of the capabilities and proven track record. Sure if your apple or IBM you can build a machine with a custom tightly bound firmware/OS/hardware interface that works well. But, if you look at for example the Arm ecosystems attempts to use DT in the linux kernel and manage everything from a heavyweight kernel it is a fools errand. It has successfully been failing for fully 1/2 the lifetime of ACPI. Never mind no one in their right mind designs a system these days where the big cores try to manage their own power/perf curves, nor even modeling the system power/cooling without using a small low power core continuously monitoring and adjusting the hundreds of fine grained power domains on even smaller machines needed to achieve modern perf/power parity.So please explain.
I don't believe it's that simple (or obvious, since you seem to be mistaken).
> 16-bit and 32-bit protected mode are not supported anymore and cannot be entered. The CPU always operates in long mode. The 32-bit submode of Intel64 (compatibility mode) still exists.
There is SO MUCH silicon these days, honestly I think CPUs should start putting instruction set compatibility cores in, so you still dedicate .5% to a 16 bit compatibility core, 1% to the 32 bit mode, and maybe even an ARM mode.
Honestly, since all CPUs are basically fronted by microcode, why can't modern CPU microcode be compatible with multiple instruction sets? Or provide some sort of programmable microcode section so emulators aren't done in software?
While I'm bitching, what happened to actually running multiple OSes at once? I have like 16 cores, gigabytes of RAM, multiple display outputs ... why can't I actually run both Linux and Windows and OSX all at once without some bad virtualization container inside one overarching OS? Like, can't Intel manage that with chipset and CPU logic?
Intel (and AMD) are DESPERATE to figure out things to get people to use all their extra cores and silicon for. Well, how about making multiple OS on the same machine an actually pleasant experience? Hell, I would love to run a smartphone OS as well.
why can't I actually run both Linux and Windows and OSX all at once without some bad virtualization container inside one overarching OS? Like, can't Intel manage that with chipset and CPU logic?
Sounds like what you want is a LPAR or proper type 1 hypervisor, which despite everyone claiming to be type 1 none of them really are. And this is largely caused by the various device HW standards on the PC not natively supporting virtualization (or like nvidia charging extra for SRIOV functionality on their GPUs). So what happens is that all these hypervisors need something like VirtIo or hypervisor emulation of devices (ex e1000/ne2000, sb16, etc) to provide generic storage/networking/display/input devices on top of devices which don't themselves provide any virtualization support. Which in PC land is pretty much everything that doesn't support SRIOV. Which in turn means that a heavyweight OS +driver stack is required to manage the HW.So, you could build a PC based LPAR hypervisor. It would just be limited to the SRIOV capable devices, or plugging in piles of adapters, each one uniquely bound to a single partition. Ex: https://www.hitachi.com/rev/pdf/2012/r2012_02_104.pdf
Nvidia did something sort of similar to this, but couldn't get licensing to use x86. https://en.wikipedia.org/wiki/Project_Denver