Intel Explores Transition to 64-Bit-Only X86S Architecture
tomshardware.com
tomshardware.com
Intel has been burned in the past, though. The Pentium Pro (1995), the first good superscalar microprocessor, was a 32-bit machine that could also run 16-bit code. Intel thought 16-bit was on the way out, so the 16-bit hardware was not fast. Unfortunately for Intel, Microsoft Windows users were still running too much 16-bit code. It ran Windows NT fine; Windows 95, not so much. As a result, the Pentium Pro did not sell well. Intel had to come out with the Pentium II (1997), which was similar to the Pentium Pro but had better 16-bit performance.
I thought the i960CA was good.
Wouldn't that depend on what they do with resulting available space?
If that 5% all became dedicated L1 cache instead, that sounds like it'd have a speed impact. :)
Just allowing booting directly in 64-bit mode and making 64-bit SMM mode would be enough to avoid boot software stack problems. Riddance of "legacy mode" doesn't require all it.
The proposal [0] still keeps support for 32-bit user space – or at least as much of it as is commonly used by mainstream operating systems (Windows and Linux). In this case, they are removing obscure features which aren't used by current versions of mainstream operating systems, so the only real impact is going to be (1) people running non-mainstream operating systems, (2) people virtualising legacy versions of mainstream operating systems, (3) apps that try to directly access IO ports from user space – which is a supported feature (albeit very rarely used) on Linux, not officially supported by Microsoft on Windows (although third party kernel drivers have been used to implement it)
Virtualisation of legacy/non-mainstream OSes which use these features will be possible through software emulation, although obviously that has a performance cost.
I wonder what the impact will be on OpenVMS for x86-64?
Is Intel going to drop these legacy features across their entire product line? Or maybe their low-end SKUs will drop them, but they'll keep some high-end SKU which retains them? Although one wonders how long that will last.
Depends on whether you view MMX, and SSE as baggage. They were OK at the time, but I guess we could do a lot better today.
Having said that I still wish they open up X86s. Or may be an even new 86 breaks most of the backward compatibility ( pure 64 bit mode ) but freely open to implement. Meaning if you want backward compatibility you will need to pay AMD or Intel for it.
I wonder how this will affect quirks such as A20, IRQ remapping, etc. All oversights and mistakes made by Intel over the years - none of which really hurt development or performance these days per se, but are definitely not fun to work with either.
Something like this is the only thing that would save the x86 architecture before ARM inevitably took over.
Engineering is not about achieving perfection. It's about organizing compromises to create a product that provides good utility at a fair price. These were not oversights or mistakes they were intentional design decisions.
Likewise.. you can see segmented memory as "nonsense" and a "waste of development time" but Intel clearly didn't fail at anything. They've been one of the largest and most successful chip manufacturers for decades, because their products provided real utility at affordable consumer prices and software developers gladly put up with the technological choices to be a part of that giant market.
(I'm still shocked how they've begun losing to AMD. Having all of the money, dirty benchmarking tricks, and "let's just hire all of the people" haven't worked out, incredibly.)
Do not forget all of the anti-competitive AMD nonsense over the years. Sabotaging AMD performance on benchmarking is my stand-out, but there were also the payouts to Dell and others to purchase exclusively Intel chips.
Probably the 2 biggest reasons for this are that only 2 companies are allowed to "touch" x86, and that the second company was easily undermined by nefarious business practices. This resulted in a lot of things that were not "oversights or mistakes" but rather for the longest time lax approaches under a distinct lack of pressure.
It's hard to say how the x86 world would have looked like today if someone like Nvidia had a license.
There's some alternate universe out there where x86 died because Nvidia was charging $1200 for a mid tier CPU so arm was able to creep up even faster :/.
Eh, this feels like a rounding error in ISA desirability compared to pointer authentication, or something like CHERI. Sure, it'd be _nice_ to have the legacy stuff cleaned up, but it's not costing much human effort to deal with compared to the effort spent securing native code.
I could imagine a more user-noticeable effect in VM cold-start times, perhaps, though I'd think physical hardware boot times would still be dominated by other things (e.g. actually loading the kernel into memory); the minimal implementation of transitioning from real mode to long mode is a few hundred instructions.
See all the revolutionary ISAs from the 80s, 90s and 00s that did away with all that legacy crap. Where are they now? I think the only one still working is POWER.
To be fair, both were always "optional" features; but yes, there are thousands of devices that shipped with them (one of the biggest being the GBA, with the majority of games heavily relying on THUMB).
A20 was removed in Haswell (released pretty much exactly ten years ago).
Well, the A20 gate was actually IBM's creation, in order to make the PC-AT system compatible with running real-mode DOS programs.
Intel simply, later, incorporated it into the CPU because it was already part of the PC architecture (and had been for years) at that time.
Now, one could argue Intel had no business incorporating it into the x86 ISA, but it wasn't ever their oversight, they were just reacting to the reality of the systems in which the vast majority of their CPU's were used.
So, feature or bug? Heh.
People tend to remember mostly how difficult it was to program with segmentation in 286's 16-bit mode and earlier. But in 32-bit mode, user-space code didn't need to bother if the OS didn't use it.
I've read several papers lately about schemes with compiler/OS co-design for securing the return pointer on the stack, function pointers and other sensitive variables from hacking attacks. This problem was already mostly solved on x86-32 where the stack can be in its own segment.
But to get achieve this in 64-bit mode (and on ARM and RISC-V), researchers have resorted to various tricks such as randomising addresses regularly [SafeHidden], switching access to the (safe) stack on and off with Intel MPK, using the Intel shadow-stack (in CET) for storing variables [CETIS], and even running user code in privileged mode [Seimi], [InversOS](ARM) ...
Some of these these schemes are tricks using the CPU in ways it wasn't intended, and therefore perhaps broken on future CPUs. Several are only available on newer Intel processors with certain extensions, and most have considerable run-time overhead. Therefore, you will never see any like it in a mainstream OS. But the x86-32's safe stack would have been, as "Safe Stack" schemes relying on randomisation already got widespread adoption.
[SafeHidden] https://www.semanticscholar.org/paper/SafeHidden%3A-An-Effic...
[CETIS] https://www.semanticscholar.org/paper/CETIS%3A-Retrofitting-...
[Seimi] https://www.semanticscholar.org/paper/SEIMI%3A-Efficient-and...
[InversOS] https://www.semanticscholar.org/paper/InversOS%3A-Efficient-...
Burroughs B5000 had segments. Roger Schell said it was added to Intel by a Burroughs guy when they requested hardware-enforced security to protect memory. They wanted to use it in upcoming OS's like GEMSOS. Rings came from SCOMP which also had an IOMMU in the 1980's. Descripters evolved into capability-based computer systems. Intel tried with Intel i432 APX and i960MX (good one). They lost billions. CHERI with RISC-V is the best of that lineage right now.
https://www.cse.psu.edu/~trj1/cse443-s12/docs/ch6.pdf
https://web.archive.org/web/20220701015547/https://www.mrhec...
https://homes.cs.washington.edu/~levy/capabook/
(Note: Aesec might still sell GEMSOS today.)
More recently, schemes like Native Client and Code Pointer Integrity used segments. The reviewer who broke the software versions of Code Pointer Integrity couldn't bypass the segment-enforced version. There was probably a lot of unexplored territory in combining modern, memory protection with segments. I'd like the alternatives to have as much assurance as designs on segments gave us. They're getting hacked more often, though.
On the contrary, it will make it worse. People choose x86 and the PC for the backwards compatibility, and Intel throwing that away is not going to make them any better than ARM.
- Foolhardy snatch 1: iAPX 432. - Degradation in first half of 1990s, concluded with FDIV bug, and marvellous savior "ASCI Red" which paid all OoO development (PPro and following). - Foolhardy snatch 2: RAMBUS + Itanium. Nearly collapse with P4 and miracle #2 with Israel team, Core, and stealing AMD64.
I can't imagine that a person adopted EPIC-based Itanium had had a real engineering competence. (Not sure for iAPX 432, this is a complicated issue.)
On a lower layes, as ISA details, every Intel move looks like using prediction less than for 1 move forward! Maybe 1/2 move sounds more veracious. Loads of examples. And, this also seems a guy that just remembered he was an "engineer" 30-40 years ago is now a decision maker.
Nor does it get in the way of software - as soon as you've switched to 64 bit mode, you can ignore all legacy stuff.
Thats the right way to maintain legacy compatibility - make sure it is all contained and doesn't get in anyone's way, and you can leave it there forever.
2. The nefarious answer: maybe they found a way in their agreements with AMD which would exclude this X86S somehow.
Segmentation also has its uses, and it's a shame that they are steadily removing it. Having a way to set up a 1:1 memory range without even touching the page tables has always been faster than messing with pages. Indeed, it's always been more performance to either have a fully static page table or just disable it completely and run in 32-bit mode. For certain tasks, it's still the fastest.
But these are all niche uses. These days, who thinks about anything other than Windows and Linux use cases?
To me though, AMD64 made most problems go away, and the architecture is just fine. Perhaps they want to do away with many of the now rarely used instructions?
That, I assume, is the point. If it isn't used by the vast majority why spend the silicon keeping it present?
> all the 16/32 bit mode stuff is all 'emulated' with microcode
That microcode still needs to be maintained, tested, & verified, with each chip iteration, so it is more than just a tiny bit of silicon that could be repurposed.
> and you can leave it there forever
If it is there it will be expected to work (otherwise why keep it?) and they need to keep verifying that containment. I'm not much of a hardware guy myself so maybe the risk profile is different - but I've tripped over, or been mugged by, so-called dead code that ends up having “interesting” side effects, enough times for me to not want a bunch of it in my CPU's design if it is avoidable.
Removing it has costs somewhere of course, both in Intel's design work and potential compatibility issues in minority use cases. But deciding to leave it in has costs elsewhere, so the options & risks need to be looked into in either case.
Is that true though? 32bit Windows 2003 Enterprise or Datacenter editions were perfectly capable of accessing 64GB or RAM [0]. Even so, that difference from the usual 4GB and 3.xGB was any kind of memory usage for integrated GPU or virtual address space mapping for different hardware.
[0] https://learn.microsoft.com/en-us/windows/win32/memory/memor...
P.S. I think the "3.2" figure irked me more than the usual "32bit OSes can't use more than 4GB". That's not a real limit and never was. Even people with no IT knowledge would have noticed it varies quite a bit anyway. I guess I expected a little more from a tech reporter.
Also, I'm not perfect, I have an attachment to the physical realm, and I desire not simply to run software, but to run it on hardware, whatever that means. Oftentimes, I desire simply to run hardware for its own sake, because it brings me pleasure.
It also looks really weird compared to other ISAs as System/Z or SPARC which really benefit of not using page translation in kernel mode. I'd expect alleviation of switching paging on and off...
2. Looks it's a good moment to relax memory ordering from TSO, with easy control using Flags. (Using it will gradually soar in years, yep.)
Stripping out x87 and the associated 80 bit registers, weird modes, etc would seem right in line with this effort. If someone really needs x87 support, trap those instructions and emulate in kernel. Nothing with truly high performance requirements will be using them. Intel could even provide the code.
Presumably a few less transistors, but a small % I'd guess.
Will it make anything run faster? If so, what?
If C++ would have cleaned up legacy, Rust wouldn't need to exist for ex.
Java has long had a policy of never breaking BC, but few years ago they started marking packages deprecated for removal. They're very slow and systematic about it, as they should be, but the point is, eventually those packages will be removed. And that's good. It means Java has a future.
Rust is a famous example; there's also Carbon[1] and cppfront[2].
Even so, this may be a controversial opinion, but with C++20, it has become fairly straightforward to program relatively safely and avoid some old footguns.
I recently wrote a server-client framework from scratch and used protobuf for messaging. The server is heavily multi-threaded with a lot of complex asynchronous rules, and there are no issues. If anything, I would hope for more concurrency primitives, as the C++ standard library is very high quality, and I would like concurrent hash maps and such.
Maybe that works for some new codebases.
When you need to integrate with other code, you can't necessarily understand the code semantics without knowing a lot more of the spec.
Also, I challenge any normal programmer to remember all of the "good" C++ rules regarding overload resolution in an arbitrary context.
I mostly design and create new products but on some occasions I've had situations like this. There was not a single time that quick Google search and now ChatGPT did not shed the light.
>"Also, I challenge any normal programmer to remember all of the "good" C++ rules regarding overload resolution in an arbitrary context."
Knowing / remembering every nook and cranny does not make for a good programmer (except language experts who write standard libraries and design / implement language itself of course). There are way more important things to consider.
So bottom line is that no matter which way you try to do the userspace emulation, some parts are going to be ring0, and thus the answer to your question is indeed yes.
To make your life easier? The customers may have totally different opinion.
I am wondering how these numbers stick together. I especially wonder why those extra 0.8 GB are unaccessible, when did that happen and who's the thief
In researching this I learned that some 32bit hardware could utilize 128GB of ram using Physical Address Extensions. Each process is still limited to 4GB though.
3.2GB is an example figure - it could be 3.5GB, 3.6GB... - but reflects that RAM can't be mapped at all of lower 4GB. Memory controllers typically remap part of it to a range upper than 4GB boundary, so 32-bit PAE-capable or 64-bit OS can still use it.
OTOH using 64-bit OS is recommended with RAM >=1.5GB because if OS can't map all RAM _twice_ and all device memory to virtual address space, tricks with memory banking (frequent page remapping) are getting needed to deal with data move.
I feel like most retro games work fine in a 64 bit OS, but maybe you can chime in if that's true or not?
EDIT: https://qr.ae/pyvelA
https://stackoverflow.com/questions/5806589/why-does-intel-h...