x86: Approaching 40 and still going strong
newsroom.intel.com
newsroom.intel.com
At the time, Intel seemed to have pinned the 64-bit future on Itanium. Which, on paper, is a much better architecture. Of course instruction parallelism should be in the hands of the compiler -- it has so much richer information from the source-code so it can do a better job. The MMU design was very clever (self-mapped linear page tables, protection keys, TLB sharing, etc). Lots of cache, minimal core. Exposing register rotation and really interesting tracing tools. But AMD released x86/64, Intel had to copy it, and the death warrant on Itanium was signed so we never got to really explore those cool features. Just now we're getting things like more levels of page-tables and protection-key type things on x86.
Intel missed virtualisation not only in x86 (understandable given legacy) but on Itanium too, where arguably they should have thought about it during green-field development (it was closer but still had a few instructions that didn't trap properly for a hypervisor to work). The first versions of vmware -- doing binary translation of x86 code the fly, was really quite amazing at the time. Virtualisation happened despite x86, not in any way because of it. And you can pretty much trace the entire "cloud" back to that. VMX eventually got bolted on.
UEFI is ... interesting. It's better than a BIOS, but would anyone hold that up as a paragon of innovation?
I'm not so familiar with the SIMD bits of the chips, but I wouldn't be surprised if there's similar stories in that area of development.
Low power seems lost to ARM and friends. Now, with the "cloud", low power is as much a problem in the data centre as your phone. So it will be interesting to see how that plays out (and it's starting to).
x86's persistence is certainly remarkable, that's for sure
It sure would be interesting to see something else challenge x86, but since ARM/Power/RISC-V are all within the same architectural tradition I doubt they will provide it. SIMD as used by GPUs is useful only when you have very structured parallelism like arithmetic on large data sets.
I think if we're going to see a challenger worth noting they'll look like they're crazy with the hard uphill battle and novel tech they will have to produce.
True of course, but Itanium did think of this and had an "abstracted" VLIW where you worked with generic instruction bundles which the hardware was then more free to pull apart in different ways underneath. I do wonder, if we'd all taken a different path, if we'd be in more of a JIT-compiled type world. who knows?
They say that JITs will have to link to their proprietary system library in order to work, which seems like it could be devastating for uptake.
Vector Microprocessors
Can you elaborate on what you mean by "architectural tradition"? I am not familiar with this term. What is the commonality of all those different architectures?
ARM/Power/MIPS/RISC-V are all examples of a reduced instruction set architecture, RISC.
Basically, despite how sophisticated the CPU of these machines becomes it is still trying to make the appearance of executing instructions in sequence.
Current x86 processors can execute many instructions in parallel under the correct conditions leading to Instructions Per Clock (IPC) greater much than one. However typical IPC hovers around 1 and some change, unless a lot of optimization effort is invested.
In order execution of an instruction sequence (or the appearance to the user of it) is the bottleneck that has led to single threaded CPU speed increases to all but disappear now that CPU frequency scaling has more or less halted.
I think that's all the OP was getting at.
Rather the DEC VAX and perhaps MC68k are poster childs of CISC. For a discussion of RISC vs. CISC see
> http://userpages.umbc.edu/~vijay/mashey.on.risc.html
or its extended version
> http://yarchive.net/comp/risc_definition.html
and look at the tables there. You will see that among the processors on the "CISC side", the Intel i486 is nearer to RISC than many other CISC architectures.
With it being borderline impossible to write good VLIW compilers, most of the instruction word was full of NOPs, which meant terrible code density, which meant I$ full of NOPs. A recipe for hot, slow computers which are depressing to program for.
Wow, might you have any links regarding Knuth commenting on VLIW optimized compilers? I would love to read more about this.
What was Intel's view on the viability of it? It seems like quite a shortcoming.
https://en.wikipedia.org/wiki/Channel_I/O
Even some embedded SoC's are doing this now where there's a strong core and a cheap core w/ main function and I/O split among them.
People called the IO processor "the real OS" because they had never seen dedicated I/O processors before.
I didn't know that. That's funny as it's the confusion I'd expect where the old wisdom of I/O processors in non-server stuff was lost for a few generations. Then, the new one sees them to wonder if it's an extra core for apps, the OS, or whatever.
But it's over my head. I lack knowledge and equipment for measuring RF leakage among SoCs. So it goes. Maybe the Qubes team will do it :)
Our discussions about physical and diode security on Schneier's blog were a small part of inspiration for Tinfoil Chat. Markus Ottela then showed up to discuss the design. He wisely incorporated each piece of feedback like switching from OTP to a polycipher and covert-channel mitigation at protocol level. Loved reading his regular updates he posted of the things he added. The design, not implementation, is one of the only things I'll call NSA proof for confidentiality and security but not availability. By NSA proof, I mean without TAO getting involved with really hands-on stuff.
One drawback was he didn't know system programming. So, reference implementation is in Python. People that know statically-verifiable C, Ada/SPARK, or Rust need to implement it on a minimal TCB. I'd start with a cut down OpenBSD first just because the NSA will go for 0-days in the lower layers and they have less. Make the implementation as portable on MIPS and ARM as possible so the hardware can be interchanged easily. Trimmed Linux next if drivers absolutely need it w/ Poly2 Architecture-style deletion of unnecessary code. If money comes in, implement it on a secure CPU such as CHERI w/ CHERIBSD.
That was my plan when I was talking to him. Lots of potential with TFC after it gets a proper review of protocol and implementation.
[1] https://pdfs.semanticscholar.org/5698/09af0fcbe5a42371cea8d3...
In 1995, the Pentium Pro featured out-of-order and speculative execution[1]. This has two main advantages over the VLIW architecture (Itanium in particular):
1) The processor can use runtime information (such as past history of branches) to speculatively execute hundreds of instructions before they are needed.
2) When a stall happens due to a cache miss, in some cases the CPU can continue to execute instructions while waiting for the stall to end. This isn't possible with VLIW because the compiler can't predict cache hits/misses, this information is only available at runtime.
These two advantages combined made Itanium much slower than the then available x86 processors.
[1]: Note that Intel didn't invent these ideas. OoO was already known in the 1970s, although it was applied to mainframe supercomputers. I'm just using Pentium Pro as an example of a similar microprocessor.
This is not true. Yes, the compiler has a lot of useful information, but it does not have access to runtime information, which the CPU does.
A modern superscalar CPU with out-of-order execution can execute instructions speculatively hundreds of cycles before they are even used.
Instructions after conditional branches can be executed speculatively using branch prediction and loads can be executed before stores (even though they might access the same address!). If the processor later detects that the load was dependent on the store, it can "undo" that operation, without any of these effects ever being visible to the running program.
I've read about this promise of JIT optimization being able to generate optimal code on hot paths, but in practice I've never seen it work out that way. Good AOT compilers exist today and consistently generate optimized machine code.
The security stuff was explored pretty well by Secure64. I've always given them credit for at least trying to build bottom-up on it with above-average results in pentesting.
https://www-ssl.intel.com/content/dam/www/public/us/en/docum...
So it failed, because Intel wasn't the only one doing x86 processors, otherwise they would have been able to push it down our throats one way or the other.
- initialize the RAM and other essential subsystems
- initialize whatever is needed to access mass storage devices
- copy X bytes from a specific location on a specified mass storage device into RAM
- set the instruction pointer on the copied data
Everything else is the OS job. Even the BIOS does too much.
Was I foolish ?
It's difficult to take this threatening language seriously when it's exceedingly obvious that Intel is very suddenly (and very rightfully) terrified of losing their monopoly position. It's easier than ever to recompile and switch to ARM, and people are tired of paying the Intel tax.
They should have seen this coming well before Otellini retired. Lawsuits aren't going to save them here, and Rodgers comes out looking like an asshole for trying.
I'd say watch out for Intellectual Ventures on the Transmeta patents. There might be some still in effect that could mess with dynamic compilation or who knows what if broadly written.
[1] https://www.law360.com/delaware/articles/16290/intel-files-s...
Even LLVM bitcode might eventually be cleaned up to make it properly architecture independent. There was a talk about it at a past LLVM days conference.
Indeed, more and more software runs on CPython or Java, so some developers might not even notice!
>Launched on June 8, 1978, the 8086 powered the first IBM Personal Computer and literally changed the world.
It was the 8088 that powered the IBM PC. (Introduced on July 1, 1979, according to Wikipedia.) Very similar but with an 8-bit bus instead of the 8086’s 16-bit.
Too bad I can't remember the source.
Article is written by the Intel General Counsel...
> Only time will tell if new attempts to emulate Intel’s x86 ISA will meet a different fate. Intel welcomes lawful competition... However, we do not welcome unlawful infringement of our patents, and we fully expect other companies to continue to respect Intel’s intellectual property rights.
ARM uses a weak memory model. x86 has a strong one. Also code containing optimized SIMD code (e.g. via compiler intrinsics) is a lot more complicated to port.
It might matter for lock-free data structures and other potentially hardware sensitive code, but it doesn't matter for 99% of code.
I have actually coded for 320 GFLOPS on a Broadwell, so I know how this works, and I can tell you it's a lot easier to get vastly higher performance on a GPU.
You have to try it for a while to get used to it and sort of understand how to do it since it's slightly different… But it works SURPRISINGLY well. And the amazing size/weight/battery life of an iPad as well as the cost relative to most Apple laptops makes it attractive.
Since I got a good iPad I really do rarely touch my Mac at home. Mostly for games that aren't on consoles.
I'm not sure why people are so hostile to iPads. For what a TON of people use computers for they are easily past 'good enough'.
Also, phablets have eaten into some of the demand for them.
Especially with the new iOS 11 changes. I mean, the only thing keeping me from using an iPad for everyday use is I can't do any code in it.
Aside from that, everybody is already using their phones more than their desktops, and even when they do use their desktops, it's to access the slightly uglier version of Youtube on Chrome...
Probable future will be entreprise and schools using Chromiums/Windows 10 (very closed platform, but that's not a bad thing), and iPad/iPhone (or not iOS) for every other use.
Here I say: "[citation needed]". It surely is rather the complete opposite of how I use the devices (phone: about 0 hours (I don't own a bugging device with telephony functions), PC - hell a lot).
Edit: Seems odd to use Transmeta as a comparison if they are talking about software os level emulation. Wasn't Transmeta's emulation all built into the chip, no os level support/software needed?
Oops, they are probably referring to Microsoft/Qualcomm's announced project here: https://www.extremetech.com/computing/249292-microsoft-decla...
If they were to switch processors and it was only (say) 10% faster, the hit from the transition layer may make it a hard sell.
Apple has been knocking it out of the park with the A-series processors though, maybe they could make a big enough difference that it would work out again.
Applications where CPU time was mostly drawing on screen (e.g. text editors) worked a lot better than those where most CPU time was in app code. Today the GPU and graphics code is an even larger amount of the processing.
So for a lot of apps it will not be a big problem if the emulation is a bit slow.
Both Mac processor architecture changes so far have included emulation or binary translation layers for this purpose.
68k -> PPC: https://en.wikipedia.org/wiki/Mac_68k_emulator
PPC -> x86: https://en.wikipedia.org/wiki/Rosetta_(software)
I can't imagine Intel going to war against Microsoft, both sides have too much to lose if they start seriously fighting. It's not called the Wintel monopoly for nothing.
But if Apple were to stop buying Intel chips, and start using their own chips in their Macs, Intel could become very litigious.
Did any publicly evidence appear officially or accidentally on the internet that Apple is working on such a project (for Microsoft the situation is known)? I am just a kind of person who prefers strongly to look at the evidence instead of rumors, speculations and hopes.
In the next version they may run 64-bit clean and use emulation in some sort of Rosetta style environment for "legacy" 32-bit software.
Or maybe you could make ARM chips and run a Rosetta like thing to run the 64-bit code, but since 64-bit x86 is so much cleaner it's not much of a problem compared to the whole 32-bit mess.
This is all pure speculation. I thought the way they phrased things with a little odd. It may simply be that they don't intend to update any libraries in 32 bit mode so stuff will just stop working as libraries change. It may be everything's going to start going through some layer that thinks 32-bit calls into 64-bit calls thus possibly reducing performance.
I don't know… I just feel like that statement meant something interesting and I am coming up with fun guesses.
I would not call x86-64 "cleaner" than x86-32. In many senses it is much more messy than x86-32.
They may have been referring to the fact that 64-bit runs in 65-but only, no choice to drop into/out of 16-bit, etc.
Even in 64 bit mode it is possible to drop into 16 bit (protected) mode (code) if one wants to. What does not work is dropping into Virtual 8086 mode.
GNU/Linux never supported 16 bit code. For Windows it would be possible in principle (and it even would IMHO sense to support this feature on 64 bit versions of Windows). On
> https://news.ycombinator.com/item?id=14246521
I wrote something about this topic. TLDR: 64 bit Windows uses 32 bit handles, which will cause problems with 16 bit applications.
But there are also people on HN who stated the opinion that while this makes implementing support for 16 bit applications on 64 bit Windows harder for Microsoft, it would have been far from impossible, thus this technical reason is a mere convenient pretense not to implement this feature.
It's borderline a non-starter.
A couple of demos of Microsoft's x86 emulation. It seems to work well:
For those paying attention to the BUILD talks, those developers will be gently dragged into UWP thanks to Desktop Bridge, regardless how they feel about it.
New Office version for Windows 10 will be UWP only.
CP/M, MS-DOS and Win16 compatibility were eventually removed.
I don't think about 32-bit Windows since I left XP.
Chromebooks are having somewhat of the opposite battle. Intel's trying to emulate arm to run android apps on chromebooks. The reports I've seen is that intel's not competing so well. Many popular android apps run better on an arm chromebook than an intel chromebook.
In both cases I think the consumer wins, competition is good, and I'd consider a new chromebook just for the hardware and then I'd install some flavor of linux on it.
Not only that, they are striking fast at high performance. Apples top end iPad chips are close to/better than recent laptop chips at this point. And clearly Apple can do that at scale.
The x86 world is having a harder time getting more performance out of each revision, and there's enough going on on the ARM side that we could be within reach of something interesting happening. This isn't like Transmeta or the G series chips from Motorola/IBM.
On UWP applications, if written in .NET Native, the MSIL gets uploaded to the store and is AOT compiled to native code for all existing supported architectures.
Apple has LLVM bitcode, which while still having a few leaky abstractions, there are plans among LLVM devs to eventually have a more clean architecture independent version of it.
When using a .NET based engine like Unity, the processor specific bits are part of the engine itself, so the majority of developers usually don't t write native plugins themselves.
I find it hard to frame that as "ARM can't compete with Intel", especially since PCs aren't a lucrative market anymore. You want to be in mobile, IoT, etc., and Intel threw in the towel there a while ago.
And until that happens my original point was: business software, operating systems, etc., have been written and / or optimized for the x86 platform so much so and said software is so pervasive that leaving x86 will take a revolution or a big player to push it, again maybe Apple.
This is great for consumers, by the way: we're long overdue for healthy competition in the CPU market. ARM is not an open architecture, so it's not ideal, but at least the various licensees compete with each other, which is an improvement over what we have with Intel.
> And until that happens my original point was: business software, operating systems, etc., have been written and / or optimized for the x86 platform so much so and said software is so pervasive that leaving x86 will take a revolution or a big player to push, again maybe Apple.
The operating systems that matter are all ported to ARM. And perhaps the most important consumer OS today--Android--has virtually no x86 penetration. In 2017, looking at the overall market, ARM has the advantage in terms of software compatibility, not x86.
Using such mainframe model, makes the actual processor only relevant to the OS vendors or the developers using "down to the metal" toolchains.
Which is something that obviously weakens Intel's position.
As long as the law is in their favor...
The whole thing is a fluff piece that is best ignored.
You do realize that your ISA is literally one of the worst things about your processors, right Intel?
A 64-bit x86 computer with UEFI and wihout a CSM, for example, has no way to boot DOS or even a 32-bit OS (which could then use v86) to support the really old legacy. You'd have to resort to using a hypervisor, but considering the practical difference in performance, you could do the dumbest possible emulation and still run those 16-bit tasks just fine.
But that's just details. For the massive scale-out farms that power today's world with solutions based on OSS components, running proprietary legacy code just doesn't matter, so the backwards compatibility is irrelevant.
Like hell. Just tell all Windows users they should just replace all their software on Windows/x86 with an ARM or RISC-V box. The software will not work or will run with crap performance. They'll ditch the alternative for Intel/AMD x86. The End.
The only time they move is for something new that doesn't depend on their legacy software. Also, for stuff where they can transition unlike many enterprises locked into Windows or other x86 tech.
Or it mostly stays the same for those locked in with the current rate being billions a year for the companies that locked them in. Hard to tell how long it could be before they can get off. The IBM/COBOL crowd is still locked in after 30-40 years.
But overall, the world has changed quite a bit. We're suddenly not in a world ruled entirely by DB2, MSSQL and Oracle, running on AIX, Windows and Solaris. The software stack has opened up, and is no longer bound to any specific platform. The few big software names don't get to decide what kit to run them all...we're approaching a world where ISA won't matter. You'll buy hardware because it fits your performance and TCO targets, not because you're locked in.
What sucks for Intel is that they've failed to identify mobile devices as the "next PC", in the sense of the disruptive potential...the PC killed boutique workstations and servers and it killed them from the bottom. Along with IoT, that's two huge opportunities missed. Itanium was a pincer attempt, but no one targets the top and succeeds... you have to start at the bottom (this is btw why Power needs to scale and price down if it hopes to grow beyond its niche).
They might be in terms of getting locked-in to a proprietary vendor using proprietary language, extensions, libraries, and/or OS's that are hard to impossible to get off of. That's exactly what Wintel has been doing with enterprise software. Microsoft is still pulling in billions despite people saying other stuff would kill them for some time now. It's a combination of their market hold, lock-in, and patent suits. They're pulling $1+ billion from Android with patents despite not contributing jack to it. Oracle and SAP are still bringing in billions due to market hold and lock-in. This is a steady thing.
"What sucks for Intel is that they've failed to identify mobile devices as the "next PC", in the sense of the disruptive potential...the PC killed boutique workstations and servers and it killed them from the bottom."
That's true. That it didn't need compatibility with Intel is exactly how it killed the need for their ISA in those sectors.
"you have to start at the bottom (this is btw why Power needs to scale and price down if it hopes to grow beyond its niche)."
I agree. Great as it is, it's way too expensive. The POWER/PPC model was doing a lot better with Apple on it selling them at reasonable price. The ecosystem was better anyway with all kinds of software made for it. It had a better ROI for buyers. They need to do that again for POWER(number here) even scaling down the capabilities as the price goes down if necessary. I'd say their accelerator interconnect that can bring in stuff like FPGA's would be a winning differentiator but Intel outsmarted IBM again with the Altera acquisition. I'm expecting great things in offloading out of that if they haven't shown up already.
(The first Datapoint 2200 prototypes were shipped in April 1970, official introduction in Nov 1970, Pat. 224,415 filed Nov. 27, 1970. The Texas Instruments microchip implementation was announced in June 1971 and delivered during the summer. The Intel version was filed for patent in Aug 1971 and the chip presented to Datapoint/CTC in fall 1971. Intel's first assignments to the 1201 chip, which eventually became the 8008, were apparently in spring 1970 with the project temporary stopped in July. [Source: Wood, Lamont, Datapoint; Hugo House Publishers, Ltd.; Austin, TX, 2012] — So for the architecture in general it's 1970, for the microprocessor implementation 1971.)
The Intel 8080 used a different instruction set (google it). The basic instruction set was introduced with the Intel 8086 and Intel 8088 and was not binary compatible with the Intel 8080. Though Intel claimed that 8080 software could easily be ported on a source level by search and replace in the assembly code.
The Zilog Z80 extended the instruction set of the Intel 8080 and defined a new assembly language for it. I heard one of the reason that Zilog invented the new assembly language was also because of intellectual property - nevertheless one can rather objectively say that Zilog's assembly language was better than Intel's. For the assembly language for the Intel 8088/8086 Intel took lots of inspirations from Zilog (google and judge for yourself - in particular compare 8080 assembly code (Intel's assembly language), Z80 assembly code and 8088/8086 assembly code).
F00F bug? ;-)
> F00F bug? ;-)
For those who are out of the loop: https://en.wikipedia.org/wiki/Pentium_F00F_bug
Though since your parent was talking about the Intel 8086/8088 and 0x0F: There is an incompatibility 8086/8088 and newer x86 processors: On Intel 8086/8088 the opcode 0x0F was "pop cs". From 80186/80188 on, 0x0F is instead an opcode expansion prefix (cf. https://stackoverflow.com/a/12264515/497193).
Thanks ;-)
It's the old FUD technique, make it appear that there's some legal doubt about new technologies to slow the uptake.
It's basically equivalent if there's no alternatives in microarchitectural details that achieve the same thing without violating a patent. There's a crazy amount of patents on that stuff. So, whoever builds it needs to sell enough to pay the lawyer fees on top of the catch up fees.
In other words… They're about to become a PC gaming has had for 30 years.
And that makes them LESS likely to switch to a new chip instruction set.