In the 80s and 90s academia and high-end computing was taken by a RISC illusion. RISC architectures were simple and clean and easy to implement at higher performance than any CISC processor of the time, so people assumed that of course RISC would be the feature, right? Well, no. CISC, or, to be more truthful, x86 was a mass-market item with competing manufacturers with deep pockets (like Intel and AMD).
In the mid 90s both Intel and AMD broke RISCs neck (though it was not immediately clear perhaps what a groundbreaking innovation this was): They introduced processors that translated x86 instructions to other instructions (µops) that were then scheduled and executed. Turns out, this is pretty efficient for a number of reasons (1- you get existing software 2- x86 code can be really dense; RISC had instruction bloat issues 3- decoding insns is only a very small part of what a CPU does and does only consume a minority in transistors 4- you can change the µops, they are an implementation detail, to best suit the exact core at hand; something impossible for RISC). This allowed x86 all benefits of what RISC previously had for itself.
By ~2000 most RISC architectures were dead (Alpha, SGI, ...) with only POWER and SPARC surviving in their niche markets.
In summary: The importance of the ISA was vastly overestimated by researchers, who assumed that this was a blocking issue to make x86 go faster. In the end, x86 won, because x86 had the bigger market with more money in it.
You forgot ARM, which is a major omission.
> In the end, x86 won, because x86 had the bigger market with more money in it.
It's not the end yet. :) x86's dominance in the future is hardly assured, and in fact if you measure by total number of microprocessors sold, x86 is a small niche--300-400 million sales per year compared to 15 billion ARM chips sold.
> Why are x86/64 processors used for most high performance applications then?
Also ARM, the architecture where most cores have between three and four distinct frontends, hardly qualifies as R ISC.
More importantly, ARM has a way to massively reduce its decoding complexity, via AArch64. AArch64 is a separate RISC instruction set, with room explicitly left in the design for dropping 32-bit ARM in the future (32-bit support is optional, though in practice today everyone includes it). ARM is effectively in the process of making a transition to a more RISC-like ISA, unlike Intel which has no realistic way to get there.
If you write down the x86 instruction encoding in octal, it makes a lot more sense:
There are probably more PowerPC chips sold via laser printers and such than x86 chips. Conversely, 8051 derivates probably outnumber all of these. So what. Meaningless argument.
Because (1) you can overcome those issues by throwing enormous manpower at them; (2) the business success of Windows, coupled with Intel's patent portfolio to keep out competitors, historically generated enough revenue to fund that manpower.
(There are signs that this may not sustain itself forever, with Intel's failure in mobile and the performance disappointment of Kaby Lake, but regardless of how well the strategy will continue to perform in the future, it worked for decades.)
Modern Intel cpus are actually risc under the hood. https://en.m.wikipedia.org/wiki/Microcode
First: Modern ARM chips (i.e. what you will typically find in a mobile phone) support three instruction sets: A32, T32 (Thumb-2 on modern ARM chips) and A64 (if you don't know what they mean, google it up). So it hardly makes sense to talk of "ARM assembly" if you don't specify which instruction set you mean. I know there is UAL, which in some sense unifies the instructions of A32 and T32 on assembly level.
There are very central differences that you will observe when looking at the disassembly of all ARM instruction sets vs. x86-32 or x86-64.
- Typical x86 ALU instructions have 2 operands. Typical ARM ALU instruction will have 3. This is very typical for CISC vs. RISC instruction sets.
Side note: A compiler often does register allocation by graph coloring. If 3-operand instructions are used this is an NP-hard problem. On the other hand if 2-operand instructions are used, we obtain a chordal graph, for which the graph coloring problem can be solved in polynomial time: https://en.wikipedia.org/w/index.php?title=Chordal_graph&old...
- x86 ALU instructions also accept a memory both as source and destination (side note: if we add a register to the value at a memory address, on modern x86 cores this is mapped to 3 mu ops: load value, add, storeobly value). On the other hand (typical for RISC) ARM is a load-store architecture. So ALU instructions don't accept memory address contents as operands - instead one has to load the content of the address into a register first, do the ALU instruction and use a store instruction to store the result in memory.
Side note: On x86 "{alu_instruction} r/m, r" (e.g. add [edx], eax) is atomic with respect to interrupts, e.g. when an interrupt occurs either this instruction has been executed completely or not at all. You want to make it also atomic with respect to code executing on other cores? Just add a LOCK prefix.
- Even in 32 bit mode ARM already offers many registers (16 GPRs) and in A64 even 32. x86-32 only offers 8 and x86-64 16. Having more and orthogonal registers is typical for RISC vs. CISC. Thus ARM's call convention uses registers for passing the parameters (again very typical for RISC) and uses a link register for passing the return address. x86-32 instead typically passes return address and parameters on the stack (I know that under 32 bit Windows there also exists the __fastcall call convention that uses registers for passing parameters, too. But this is an exception). But admitted: under x86-64 on the Linux and Windows call convention they also use registers for passing parameters.
- At least A32 and T32 cannot directly call functions that are too far away from the current instruction. So the linker has to insert veneers, which make this possible. One of course will not find veneers on x86.
- It is not easy to access memory addresses relative to the instruction pointer on x86-32 (under x86-64 it is easily possible), while on ARM it is. So ARM assembly code will often contain small const variable values next to the function. One hardly sees such code on x86-32 (though on x86 it is not really necessary, since there are instructions that put some arbitrary constant into a register; A32 and T32 on the other hand allow only to encode some kind of constants, bases on a rather complicated scheme that I will not explain here; so this trick is rather helpful for writing ARM code).
TLDR: Better look again at the disassembly carefully.
Oh, it's not that bad: It codes an 8-bit value and a 4-bit circular shift-by-2-bits.
Well, you should start from this and then work backwards because it's just an empirical fact that x86_64 is fast and that RISC isn't. There are many reasons for that, none of them fair, but in the end, it's still just a fact.
Yes, RISC-V is simple+elegant. If that's what you want, great. True, RISC-V is open source. If that's what you need, awesome. Now weigh those qualitative features against the cold hard $1B/year that Intel invests in making the mess that is x86 run fast. Add to that the billions that ARM+Samsung+NVidia+Apple+... spend on making ARM fast. It's not a fair fight. It really isn't.
But know this. If RISC had anything tangible, any fundamental advantage, then we'd be living in a RISC world by now and we aren't (Patterson+Ditzel was 37 years ago). If you want to think worse is better, sure. Whatever. If you want to call ARMv8 MIPS-ish, that's your constitutional right.
Skylake is a superscalar, speculative, out-of-order, renaming, hyperthreaded multicore beast. You may program it in x86 but those instructions get translated+cached as μops, very wide microinstructions.
BTW, anyone who thinks that μops are actually RISC under the hood severely needs to re-read The Case for the Reduced Instruction Set Computer and then say, Inside Nehalem [1] or Micro-operation cache: A power aware frontend for variable instruction length ISA [2]; microprogramming (vertical+horizontal) predates RISC by decades [3]. These μops are wide, 150+ bits. Ain't nothing reduced about that.
In 2017, believing in RISC is only slightly more acceptable than believing in Mill. Moreover, I say this as someone who went to Berkeley and read Patterson+Ditzel in 252. I should preach the religion.
[1] http://www.realworldtech.com/nehalem/
[2] https://www.researchgate.net/profile/Ronny_Ronen/publication...
(I may have misquoted badly, in particular I'm not sure about the dates)
(edit: I found the talk, it's the last couple minutes of this: https://www.youtube.com/watch?v=LgLNyMAi-0I)
So what could you do with X transistors? But then X became stupid large. The 68000 (1979) had 40,000 transistors. Now the Apple A10 has 3.3-billion transistors. So you can imagine that architectural design assumptions dating from 1980 will need to be revisited.
Where did you get that figure? I'm curious because the figure I've seen tossed around is 68,000 transistors (the story going that this is where the model number came from).
https://www.motorolasolutions.com/content/dam/msi/docs/en-xw...
The Geometry Engine had 40,000.
https://books.google.com/books?id=gD4EAAAAMBAJ&pg=PA17&dq=68...
This was also the very brief time when memory was actually faster than the core, making it possible to contemplate things like large fixed-with instruction encodings, since the main bottleneck with the early CISCs was instruction decoding and not fetch bandwidth; it is unlikely that, had this period not existed, RISC would have developed in the way that it had.
[1] http://mprc.pku.edu.cn/~liuxianhua/chn/corpus/Notes/articles...
http://patents.justia.com/assignee/mill-computing-inc
I'm impressed at their doing something completely different and do wish them well. I know someone on the team but I'm impatient.
I wonder how hard it would be for Intel to substitute an alternative decoder in place of the current x86/x64. Given the existing micro- and macro-fusion capabilities, is there anything that would prevent them from reusing the loop streamer and decoded instruction cache and entire back-end for a different instruction set? While there are probably legal issues with ARM or Power emulation (are there?), it seems like it wouldn't be too hard for them to quickly put together a high-performing hyperthreaded superscalar out-of-order alternative for RISC-V should they ever want to.
BTW, anyone who thinks that μops are actually RISC under the hood severely needs to re-read ...
Could you expand on this? I take you to be saying that Intel µops aren't really the same as classic RISC at all, but you seem to be the only person in this thread that holds this position. The majority appear to be saying straight-out that current Intel is actually RISC under the hood. Personally, I suspect that you are right, and that it would be useful for you to make this claim more forcefully. Or maybe I'm the only one that believes this, and I'm misinterpreting you?
The big idea of RISC was to avoid microprogramming and have a simple to decode ISA which was directly executed by a short simple pipeline. ld/st + lots of registers + C compiler. Awesome.
Indeed in Patterson+Ditzel's version of Luther's 95 theses, they say:
Microprogrammed control allows the implementation of complex architectures more cost-effectively than hardwired control
They say this in the section Reasons For Increased Complexity. RISC is revolting from that complexity, from microprogramming, so how can μops be RISC if that's what they were revolting from?I think people say this because no one knows what microprogramming is anymore. You might read one paper in a graduate architecture seminar. And no one but no one writes microprograms.
The decode stage of Haswell translates add RAX,RBX into a '150b' μop. If can you read Agner Fog and you'll know how many μops, what the latencies are, etc. These are empirically determined properties of the μop but that's it. You know the Decoded ICache characteristics. But you don't know the ISA.
After all that's done, you're left with maybe a 14-pipestage data path. That's not RISC. Complicated instructions (AAA) get interpreted. That's not RISC. Multiple functional units can calculate 2+3 operand effective addresses. That's not RISC. This all is REALLY not RISC.
μops are micro-instructions, not RISC instructions.
BTW, the alternate decoder kinda wouldn't work so well because the microarchitecture, functional units, datapath, caches and registers are really set up for x86. Folks wanted something like that at Transmeta, a Java processor, but the HW was designed for x86.
Thanks for the clear restatement!
Complicated instructions (AAA) get interpreted.
It may be worth mentioning that the BCD suite of instructions are not supported in x64. But if you are looking for an example of modern microcoded x64 to avoid, BTR/BTC/BTS with a memory operand is a prime example. Which does bring up the point that talking about "x86/x64" can be misleading in a similar way as discussions of the "C/C++" language.
the alternate decoder kinda wouldn't work so well because the microarchitecture, functional units, datapath, caches and registers are really set up for x86.
Are there specific areas where you see problems? For example, I'd think the existing abstraction between limited number of architectural registers and the hundreds of physical registers would work fine for alternative ISA's. And I don't immediately see why caching wouldn't work unaltered.
Folks wanted something like that at Transmeta, a Java processor, but the HW was designed for x86.
I searched for more information about what became of Transmeta after I posted last night, but didn't find much about the technical issues they encountered. Do you know if there is a good post-mortem?
While searching, I did come across a couple interesting CPU's that are going the other way:
Loonson 3 is Chinese MIPS processor that emulates x86 and ARM: https://venturebeat.com/2015/09/03/chinas-loongson-makes-a-6...
Elbrus-4S is a Russian VLIW design can emulate x86: https://www.extremetech.com/computing/205463-shadows-of-itan...
The main problem I see mentioned when going the other way (ARM emulating x86) is dealing with "flags". Since recent Intel processors already support "flagless" variants of most of the instructions (SARX, MULX, etc) this doesn't seems like it could be insurmountable.
Also, register renaming helps a LOT; so you don't need a 32 direct registers when 16 registers renamed to 180 internal registers will more than do. Register renaming predates RISC by 15 years (Tomosulo). I'd almost say it's anti-RISC.
BTW, Elbrus + Transmeta are kinda joined at the hip. Babayan consulted at Sun with Ditzel and is now at Intel. Half the Transmeta people went to NVidia (eventually) and did an x86 before switching to ARM (Denver).
Eventually people will see RISC as the provisional idea it is. Register renaming is a Great Idea. Uop translation is a Great Idea. RISC is provisional.
I see the much larger problem in ARM emulating the memory model of x86, which gives much stronger guarantees on ordering and synchronization than the weak memory model used by ARM processors.
x86 took over because of a number of factors, and over time improvements made it more RISC like internally. Read up on modern CPU design to see how RISC, CISC or whatever is just an API of sorts to the internal workings of the CPU. The two philosophies have become so blurred few people even use the term RISC or CISC any more. It's irrelevant.
Often when an "inferior" technology wins, it's because you're focusing too narrowly, and the winner is outright better in some way.
For x86, Intel's process and design expertise overcame the shortcomings of the ISA, and then some. You're not necessarily buying x86, but rather buying the best silicon that happens to speak x86.
Intel also made, and probably still does make RISC chips. The i960 and i860 are just two examples. https://en.wikipedia.org/wiki/Intel_i860 They never took off on any appreciable scale.
> You're not necessarily buying x86, but rather buying the best silicon that happens to speak x86.
Or more specifically, you're buying the best chip that runs your existing software and works with the toolset you're familiar with.
This is why Intel's Itanium project was doomed from the day they announced it. Nobody was going to re-write everything to work with their new instruction set.
It's notable that Apple managed to go from 68K to PPC to x86 to x64 almost seamlessly, but they did that at the expense of backwards compatibility. You can run many Windows apps from the 3.1 days in current versions of Windows without emulation, but you can't run Mac software from the 68K days without it. Two different approaches.
The Intel world is ruled, if not held hostage to backwards compatibility concerns. "Better" means "more traditional".
VHS looked like garbage, Beta always looked better, but VHS was cheaper, more convenient, and ubiquitous.
You're right about the length thing being a hassle. Friends who had Beta decks always had to jockey tapes in the middle of a movie, not unlike later when you had to flip a laserdisc. It was always a point of ridicule. Looks great, if not amazing, but you were always X minutes away from having to get up and change tapes.
VHS could handle even the longest movies with ease, though often the quality would suffer accordingly.
Sony had actually bet on one hour initially because TV shows were mostly half an hour or an hour. This left movies and sporting events mainly as the issues with recording time. They eventually brought out longer tapes, but it was a combination of factors up to that point that kept Beta from leading.
Now, compare and contract MiniDisc vs. CD, Memory Stick vs SD, UMD vs. Download vs. SD, and where Sony actually won with Blu-Ray vs. HD-DVD. MiniDisc went away. Memory Stick mostly went away. UMD went away. But where Sony widely licensed the superior technology early on, it won.
Only on 32 bit versions of Windows (and even this only sometimes works by this time). But note that the main reason why you cannot run Win16 applications on 64 bit Windows is not:
- that you cannot run 16 bit applications in long mode (you can - look up the details of in the Intel documentation if you don't believe; just generate a suitable segment descriptor)
- Virtual 8086 is not supported in Long Mode (it is indeed not - but this is only important for applications running in real mode (e.g. DOS applications) and thus is not of importance for Win16 applications)
as it is often claimed in the internet, but has to do with Windows' architecture:
> https://msdn.microsoft.com/en-us/library/aa384249.aspx
"Note that 64-bit Windows does not support running 16-bit Windows-based applications. The primary reason is that handles have 32 significant bits on 64-bit Windows. Therefore, handles cannot be truncated and passed to 16-bit applications without loss of data. Attempts to launch 16-bit applications fail with the following error: ERROR_BAD_EXE_FORMAT.".
Not to toot my own horn here, but we're one of the hopeful deep learning chip startups, I'd be happy to answer any more questions or perhaps elaborate on why I think ZISC is better for neural nets.
If you burned an ASIC for a single type of network architecture, then that would be something else.
The result is called a Tensor Processing Unit (TPU), a custom ASIC we built specifically for machine learning.
https://cloudplatform.googleblog.com/2016/05/Google-supercha...
Google's TPU architecture is CISC type instead of ZISC type: https://www.nextplatform.com/2017/04/05/first-depth-look-goo...
I don't understand what you want to say: ZISC is a rather different concept: https://en.wikipedia.org/wiki/Zero_instruction_set_computer
In any case, that being said, the TPU is less of a CISC and much more like a lightweight RISC control mechanism.
I tend to think of "CISC" as complex instruction encodings, smaller register files, complex sets of addressing modes, lack of orthogonality in instructions, etc. None of these aspects are gaining in popularity for new designs. The only one I can think of that gained any popularity is Thumb-like compressed instruction sets, which aren't so much "CISC" as "don't spend opcode space for no reason".
FWIW, Intel uses neural nets to speed up its branch prediction. Yes, you read that right.
[0] https://groups.google.com/forum/message/raw?msg=comp.arch/9T...
[1] https://en.wikipedia.org/wiki/John_Mashey
Edited to remove a duplicate footnote.
There is no true CISC anymore, everything is RISC under the hood with a CISC to RISC conversion layer.
Cool to know the story but don't trust what a career prof spouts. Often it's not actually true.
x86/64 processors features very refined hardware prefetcher, caching system, jump prediction.
Such processors will run average code runs more effectively than it deserves.