A New Golden Age for Computer Architecture
cacm.acm.org
cacm.acm.org
> Concluding this historical review, we can say the marketplace settled the RISC-CISC debate; CISC won the later stages of the PC era, but RISC is winning the post-PC era
It is clear that his assessment is right, but isn't the 99% number too high ? Servers, laptops and desktops still run x86 and they are CISC ( unless you are counting x86 as RISC based on microcode )
> Many researchers assume they must stop short because fabricating chips is unaffordable. When designs are small, they are surprisingly inexpensive.
> High-level, domain-specific languages and architectures, freeing architects from the chains of proprietary instruction sets, along with demand from the public for improved security, will usher in a new golden age for computer architects. Aided by open source ecosystems, agilely developed chips will convincingly demonstrate advances and thereby accelerate commercial adoption.
It will be interesting if manufacturing also gets open sourced. There already seems to be a project attempting this : http://libresilicon.com/
In fact now that I think about it, pretty much every SSD contains a controller SoC which is almost certainly RISC, so even in a standard laptop you get at most a 1:1 ratio of RISC:CISC. And modern GPUs (including IGP) would be RISC in a VLIW configuration, if each computation core is counted as a RISC chips the ratio quickly gets ridiculous.
If you manage to find all the processors I bet it's more like 10:1. To a first approximation everything on the PCI and USB buses will have its own processor, even if it doesn't accept external firmware.
I feel there's a long way to go there. The economic structure of the industry is against agility because "deployment" remains stubbornly expensive, and the product culture is also much more conservative.
> manufacturing also gets open sourced.
It's one of the most capital-intensive industries in the world, so I don't quite see how this would work? Libresilicon are offering a 1000nm (not a typo) process.
My point? Open sourcing stuff with ridiculously huge capital requirements is not that useful, I guess?
Also I do not think uops are fixed size as IIRC they can take a variable number of slots in the uop cache, and fix size instructions is pretty much one of the only two remaining differentiating RISC features. The internal x86 microarchitecture is also not load-store, the other one RISC feature, at least in the fused domain, and as far as I understand, in the uop cache.
So, even if we want to abuse the RISC term to describe the microarchitecture, I do not think it cleanly apply to the usual x86 implementations.
edit: this is a pet peeve of mine. It seems I have this discussion every 6 months on HN :)
But the important thing is that uOps are much higher level than microcode instructions. Except for the odd encoding size they would make a lot of sense as an early RISC ISA. Now, they expose a lot of the odd corner cases of the underlying architecture in a way that no modern ISA would but the original Berkeley RISC had branch delay slots and followed the philosophy that you'd just recompile the code when the ISA changes.
I'm at the edge of my knowledge here but I understand that microcoded instructions would tend to be much lower level, being things like read from memory to such and such an internal buffer. By contrast uOps do specify registers or constants, though they do so (post-rename) in terms of physical rather than architectural registers. But the decision on whether to get that arguments from the physical register or the bypass network is still made further down the pipe as with a RISC processor.
Is the analogy perfect? No, of course not. No analogy ever is. But I do think it illuminates more than it misleads for people learning about the evolution of processors - just as long as people can keep architecture and micro-architecture straight.
[1]https://en.wikichip.org/wiki/micro-operation for instance. [2]https://www.realworldtech.com/haswell-cpu/2/ [3]https://en.wikichip.org/wiki/amd/microarchitectures/zen%2B#M...
Encoding constants in the instructions themselves is a very non-RISC thing BTW.
Not really. Micro-Ops are typically very large (100+ bits wide) where each bit can be thought of as directly controlling a specific function in an EU. They can do things in parallel; the frontend may emit only one uop for more than one ISA instruction, they can contain constants, they're kinda-of variable-length in some microarchitectures. Overall they're very un-RISC-y.
Overall the whole RISC/CISC debate is pretty much meaningless and has been for decades. Many folks superimpose their own superstitions about unrelated issues (e.g. "PC server" vs "UNIX server" seems a popular one), but at the end of the day pretty much all high-performance cores look fairly similar, regardless of ISA.
Now if one would take what you said at face value then it would actually imply that RISC as an ISA is completely irrelevant because it is possible to achieve the same benefits or even beat it despite CISC having the inefficiency of hard to decode instructions and the high cost of a conversion step in the micro architecture. One suddenly realizes that this battle of ISAs is completely futile and the secret sauce is in the micro architecture which is completely divorced from the ISA.
Here are some examples: ARM makes slow ARM chips. Apple makes fast ARM chips. AMD made slow x86 chips in the past but now adopted a faster micro architecture. Intel made Itanium but the chips didn't have any sort of dynamic scheduling so they couldn't deliver the promised performance gains.
There’re instructions combining a dozen of math operations and also RAM loads, that on modern CPUs decode into just a single micro-op.
Example: https://www.felixcloutier.com/x86/vfmadd132ps:vfmadd213ps:vf... The AVX version computes x=a*b+pointer[i], for 8 independent SIMD lanes.
If you call that amount of stuff in a single micro-op “reduced instruction set” I wonder what exactly is the meaningful definition that you referred to?
Add in all the microwaves, routers, the many processors in your car, and so on and 99% seems a bit high to me but not unreasonable.
Anything (like a Printer) with a web based interface, having a micro web server within the device.
There are VASTLY more ARM and other architecture chips than there are x86/64 chips around you right now. Possibly even in the computer monitor you're reading this on. The desk phone in your office.
Just about anything that has any kind of a screen with menu system.
The SD card has an ARM chip, usually, in addition to whatever is in the camera itself.
I'm sure you can think of a million other markets like this- sedans vs race cars, fighter jets vs puddle jumpers, etc.
But really I think we really do overemphasize the importance of x86 because that's the architectures we have the most experience working with directly.
"The Machine will be a complete replacement for current computer system architectures. There will be a new operating system, a new type of memory (memristors), and super-fast buses/peripheral interconnects (photonics)."
"HP says it will commercialize The Machine within a few years, “or fall on its face trying.”"
It seems later happened...
[1] https://www.extremetech.com/extreme/184165-hp-bets-it-all-on...
So they will deliver it, it just won't be anything like what they promised.
You can buy NVM today [1], and building systems for resource disaggregation work is an active problem [2].
[0] https://www.usenix.org/conference/fast14/technical-sessions/...
[1] https://www.intel.com/content/www/us/en/products/memory-stor...
Computers, like guns, drugs, and any other invention of man are not inherently evil. All these things can be used for good or evil. Unfortunately, the fly in the ointment is human nature. With the convergence of cheaper but increased computing power and the monetization of personal information, I fear what the future holds. I hope I'm wrong, but it looks to me that humanity is doomed to forever live in a state total surveillance and control. We're seeing it happening already. Just the other day I saw an article that said that Sweden is going to tax people on the miles they drive. The very next day I saw another article saying Los Angeles is planning to do the same thing.
Sigh... I'm glad I'm old.
You essentially have heavy-duty or commercial trucks that don't "pay their weigh" and hyper efficient hybrids and EVs that don't "pay their way" when it comes to infrastructure costs.
So fuel tax isn't equivalent to a tax on "miles driven," it might be better suited to off-setting air pollution (pure conjecture on my part), but if you wanted everyone to be responsible for the damage they cause to public infrastructure, you would need some weird calculus of axle-weight/mile driven tax.
Or we could decide major roads/infrastructure are an economic public good worthy of paying taxes on.
If having a GPS device on your person becomes law, then avoiding the tax by riding a bicycle, walking, riding a horse etc is defeated.
Chuck Moore's "Green Arrays" is kinda cool and so is the Parallela board.
https://en.wikipedia.org/wiki/Transputer
https://en.wikipedia.org/wiki/XCore_Architecture
SOAR (Smalltalk On A RISC), though the conclusion there was mostly that a plain old RISC will do. I wonder if that is still true today.
Rekursiv OO computer https://en.wikipedia.org/wiki/Rekursiv
NEC dataflow processor. https://books.google.de/books?id=qRrlBwAAQBAJ&pg=PA152&lpg=P...
https://www.deepdyve.com/lp/association-for-computing-machin...
http://digitalassets.lib.berkeley.edu/techreports/ucb/text/E...
Looking at the results, they say that hardware tag-checking for integer arithmetic and register windows for fast method calls were the two most important features of the design, nearly doubling performance.
I wonder if that still holds today, with the memory wall so dominant that CPUs tend to be stalled quite a bit (therefore enough time to do tag checking in software).
For a time the T800 was the top of the FP pile. Not for long, though, and then the long, long, long wait for the disappointing T9k doomed the whole architecture.
"A Transputer had a number of simple operating system functions built into the hardware. These included hardware multitasking with foreground and background priority levels, hardware timers, and hardware time-slicing of background tasks."
The Hardware and Software Architecture of the Transputer -- https://archive.org/details/Xputer
There are special instruction to start and end a process, and the fact that it was a stack machine means context switches were extremely fast, almost no registers to save/restore.
> so much as "integrated communications network."
It had both, and both were integrated. IIRC, the instructions to send/receive on the links were integrated with the multitasking hardware.
"The first 16 'secondary' zero-operand instructions (using the OPR primary instruction) were:
Mnemonic Description
REV Reverse – swap two top items of register stack
LB Load byte
BSUB Byte subscript
ENDP End process
DIFF Difference
ADD Add
GCALL General Call – swap top of stack and instruction pointer
IN Input – receive message
PROD Product
GT Greater Than – the only comparison instruction
WSUB Word subscript
OUT Output – send message
SUB Subtract
STARTP Start process
OUTBYTE Output byte – send one-byte message
OUTWORD Output word – send one-word message"
https://en.wikipedia.org/wiki/TransputerAlso, you could designate memory locations as local communications channels, and the same instructions would work. So the same binary could run locally or distributed.
I hope this won't mean the Mill has no chance of success, or at least influencing the industry for the better. Every video seems to introduce interesting new ideas, or new takes on old ideas[0].
To contrast my fantasy for "Golden Age" would be multiple viable replacements for CMOS that were actively being used in a variety of processors.
It will cause a crash in the entire tech industry, which will throw the global economy into a depression.
We’ve gotten a bit of a reprieve because things like GPUs turn out to be good architectures for many compute intensive workloads.
Do you mean that future requirements will outstrip capacity? I don’t think this has been true for some time, we run a lot more systems and tend to scale horizontally; doubling compute power every n period isn’t a hard requirement imo.
For example Tesla is betting on conventional cameras and tries to overcome their shortcomings with computationally intensive machine vision. Tesla's custom chips allows them to have more processing power which means they can install more high resolution cameras and other sensors.
Between GPUs, more power efficient designs (due to heavy mainstream interest in mobile technology), more work put into algorithmic efficiency, and promising early developments in quantum computing, it appears the focus has shifted away from relying on more transistors per square inch for ensuring the industry's future growth.
It was always going to end at some point, because we're down to just above 100 atoms gate thickness. What we're seeing is not an "end" but a soft landing that really started some years ago with massively multicore processors. The industry has started adapting to horizontal scaling rather than "free" process improvements. There is still an awful lot of low hanging fruit on the software side for performance improvements, but that's harder because it involves removing or adapting abstraction layers.
Why?
Well, what happens when next year's smartphone really isn't any better or cheaper than last year's smartphone? Sure, there will always be some business because people drop their phones all the time, but sales won't be as high without upgrades driving it.
Lots of people (investors) will be unhappy with the tech sector at that point.
I also dispute the scope you assign to this "upgrade" cycle. It's a huge overstatement to claim that the whole tech sector would "collapse". The tech sector is far larger than phones and personal computers.
Slightly cheaper, maybe, as the fab equipment depreciates.
> I also dispute the scope you assign to this "upgrade" cycle. It's a huge overstatement to claim that the whole tech sector would "collapse". The tech sector is far larger than phones and personal computers.
It used to be that new hardware would drive new software sales, and vice versa. You'd buy a new PC, because Windows XP ran too slow on your old one. And a couple years later, some awesome game comes out, and you'd need to buy a new PC to play it.
Expansion, not replacement (of broken systems) is what drove the PC segment. It was the foundation of the growth that affected everything that used or could benefit from PCs. And we've seen that with the mobile segment.
And we may yet see that with the VR/AR segment, but that's going to be a tough hill to climb if we can't count on continuing IC improvements.
Sure, but the same equipment would also be cheaper to replace than the capital required for the next generation fab, for the same reasons.
> Expansion, not replacement (of broken systems) is what drove the PC segment. It was the foundation of the growth that affected everything that used or could benefit from PCs. And we've seen that with the mobile segment.
Expansion into new sectors, not expansion into the same sectors. The same will still happen. Solutions will become more customized rather than remain general purpose. There remain considerable gains in parallel computation for instance (GPU), and ASICs will become the new hotness (again).
Furthermore, if clients cease to expand, then more computation will happen on server infrastructure. Your fat clients will become more thin clients again, repeating the same cycle that has happened multiple times so far.
You're also neglecting the investment in infrastructure (data centers, networks) that will continue to grow, if not accelerate. Current hardware is more than sufficient for most of this.
I'm a bit skeptical of the promise of DSAs, though it does seem we're already going that way. Curious what others think on that point.
I think the real question is if it will make sense to have FPGAs in wider use. Certainly not until the programming model improves...
RISC makes a lot of sense given compilers and many other possible optimizations?
However, there are increasing trends to put functions right into silicon too. Those frequently replace many instructions.
Will ML tech somehow make better sense of those things, and CISC in general?
That depends on the number of cores.
The power of computing is redefining civilization, humanity.
[1]https://www.destroyallsoftware.com/talks/the-birth-and-death....