RISC creator is pushing open source chips
gigaom.com
gigaom.com
In 2001, an unexpected delay of 18 months could mean that you release to competition running at 2ghz instead of 1ghz. These days, Intel's 32nm, 22nm chips (and 14nm, from what is known about Broadwell) are more or less identical from a consumer's perspective.
As an aside, Patterson's book (coauthored with Hennessey), Computer Architecture: A Quantitative Approach, is one of the best I've ever read. Right up there with SICP.
I think the reason that the 32nm, 22nm and 14nm chips are identical from a consumer approach is because of two issues.
One is marketing - we no longer have convincing benchmarks that are as persuasive as the GHz a chip would support. It was a questionable metric anyway, but it now isn't being pursued (for a whole number of reasons).
The second is an application space problem - consumer applications are no longer exploiting CPU improvements in the same manner. That's partially because we now get more parallelism rather than single threaded speed, but it's partial a testament to the success of the previous generation as to the breadth of tasks a current PC will happily perform in a reasonable time frame.
Maybe when we finally get to quantum computers a similar cycle will happen, but to the dismay of OEMs the cycle is gone.
Even mobiles already reached the same peak and are mostly driven by contract renewals nowadays, not features.
Desktops are as powerful as anyone needs - largely because you just buy the appropriate size - at least at the consumer level.
Laptops, phones? Definitely not. But the push for more isn't being limited by our performance, it's being limited by energy demands. I'm shopping for a new laptop at the moment, and my overriding requirement is to get the most WHr's of battery into a low energy configuration.
And even there, to some extent the problem is also not where you'd think - you can take a modern i5 down to something like 1.5W idle consumption. Most of the losses are switching circuits or the screen.
Other than that, we also have same hybrid multi-core devices on our pockets that we have on our desks. Coupled with all the hardware sensors one can imagine.
So besides the power consumption point you rise, there is hardly anything else to improve at hardware level.
Except maybe, for things like better voice recognition experience or holographic displays a la Star Wars.
Don't worry, they won't give up without a fight. I bet they'll try to push some artificial reasons for people to keep throwing away perfectly good machines every two years.
This related paper is an interesting read: http://research.cs.wisc.edu/vertical/papers/2013/isa-power-s...
I think x86 still has plenty of room for optimisation, and Intel are just doing it slowly according to market conditions. An open-source x86 core would be great - in fact, pre-MMX-level x86 is fully open, as all the patents on that have expired now.
x86 are no longer real CISC, as they moved the internal architecture to RISC, starting with the Pentium 4.
And all the other processors that matter are also RISC.
The real loosers were the VLIW proponents.
I'm curious which processors that matter are all RISC. I mean, I realize there are a fair bit of chips out there. I wouldn't have thought many of them really fit the original RISC mold. Sure, they have less instructions than some of the old behemoths. But, that is a far cry from the aims of RISC.
There is also the rise in specialized processors. Combining that with the increased focus on consolidating everything into one chip, and much of the debate is just weird nowdays.
Easy way to consider it, a GPU would probably not qualify as a RISC processor, but it also probably has fewer overall operations than a CISC processor. (Or, do they qualify nowdays?)
Now, ARM definitely is. It remains to be seen whether they can continue to be dominant once Intel enters the same arenas.
The upshot is that simple scalar code tends to be somewhat more compact on x86, and tuned vector code (likely to be found in perf-critical routines that are CPU-bound) tends to be somewhat more compact on arm.
More to the point, loop buffers and µop caches on the last couple generations of processors make encoding density mostly irrelevant to performance (though occasional pathological examples do still exist).
As in, the part where x86 processors get to fully enable the "operate as a RISC" processor mode.
This makes ARM (especially ARMv7 and newer) a mixed setup like x86.
I suppose my point is that ARM is going towards the same inbetween-ism that x86 has moved to.
These custom chips have the potential to be much more power efficient (for their special class of applications) than off-the-shelf general purpose chips.
With RISC-V it will be easier to start your own CPU design...
A few examples of this (none of which are unique to RISC-V except in aggregation): no delayed-branch semantics, source and destination registers are always in fixed location (see then store vs. load encoding) and are always explicit, sign bits are immediate fields is always in the same location, etc.
1. Nobody can put more than XXX transistors on a square millimeter (where XXX is some physical limit to chip material)
2. "Everybody" can put YYY transistors on a square millimeter (where YYY is a number sufficiently close to XXX so that the difference does not matter much)
The first means some stagnation in the chip market: not only does Moore's Law not apply anymore, but not even does a weak version of Moore's Law ("chips tend to get faster/cheaper over time") apply. But if that scenario becomes true, it doesn't necessarily mean that the democratization of the second scenario becomes true. The ability to put YYY transistors on a square millimeter could still be limited to a select few companies investing billions in the production of these chips. After all, how many fabs worldwide are currently capable of producing 45nm chips, which could be considered "commodity" [1]?
Is it really true that it becomes much cheaper to produce chips on smaller process nodes over time? Sure it becomes cheaper since less original R&D is necessary, but is it reasonable to think that producing 20nm chips ever becomes a question of investing anything less than, say, 100s of millions? Does the second scenario ever become true, at least in the forseeable future?
I don't know much about chip fabs so I wonder if someone more knowledgeable could share some insight on this.
[1]: I genuinely don't know the answer to this, but I don't suppose it's very widespread?
To get your own chip produced using a 'pure-play' fab (which, these days, basically means TSMC, and to a lesser extent Samsung and GloFo), however, is much cheaper. If you want a 20nm chip, you design your chip (basically the cost is just engineering effort + CAD tools..a few million $$), send it to TSMC (getting masks made is another few million $$), and they send you back chips.
Reality is slightly more complicated, in that you're using using an integrator to manage the relationship with TSMC, the packaging company, the testing company, etc., but a startup with $10-20 million could probably scrape together a decent 20nm chip.
Once TSMC has built a new fab though, the longer it is 'relevant' the cheaper they can make the wafers, as they have more wafers to amortize the ~$10 billion or whatever it cost to build the fab over.
Although I guess the latter really isn't that relevant when it comes to the discussion of chip design.
The reason it becomes cheaper is that your development costs for both the process and the design get amortized over longer and longer periods. At the same time architecture becomes the only way that some chips can be faster than others.
Though I guess I should point out that progress in chip performance isn't the same as putting more transistors on a square millimeter. One aspect of this is that more economical silicon processes such as TSMC's are better at cramming more transistors into a square millimeter than Intel's, but Intel's transistors tend to have better drive current - all at a given process node. The other aspect is that ever increasing transistor leakage means that processors might get to an era when they can't afford to keep all their transistors lit up at the same time[1].
[1]http://hopefullyintersting.blogspot.com/2012/02/dark-silicon...
Ironically, FPGAs as endgame would be way more disruptive than an "ASICs to the people" endgame.
FPGA density, price/gate and power efficiency trail "native" hardware but constant factors of overhead but it should be possible to reduce the gap compared to now. More FPGA volume means better economies; the hardware architecture could be optimized further; most importantly IMHO the current design tools are heinously unfriendly and could benefit a lot from programmer attention (once programmers realize a more-or-finalized FPGA structure is the "new assembly" to optimize for).
I'm not sure if a constant factor of disadvantage would become very acceptable (because we'll drop the throw-faster-hardware-at-it mentality) or very unacceptable (because robots with FPGA brains always lose at high-frequency chess wrestling to robots with native brains).
That’s a little bit of a stretch. 32nm was Westmere -> Sandybridge, which doubled FP throughput and L1 load bandwidth, among other changes. Haswell doubled FP throughput again, doubled integer SIMD throughput, and doubled L1 load and store bandwidth, among several other significant changes, and there’s also been a steady ~10% per generation improvement in generic IPC. Consumers don’t care about those numbers, but they do care about some of the operations they enable getting faster. (In fairness, this often lags the release of the processors by a year or two as SW re-optimization is required to take advantage of new features; the full benefit of the changes isn’t available to consumers until some time after the processors are released, by which time we’ve all forgotten just how slow our old hardware was.)
The energy footprint of those processors has shrunken considerably in the same time, which is a huge improvement for portables.
Yes and no. The Westermere, Sandybridge, Ivybridge, Haskell, Devil Cannon's line isn't really targeted for the mobile market. The shinking power budget is caused by 2 things.
1) To stay competitive against 64bit ARM processors. Intel processors are barely used in mobile devices. But eventually ARM will attempt to transition into laptop/desktop/server and compete with Intel directly, this is where lower power comes into play.
2) Feature set. New features (functionality) on chip cost not just dye space, but cost power and heat. The heat and power are the far most critical issues.
The second issue is commedically addressed in, "The Cold Winter" a satirical document about hardware engineering.
"The idea of creating more and more cores ran into an issue. Most people don't use their computer simulating 4 nuclear explosions while rendering Avatar in 1080p. They use their computer for precisely 10 things, of which 6 of things involve pornography.
The other issue was the 600 core dub 'Hydra of Destiny' was so smart that its design document was the best chess player in the facility. The issue was it required its own dedicated coal fired plant, and would run so hot that it would melt its way into the earth's core." - The Cold Winter (paraphased)
Yes and no. The Westermere, Sandybridge, Ivybridge, Haskell, Devil Cannon's line isn't really targeted for the mobile market. The shinking power budget is caused by 2 things.
1) To stay competitive against 64bit ARM processors. Intel processors are barely used in mobile devices. But eventually ARM will attempt to transition into laptop/desktop/server and compete with Intel directly, this is where lower power comes into play.
2) Feature set. New features (functionality) on chip cost not just dye space, but cost power and heat. The heat and power are the far most critical issues.
If your interested in a slightly satirical look at modern hardware scaling I suggest you read The Slow Winter http://research.microsoft.com/en-us/people/mickens/theslowwi...
I write optimized compute libraries for a living. Haswell really was an enormous improvement for real workloads (actually, more than 2x for some integer image-processing tasks due to three-operand AVX2 instructions eliminating the need for more moves than the renamer could hide in some loops). Is everything magically 2x faster? No, of course not. Are a lot of important things significantly faster? Absolutely.
We did not include special instruction set support for over ow checks on integer arithmetic operations. Most popular programming languages do not support checks for integer overflow, partly because most architectures impose a signicant runtime penalty to check for overflow on integer arithmetic and partly because modulo arithmetic is sometimes the desired behavior. [1]
Please Regehr, don't hurt them. [2]
[1] Spec, section 2.4 https://s3-us-west-1.amazonaws.com/riscv.org/riscv-spec-v2.0...
Scheme and Common Lisp also rely on overflow checks to optimize small integers as "fixnums" instead of always using arbitrary precision arithmetic. Not having hardware support for overflow checks would complicate the implementation of many dynamic languages, and reduce their performance significantly.
Not sure what this guy was thinking. It can't be that hard to implement some overflow flag you can branch on, I mean, adders basically produces that information for free, don't they? Seems like a poor design choice.
> Hmm; carry is free, not sure about overflow.
Overflow is just the final carry out XOR the second-last carry bit, so it's practically free.Of course RISC-V doesn't have a carry bit either!
[1] NaR will propagate an error, but trigger an exception if it's stored or branched on.
IMHO this is cost saving going too far..
I wish that the mill succeed instead, at least it brings new things regarding security!
At the end of the day, the chip is a physical thing that have to be produced by someone. As long as there is no 3D printer, which can build nano structures to makes chips, I can only implement it into a FPGA. The main FPGA providers are either Altera or Xilinx and both a very expansive. So they are good for prototyping or very low volume. But for anything beyond, I need a specific implementation. Okay, I could go to a foundry like Samsung (or other) and order according to my design. But, even that requires a very high volume in the millions to make it affordable and viable from a business perspective. Especially, if it is intended for the IoT market. On the other side, I can buy ARM Cortex M0 and Cortex M1 for cheap from NXP and other for less than $1. They are power full and power consumption is very low.
Just to mention, the other day the WRTnode[1] was released for less den $25 or look at an Raspberry PI Compute. Don't get me wrong, Yes, I would like to see those with an open source chip. But would that be really a viable business case, except from the University context?
So imagine how cheap these new CPUs could be!
In hardware design you cannot build a "minimal viable product" and then do "continuous integration" from that on. A hardware product either works within its defined limits or it doesn't. Then, depending on your limitations, you need additional approval by FCC, FDA, and such organizations within the US alone. The same thing then with their counterparts in every different country. Chip design is then the next level up. Because changing a mask later on, because you found a bug, means you produced a million chips just for the dust bin.
That means the entry barrier into hardware, especially chips, is very high.
Perhaps the cost situation will change over time. Perhaps, in coming decades, the ability to fab microchips will itself become something achievable for thousands instead of millions of dollars. And as for regulation: That only matters if you're selling. If you're researching or making it for personal use, the FCC doesn't (or shouldn't) matter.
Also, return to the taxpayer. To me, a publicly-funded verification of RISC-V would bring wider benefits than that of a proprietary ISA like ARM [1].
Why would you need 128bit addressing? Isn't 64bit address space plenty big? This isn't a "nobody'll need more than 640k scenario, right?"
https://gigaom2.files.wordpress.com/2014/08/requirements-tab...
A choice quote:
It is not clear when a flat address space larger than 64 bits will be required. At the time of writing, the fastest supercomputer in the world as measured by the Top500 benchmark had over 1 PB of DRAM, and would require over 50 bits of address space if all the DRAM resided in a single address space. Some warehouse-scale computers already contain even larger quantities of DRAM, and new dense solid-state non-volatile memories and fast interconnect technologies might drive a demand for even larger memory spaces. Exascale systems research is targeting 100 PB memory systems, which occupy 57 bits of address space. At historic rates of growth, it is possible that greater than 64 bits of address space might be required before 2030.
Key point, and that's a huge "if". All large systems are NUMA, and trying to treat that like a uniform address space will be absolutely horrible because of the extreme latencies that arise.
2) Yes Transmeta was VLIW internally, but I see that as an implementation-detail over other forms of superscalar; either way you have a linear stream of instructions generated by the compiler, with hardware turning that into parallel execution by the CPU at runtime. Calling that "VLIW" is about as interesting as calling a modern x86 "RISC."
Seems like good forward planning but a little premature, no? These features must come at some cost. Surely you could have a re-vamp sometime around 2025?
We have off the shelf x86 machines today that can fit enough memory to use about 43 bits to address it (at least in theory - depending on DIMM availability; 1TB/2TB is trivially available on the other hand). Now add in size of readily viable storage arrays or even network file systems and expect some people to want to be able to memory map "unreasonably" large files, and suddenly we're into 50+ bits today.
It'd seem crazy to me if someone were to start planning a new processor architecture today without planning for addresses greater than 64 bits.
The Mill architecture for instance, requires each program to live in its own area of the address space, as this helps make the caches faster (TLB dont have to come before cache). This still leaves a comfortable amount of space left in the 64bit address space.. today. But maybe not in the future
It's quite possible - even likely - that low end systems will not be 128 bits anytime soon. 8 bit microprocessors are still selling in large volumes for embedded use. But we're about to see 64 bit entering phones this year, because it is becoming necessary, or at least more convenient than not.
All the evidence is that we're not heading towards a slowdown in growth in storage requirements anytime soon. If anything, the existence of super-computers that are spread over huge clusters instead of being a single, tightly integrated system indicates that there is some level of demand for systems several orders of magnitude larger than the current largest off the shelf systems even today, at the right price.
The top end of the market have increased by at least 10 bits over the last 11-12 years alone. At the current storage growth rates, we'll hit the 64 bit limit on single server systems sometime in the next 10-20 years when factoring in memory mapped IO; sooner for single-system-image clusters.
Not a chance. 64 bit is ca 16 exabytes.
Currently, the shipping volume of harddrives is about 500 million units/year. If we're generous and say that their average storage size is only 100GB, despite the large number of models in the 1TB-5TB segment, then that's 50 million TB/year, or about 50 exabytes of harddrive capacity per year. In reality it's likely much higher, and rising rapidly.
Yes, the number sounds big, but so did 1TB just a few years ago. And 1GB just a few years before that. It's not that long ago we were marvelling over even being able to buy 20MB hd's for home use. The number may sound outrageous, but my experience based on actual product availability is that we should expect a factor of 1000+ rises in storage capacity per 10-15 years, and I see no evidence to justify a slowdown.
And increases in capability causes changes in how we engineer things. When petabyte sized databases becomes possible for more people at reasonable price points, you'll see a lot of people that previously "made do" with terabyte sized databases find all kinds of uses for extra analysis etc., or simply storing more intermedia stages and being more wasteful because we can.
Purely for RAM we can survive with 64 bit for maybe a decade extra.
And it's not that big a leap: 15 years ago, I worked mostly on systems with <1GB RAM, and where to get a decent performance 1TB+ disk array, we were looking at a fridge sized monstrosity. Today, my laptop has more than that. Even my phone is closing in: 1GB RAM + 64GB storage (even some mid range Chinese phones today advertise support for 256GB SD cards, so it is already obsolete)
15 years before that again - 1985 - my machine had 64KB of RAM and 176KB floppies. A couple of years later I finally got 1MB of RAM and a 20MB HD.
So I don't consider it unreasonable to assume that we'll have systems in the 1PB RAM range and 1000+ PB storage range by 2030. That's about 50 bits for physical memory alone, or more like 60 bits to memory map that much secondary storage.
That's assuming no shared memory eating address space for example.
For most people, it will not be that relevant for a few more years, just like it took several years from 32bit became an issue on servers until it mattered for home computers, and like how 32 bit is first now becoming an issue for mobile.
But think about that for a second: We don't need 64 bit for our mobile phones. We do nothing on them we couldn't find ways to shoehorn into a 32bit address space for decades to come. But going to 64bit is convenient. It lets us shove 8GB or 16GB RAM or more into future systems and not have to think so much about memory efficiency. So we're going there.
Thus the idea that "humanity will never need more than 64 bits" I think is pretty much ridiculous, because it puts things on its head: It's not that we will "need" it, but that we will easily find ways of making use of it if we can. E.g. direct addressing all your storage is convenient in all kinds of ways. Storing all your home videos uncompressed, unedited, in 8K resolutions, in 3D, at a higher frame rate, becomes convenient when there's enough storage. Being able to write apps that can mmap multi PB monstrosity video projects instead of worrying about disk buffering becomes convenient when it's possible to do it.
It's not that long since I used to consider databases of a few hundred MB big. Now I throw around multi GB databases on a daily basis, and I know they are small for a lot of people, who deal with individual databases in the TB or PB size without blinking.
One use case for 128 bit addressing is for a single-address-space OS running on a cluster. http://en.wikipedia.org/wiki/Single_address_space_operating_...
Pointers can have 32 bits for the IP address of the hosting node and 64 bits for the address within the node, and 32 bits spare for various flags.
IBM's new openness isn't really open at all. It's just what ARM has always been doing: they allow you to pay them a lot of money so you can use their ISA in your CPU.
Still, that requires some chip maker to build a SoC around a RISC-V CPU that attains these efficiencies in the real world.
The paper makes these arguments for RISC-V:
• Greater innovation via free-market competition from many more designers, including open vs. proprietary implementations of the ISA.
• Shared open core designs, which would mean shorter time to market, lower cost from reuse, fewer errors given many more eyeballs3 , and transparency that would make it hard, for example, for government agencies to add secret trap doors.
• Processors becoming affordable for more devices, which helps expand the Internet of Things (IoTs), which could cost as little as $1
The first point is not very concrete. China has long had some of their own MIPS-based RISC CPU designs, and they are most likely to act on the transparency issue. That leaves super-cheap processors for IoT. ARM may be able to deliver pricing and value that's better than free.
And all this assumes very low friction in the form of, say, Android adding this ISA to the standard set of compilation targets for native code, and, to ART pre-compilation.
One of the challenges that Microchip has faced was that the affinity for C that the ATMega architecture had from rival Atmel meant it lost a few significant design wins (Arduino perhaps the most serious). They could use RISC-V to try to offset the Atmel SAM series. But other than that I don't see the motivation for folks to not use ARM, granted a full processor license would be expensive but if you are looking at volumes where that would be an advantage, it isn't that expensive.
I feel like the fabbing cost is so high, that at that scale, the CPU license fee is really nothing.
Correct me if I am wrong. This is coming from the mind of someone who knows next to nothing about how hardware is really made.
Perhaps the next edition of his textbook (bible for computer architecture) should use RISC-V, it would probably help as a learning aid and spread the gospel about RISC-V.
Of course, it's possible that the current edition of the text already does this.