10nm versus 7nm
semiengineering.com
semiengineering.com
I think your feeling is wrong.
Mind you, performance improvements didn't ended, we can still find other ways to do things...
But 7nm is the smallest physically possible transistor (well, theoretically it is 5nm, but that one can't be actually manufactured).
So, you hit 5nm, then what? Your only choice would be make chips larger... but if you do that, you limit your maximum clock (example: a 5cm processor, if it allowed data to travel at light speed, and had perfect cooling, would still be limited to 6ghz)
Or make them taller, but then you can't cool them easily.
That is the problem, Moore Law, as in transistors per chip, ended with 7nm, there is nowhere to go after that, the only thing you can do is add more chips and make them communicate with each other (ie: multi-cpu... see what AMD has been doing in their recent research, trying to figure how to make an API to allow infinite amount of GPUs in the same calculation in a way 100% transparent to the developer).
But making more chips is not improving Moore's Law.
The only thing that can be done now, is invent something else, silicon transistors are done for, you can't improve them much anymore, you can only change the tech (use something else other than silicon transistors) or find novel ways to use current silicon tech (better architectures, multi-chip systems, better code, etc...)
Haven't you answered your own question? Your other points are about impossibility - the one I just quoted, about making chips taller, is about difficulty.
If I were the head of Intel I would say "Well. Hmmm. All right. I guess we'd better start making them taller."
Moore's law is not a question of "how hard can it be". As you've pointed out, on the current approach it's a question of "It is physically impossible for us to increase transistor count without dropping the clock rate or moving out of roughly a plane. No thanks to Einstein. (My favorite link on this subject: https://www.google.com/search?q=c+%2F+4+ghz )
You say "Or make them taller, but then you can't cool them easily." Well then, make them taller and cool them with difficulty. Nobody said it would be easy. Off-hand you could:
-> Immerse the taller chip in insane cooling
-> Fabricate little duct work between layers and liquid cool at high velocity
-> Figure out a way to reduce heat waste usage, e.g.: different substrate, not silicon (pretty insane)
-> Possibly combining some of the above points, make it a superconductor.
These are insane ideas from me. But then again, I'm not the head of Intel and you've heard a few minutes of thought on me on this topic, where I'm not a chip designer.
"When you have eliminated the impossible, whatever remains, however [difficult], is your [roadmap.]"
And no, we can't just assume the head of Intel will conjure up a materials revolution solely to adhere to an industry observational law.
https://en.wikipedia.org/wiki/Moore's_law#Moore.27s_second_l...
This talks about an exponential increase in R&D costs which is expected by Moore. With stuff like cloud computing, I don't see the need for denser and denser CPU's going away any time soon.
It was in 2005. I lost contact with these people, I don't know how they do now.
Easy. Just have multiple synchronized clock sources.
So first it's about transistor density, hence the relation to having smaller and smaller transistor. And also and very important, it's about cost. You could go 3D to increase the amount of transistors per area, but that requires more masks/operations and impact the yield, and increases cost. If in the end the cost per transistor increases, then it's a density increase but it's not advancing Moore law.
In other words, Moore's law can be broken in two different ways: 1) we can't improve anymore the transistor density; 2) we can improve transistor density, but not in a cost effective way. If you want more you need to pay more (vs. getting more for less money when riding Moore's law).
But the more interesting effect for me has been that Moore's law has slowed down but the exponential growth it had given computers allowed them to speed past even our more amazing ideas. We haven't reached Jim Gray's "smoking hairy golfballs"[2] yet but we also don't upgrade our machines nearly as often. When I bought a new laptop, I bought it for its stylus experience, before that I bought a new lap top for its retina display, and two years before that was the last laptop I had purchased because the one before it wasn't powerful enough to do the stuff I needed to do (most e-cad, m-cad, and code development)
So how does the world change? Some software places might consider writing software that doesn't change for change sake. Every year my m-cad vendor wants me to upgrade to the latest version and every year its the same story, tweak here, bug fix there but no new features. Support for new things the OS has imposed sure, but not real features.
This has been great for Linux folks, where the same computer has been fine for 5 - 6 years now. They can look forward to actually seeing a computer fail due to old age rather than forced obsolescence. It will be harder to maintain a distro that runs on a decade's worth of hardware changes. And that latter is the killer.
So computers aren't changing but as I mentioned before the I/O around them is, better screens, better peripherals, different ways to do storage. Since that is the only way to force upgrades these days, I expect to see a lot more of that.
[1] From the article -- “Because Moore’s Law has broken down and you no longer get gains in all areas at the same time by going to the next node, each foundry customer will have a different strategy, depending on which parameter is most important,” according to a source at one large customer, who asked not to be named.
[2] http://research.microsoft.com/en-us/um/people/gray/papers/Su...
Inconsequential feature creep doesn't happen for changes sake, it happens because people want to keep making money.
That I can personally attest to:
1. Simulation games are mostly stuck in power available since the 2000s, many common game algorithms can't be made parallel, and several algorithms increase in computing cost faster than processors increased in power since then (example: Astar, its computing cost rises quadratically with the search space, and the search space when you want to increase a game map size ALSO increases quadratically, thus to double a city simulator map size, you need to run A* 8 times the processing power... if you pay attention city simulators since SimCity 4 release tended to make instead smaller maps with more miscellaneous simulations).
2. CAD, CAE, etc... In this case the recent advances, including in GPU design made those software much faster, but they are still rather slow, if you need to compute a complex project, or even a simple project with greater precision, you still need some crazy powerful computer... As personal anecdote: I am trying to design a PC case for myself, decided to try my hand with CFD, first to test I made just a empty box with a "fake" fan, that is just a number saying how much air get inside the box... got a good result in 3 minutes. decided to test what happens if you have 4 fans side by side... well, this creates turbulence, thus need a more detailed grid to calculate that, the grid generation alone took 20 minutes, I decided to make half-assed simulation and the result was mostly nonsense. I decided then to test if I could model the entire fans, including their blades, housing, etc... I got a already done precise model of a fan, tried to generate a grid... and my computer overheated after 5 minutes using an i7 at 100% Of course, I knew I was "abusing" things, still, it became obvious to me that CFD for example, for anything "serious", is mostly impractical, if I had infinite money, I would still need to setup some kind of huge cluster or super computer, and the code to run CFD in that would be very complex, the thing is possible, but still impractical, even with infinite money.
Below 28nm their numbers become really deranged. This overstatement is IMO peddled by big name chip companies trying to convince analysts that FinFET nodes give them a moat. In reality the end of Moore's law means they'll face vicious competition.
Of course nothing prevents you from spending huge amounts of money. Buy an ARM architecture license - there goes $30M and that's before you pay for developing your own ARM-compatible CPU.
For argument's sake, let's pin the tapeout cost at $5M and EDA costs at $1M. Everything else is engineering. The reason these projects are expensive is that the chips are complex (because the have 10B transistors!). It has nothing to do with technology scaling (except for the fact that smaller nodes lets you integrate more transistors). I can design a memory chip (with repeating structures) in any node for almost nothing (..besides the tapeout cost). If they would have costs have escalated from $50K at to $5M at 14nm, that would be reasonable. The $250M number is utter nonsense....
Some more background:
http://www.adapteva.com/andreas-blog/semiconductor-economics...
Yet, two of your articles suggest me doing a Snapdragon or something without buying third-party I.P. or RTL could easily cost the numbers in the article on 28nm. Or the other nm. Intended audience of those big nodes will also mostly be companies making stuff like that where they can afford development costs. Do shuttleruns even exist for 16/14nm?
http://www.adapteva.com/andreas-blog/a-lean-fabless-semicond...
Adapteva is pretty clever. Yet, they don't tell us whether they reached $2 million at the 16-core model with 64 costing much more or if that was both. The write-up also says they used third party I.P.. It implies they outsourced as much stuff as possible. You can bet it was mostly standard cell rather than custom in those. They used the shuttle runs I often advocate which are available on their chosen nodes. Biggest part, they recovered this by selling to niche markets that didn't care much about price. Try that in mass market outside being Apple you will fail.
All quite different than doing a whole SOC by yourself on 28nm. They even admitted complex SOC's can cost big money:
"Complicated SOCs are indeed very expensive to develop and there are few if any examples of complete products that have cost less than $10M from start to finish. "
So, now we've shifted from $2mil to $10mil that quickly. Gotta wonder what in-house development of full SOC a la Snapdragon with bells and whistles at 28nm costs with only packaging or PCB's outsourced. Is it maybe $10-30 million? ;)
There is no mystery here. EDA costs, tapeouts, engineering costs are all known. Tell me the chip you want to build, I'll tell you how much it will cost.:-)
People seem to have a hard time grasping that the term "SOC" is meaningless when it comes to predicting cost. It's like comparing the cost of a Ford Focus and a Ferrari. Sure they are both cars...
http://www.adapteva.com/andreas-blog/semiconductor-economics...
Full mask tapeout was done later.
And good luck with "The Brain." Saw that on your website just now. Funny that I was thinking of proposing to someone a prior product for revival that did NN's with 256 8-bit cores on 0.5 micron with good results. Then, ABC Thinking Machines Super-Connector had badass GA results back in the day with 64,000 simple cores. Your design looks like a Thinking-Machine-on-a-chip level of cores with cores being more versatile. Not 64-65k but close enough. :)
Also, the design man-hours quoted in the article must be when designing from scratch, and may or may not be accurate. In the context of moving to a new node and iterating on a previous design, the quoted numbers are junk.
If we get to the point that TSMC has the same technology as Intel and IBM and it never changes then some hardware hackers will eventually publish something like a 250MHz processor that uses 10 watts on that process, which will be "good enough" for some embedded devices to justify a fairly large production run. Then you have a community and people make incremental design improvements for the same reasons Google and others do for Linux and in a few years the free hardware design is above 2GHz and below one watt.
Open source hardware refers to the silicon itself using open source designs.
[0] http://www.eecs.berkeley.edu/Pubs/TechRpts/2016/EECS-2016-1....
My advise is still the same, though: put Gaisler's Leon3 SOC w/ key I/O on fab at 180nm-90nm using shuttleruns like Tekmos. Its been GPL with no action forever. Already has OSS software on it, too. SPARC compatible trademark only costs $99 if you wanted to be official.
So, if there's such a will or demand, where's our open-source SOC that uses Leon3? It's been one MPW, packaging, and PCB run away forever. Still don't have one...
Anyway, you know what I wanted: open-source, 28nm FPGA w/ key I/O already included so I could synthesize my own CPU's and such on it. Verify one design in hardware, worry about the rest in simulation. You never got back to me on that one. Whether it was a good idea to solve most people's open-hardware concerns or what it might cost?
(I'll say a Virtex-6 amount of slices or memories with USB 2.0, DDR3, PCI-E, SATA, and 10/100/1000 Ethernet onboard. Should do.)
I've been doing a project that's a combination of java play, with its NIO and actor model approach to parallellism, and apache spark, with the rdd approach. Either solution lets you scale horizontally without having to manage locks and threads yourself and I like both approaches, but you can tell it's still early days. These approaches need to mature, standardize, and then percolate throughout all languages and frameworks, but eventually they will and we will be writing parallel code by default.
Alternative theory: there will continue to be bad programmers, but there will also be good money in compilers that can deal with it and turn crap into optimized goodness. Midas compilers.
I also suspect a few new language concepts (like Rusts ownership model) will need to be invented as well, to enforce certain optimizations with minimal cognitive load on the developer.
No amount of compiler optimization is going to fix algorithmic failures.
Optimizing via better algorithmic choices is a semi-deterministic process performed by a human. Why can't a computer do it too?
Code is just your way of telling the computer your intentions. Once it can understand those deeply, well, it can choose a better way to achieve those intentions.
Demand for new data models is insatiable and for each one, a software ecosystem can develop around a similar degree of automation, tuned towards the specific domain. It's a very, very black-boxed future.
- Do we need much faster hardware for most tasks people are actually trying to accomplish? I don't think software is necessarily becoming more complex. It could be simplified it's just rushed out the door. I think complexity has plateaued. A common run of the mill usecase of application-layer software, a CRM, isn't limited by CPython (and definitely not by PyPy) today and there's no reason why it would start tomorrow.
- Something new will come along for hardware. So it actually is inevitable it moves forward. This isn't anywhere near the end. It may stratify hardware. If you truly need more performance, it may be very expensive hardware. The rest of us will rely on servers running on 7nm chips for ages. In connection to this point, the machine I'm typing this on is from 2009. A Q9450, it's still more CPU processing power than I need. I think that says a lot. The only reason I want to upgrade is for less power draw because that means less heat produced. If I needed more power for video, 3D rendering or compilation I'd use networked processing which would annihilate any local big iron I could put up in my basement.
- More development platforms will dump everything that isn't LLVM based. Significant gains could be had if say, Python was reimplemented on top of LLVM.
- For cases where the everyman company needs more power, they won't go for the chips made out of brain tissue or diamonds. They'll use services like AWS to gleam the power needed which will have the more expensive hardware on-top or massive swaths of 7nm CPUs to utilize.
- Chips will increasingly be built for specific tasks. Dropbox will keep having custom designs built by Intel for their specific workload. The general-use CPU will become less popular in the server room but still there for the small to midsize companies to lean on.
So while there's some individual points there, to add the always beloved car analogy. I've downgraded my car from a 325HP model 10 years ago to a 145HP model that I've been using since. Both got the job done within their environment and I suspect at worst, chips will simply become a commodity closer to modern day cars. The advance won't be more HP or better brakes, those are mostly maxed out in capacity. The paradigm shift is where the advancement will be (electrics for cars). But it's been settled for a while that either way, I really don't need over 150HP for what I'm doing.
Good news is many accelerators work well at old 90nm and 45nm nodes based on my reading of academic papers. Even startups or FOSS projects can still get plenty of mileage out of the cheaper nodes. The Core Duo 2 I'm typing this on... still running ever bloated web well... was taped out on 65nm.
Consider a simple example: CPUs have had caches basically forever. Caches are the epitome of independent development, because they require modest performance improvement with no effort. (Advanced users and compilers can of course put in more effort to get more improvement, but the maximum improvement is limited because CPU vendors have been reticent to allowing direct user control over the cache.)
What if these caches were instead scratch pads? Frankly, this sounds like painful way to program, especially in traditional languages, because now everything will have to be explicitly moved into the scratch pad. Maybe high-performance users will benefit, but the vast majority will not.
But we can reinvent the language as well. Let's propose a different programming model: tasks (blobs of code) with explicit data usage declarations. The rule is that a task only runs when its data is available. Thus the processor can copy data into the scratchpad asynchronously, and only run the task when all the data is there. Suddenly, there are no more cache misses! We just had to give ourselves permission to modify the programming model too.
Well, actually this is not a hypothetical programming model. I've described Legion, a programming system from Stanford [1]. A software-only implementation of Legion was able to scale to 1000s of nodes on a distributed system [2], so we know the basic approach works. Even better, the Legion folks described a type system which allows you to hide the gory details behind an easy-to-use apparently-sequential programming language [3, 4].
So far, these results have focused on software only. What would it look like if we were able to change the hardware as well? To answer your original question: I think this will require more expertise (and more diverse expertise) on the part of the system designers, but if we do it right, we can still design a system which is easy for everyone else to use.
That's where I think the biggest wins will be in the next 50 years.
(Disclaimer: I work on Regent, the Legion programming language)
[1]: http://legion.stanford.edu/ [2]: http://legion.stanford.edu/pdfs/legion-fields.pdf [3]: http://legion.stanford.edu/pdfs/oopsla2013.pdf [4]: http://legion.stanford.edu/pdfs/regent2015.pdf
I don't think we are going to see that. Most use cases are not performance constrained (ex: a simple app or website). What you're talking about applies to cutting edge graphics, machine learning, OS and compiler design and a few other computationally intensive domains. Our computers are going to be at least as fast as those of today, so they will be able to run even the sloppy software of today with no problem other than (maybe) excessive power usage.
I think we are going to see an even greater shift towards asynchronous and parallel programming and design of specialized hardware for speed and low power. Also, we're going to have a much more complete AI toolbox than today, containing ready made image, video, audio, text, dialogue and behavior modules. For example, there is no need to reinvent the wheel when we already have a neural net that can distinguish tens of thousands of objects. Once it's trained in one place, it can be easily copied and run everywhere else, like an app.
More transistors, larger chips.
We have bought so much physical space going smaller and smaller - and I think a lot of people will happily pay for a thicker iPhone if it's not more room allocated for a larger battery.
Eventually - when that becomes ridiculous - we will write better software. More likely we'll buy an extra couple years after the end of Moore's Law and Intel will make some unexpected breakthroughs to go smaller again.
Color me cynical?
This heat problem can be solved if they make it a priority. Until now they've been able to get away with ignoring it as new nodes have kept them competitive. If that changes, which it is, a new approach is required.
However, as the new EUV machines are being put into production. It does not make financial sense to still make the 10 nm chips. As buying the new machines and infrastructure is a huge investment, you can understand that some companies are focussing on the existing 10 nm machines (maybe focus on lower end of the market). The problems with EUV are mainly starting problems with using different technology, IBM has already made 7nm chips and ASML has 5nm and 3nm in their roadmap[1].
[1] http://arstechnica.com/gadgets/2015/07/ibm-unveils-industrys...
My mental model of this process is completely wrong. I thought they'd have a Verilog code base, and moving to a new process was basically a recompile.
What are they doing for 100/200/300/500 man-years, and why do the smaller processes take longer?
But how does it Apply to a whole SOC ? you basically go to the critical sections and optimize them by hand, better than the synthesis tool ? Or just supply some basic components to the design team , and if so isn't it the job of the fab and the PDK(process design kit) ?
What the OP showed you applies only to the Mixed-Signal IP, when you have a Digital part (developed in Verilog) and an Analog part (developed in tools like the OP showed).
When you start building a SoC, you will build it like you would be building a Lego: you contact the various IP providers (Synopsys, Cadence, ARM, Imagination technologies, etc) and you start buying IP for the CPU, for the HDMI chip, for the memory (DDR), and so on, and then you put them together.
The IP providers are responsible for making sure that you receive the layout of the IPs working for a certain process node. If you are just building the SoC, you won't need to look at the schematics or at the layout.
For custom digital design and SoC, you start with RTL.
In custom design, you read the RTL and create a schematic using logic gates which implements the RTL spec. Generally, you have something to start with from the previous project, unless your block has seen massive changes. Then you floorplan the design, placing the interface and gates, and drawing the routes of the critical nets. Then you perform static timing analysis and many other checks to converge the design. This is a very iterative process, especially as some of your timing analysis depends on other blocks.
(I don't have experience in analog, other than working closely with a few analog designers as a customer for some mixed-signal design I did a few years ago) Analog designers use similar tools to create schematics/netlist and physical design. The verification process is much different, with a lot of work going to developing simulations to ensure the circuit works as it needs to in a wide variety of conditions (process variation, operating voltage, and temperature). There's probably a lot of other stuff as well, but I am not privy to all they do.
For SoC design, you compile RTL to netlist, using a logic synthesis tool. The synthesis tools are smart enough to approximate the physical design of the block, and you can refine your recipe to guide the synthesis tool to converge static timing in the netlist itself. Once that's complete you move to the automated place and route (APR) tools. Here you will spend some time and effort to floorplan your block: mostly placing of large sub-blocks and your pin interface to the rest of the SoC. Once you have that done, you'll go though placement, clock-tree synthesis, and routing, refining the recipe that will converge your design. Again, you will iterate on the design to converge static timing analysis and other checks.
In both these cases you will also need to converge the physical design to meet the design rule checks. The automated tools available in both environments take the design rules into account, but the complexity of these rules is not fully understood by the automation, and so user intervention is required. Depending on your group, you may have a mask designer available to help with meeting process design rules, but most groups require the block owner to do the majority of work in this domain, with the mask designer available for final clean-up at tape-in. Hope this helps.
But you need to keep in mind the analog designs. While a typical CPU, a typical GPU or a typical cryptographic IP chip (for a example) can be developed in just Verilog since they are just pure logic, when you move to the Mixed Signal IP, you can't use only Verilog, you have to manually design parts of the chip (creating the schematics of the designs at the transistor level and converting them to layouts).
Examples of mixed signal IP are USB, HDMI, MIPI, Ethernet, PCIe, and so on.
We don't have tools capable of synthetizing really fast PLLs, DCOs, Band Gaps, most of the high speed analog blocks.
But I think the article was exaggerating on the "50 engineers will need 10 years to complete the chip design to tape-out". I work at Synopsys and we are already working on the latest process nodes (10nm, 7nm, 5 nm) and we already have some solutions ready for silicon and it didn't took us 10 years.
All these articles make it sound so dramatic that it's nowhere near the reality. The process of designing the analog blocks again in 10nm, 7nm, 5 nm and so on, it's exactly the same every time. It is very much like porting your C++ code to Python: sure you have to do it from scratch but the process and the fundamental concepts are exactly the same.
It's not rocket science nor quantum physics, it's easier than it sounds. The main problem is the huge cost of the tools and of the technology from the foundry. If you are past that, you just need time and some knowledge in micro-electronics.
Don't be fooled by these over dramatic articles, the hard part of digital and analog design is not getting the knowledge to do it, it's getting the money to have the proper tools to do it.
The companies that provide IP are watching the market splitting into two: the companies that want the latest process node to show off technology (for example, Apple, Samsung, Huawei, etc) and the companies that want older but stable process nodes (IoT companies and the automotive industry).
When most of the articles speak of the newest process nodes (for example, the 5nm process node) they only touch the economics part of it: "it's expensive, it's costly". But they don't mention the many problems that come with it: it's unstable, there is much more current leakage which increases power consumption unexpectedly, it is not suitable for long longevity and the architecture of your analog blocks have to take into account all these problems which don't exist in the older process nodes.
It is possible to produce solutions in 5 nm, but the question is no longer about if we can do it or not. The downsides os these nodes go beyond the engineering resources and the economic resources that are needed to achieve them.
That is why the automotive industry is not interesting in these kind of nodes. They want the stable and cheap nodes like 28nm and 40nm, they know that these nodes are well tested and have long longevity.
We are also reaching a point in which we already have enough advanced technology for our needs, we need to start to get rid of the need of reaching the next process node in order to have better hardware.
When you talk about cost, there are fixed NRE paid once per chip design, and then marginal cost per transistor (just including production). To get the full view, it's logical to consider a "fully loaded" cost per transistor, where you also amortize the fixed NRE expense taking into account the expected number of produces chips for a given design.
And that's where the increasing NREs bite and make a difference. Even when the marginal production cost per transistor decreases, to benefit from that cost decrease one must have sufficient chip production volume to amortize the NRE and still benefit from a cost decrease. With exploding NREs on more advanced nodes, this requires ever higher volume production to benefit from lower costs on more advanced nodes.
At this stage, if you're a medium chip vendor Moore's law is already over --- because with small volume the NRE leads to an increase in total cost per transistor for newer processes. Vendors like Broadcom and NVidia complained publicly about the end of Moore's law, so we're not talking small volume either.
For very large volumes, it may still be going on. But that's hard to tell. At some point, the cost may go up even for the biggest (Apple, Intel, Qualcomm) but they may still push forward a bit for improved performance --- meaning higher power efficiency --- as long as there are enough people willing to pay for it. It's no longer Moore's law (where you get more for less money), but it's still moving onto more advanced processes. Still, even assuming enough people are willing to pay more for improved performance (a big assumption IMHO) there are quite formidable challenges ahead.
Also, you forgot to mention that memory is simply not synthesizable, so on-chip cache for example needs to have a full layout from scratch.
Keep in mind that this is the case for Intel and not ARM. The latter uses standard cells, which means it's basically all synthesized.
Source: my professor.
As you said, the EDA companies and foundries always came through. It was hard work, though. Work OSS will have to duplicate at least on EDA side.
Even if the EDA companies were to provide all of the tools for free, the startups would be facing the high costs from the foundries.
But to tell you the truth, there is not much incentive to design in the newest node if you don't have money to spend. The technology not only is expensive but it is also unstable.
Or the other option is to build multi for chips like marvel with their Lego concept offered, all at advanced nodes, but most of the dies after reused from previous designs.
In addition to this, manufacturing design rules get more stringent and complicated, which result in more headaches for the designers. This is not only from the HF standpoint but things like narrower voltage windows, high power dissipation from all those extra transistors, etc.
One should know that for processes less than 90nm the shape of wire starts to play - delay begin to depend not only on the length of wire, but on the number of turns as well.
All this means that the smaller the process, the more time and computing resources will be spent on post-synthesis, place and route stages.
Even if we constrain ourselves to pure "digital" designs that somehow magically meet timing perfectly there will still need to be significant changes in architecture to the chip.
For example, presumably you went to the lower process node to up the transistor density. This will have serious impact on how charge and heat flow in the design so those systems are going to need a serious overhaul no matter what the RTL was. This is going to impact the layout and the area constraints which is likely going to ripple into the RTL hierarchy.
That said, the article's estimates all seem a bit high to me but it's impossible to be certain as they don't specify any of their terms. Is, say, a pipelined ADC in 28nm CMOS a "chip" in their book?
What is a 1D layout? Surely it must be something of a misnomer, as the only way I can visualize it is as a barcode, where each material you place has to span the entire width of the chip - and that seems utterly worthless.
I would have naively thought chip design is fairly automated, but a 9x increase makes is sound like every gate and path is delicately hand-stitched like a Persian rug.
So it's far from fully automatic. There's a lot of work done on the tools, but there are also a lot of additional constraints to deal with with each new nodes. It's a real complexity explosion, and this is what makes designs on advanced nodes so much more expensive. This a highly simplified description but hopefully enough to get a feeling for what's going on.
Depending on the complexity of the chip, the backend process takes several months. It's an iterative process. A reasonable complex chip has to be split into several partitions. The size of the partition determines the turn around. In the designs that I have worked on, we try to limit the turnaround time (RTL->GDS) to be one week for the larger partitions.
I would change your statement to say that the EDA vendors have created tools that allow Physical Design engineers to address the design challenges. It is still a gruelling, iterative, painstaking process.
The lasers have a limit of how small it can get and the smaller the chips get, the more you have to re-etch with different materials (for the laser to go through) to get it to where you want it, which also increases the amount of defects that will happen on the chip, driving up the testing costs as well.
28nm was the limit with a single pattern design. 20nm required a double-pattern and certain 10nm/7nm designs will need triple or quad-patterns which is super expensive and slows down the process.
That's what the article meant with this statement:
> But at 7nm, optical and multi-patterning are simply too complex and expensive, at least according to Samsung. So to make 7nm cost effective, it makes more sense to wait for EUV. In theory, EUV can simplify the patterning process.
You can find more information here: http://www.anandtech.com/show/10272/samsung-foundry-updates-...
This might help as well: http://www.extremetech.com/computing/160509-seeing-double-ts...
But that's for CPUs. Memory, especially mostly-inactive memory such as flash devices, still has a few iterations ahead. Memory devices can tolerate and recover from bad cells and bad rows. There's also the option of going 3D and stacking memory. That doesn't work for CPUs because getting the heat out of the middle is very tough.
The other big problem is cost. Fabs have become multi-billion dollar projects. The article says that designing an SOIC for 7mm costs upwards of $200 million. (Why? Design tools? Mask making?) If cost per gate declines, we can still get more compute power by making bigger parts with more CPUs. Single CPU performance probably isn't going to climb much more, though.
Fortunately, most of the things we want to do in machine learning and AI can be done in parallel. This problem isn't going to keep us from getting to AI.
> Samsung, for one, plans to ship its 10nm finFET technology by year’s end.
> TSMC will move into 10nm production in early 2017
> Intel will move into 10nm production by mid-2017
>“Not all 10nm technologies are the same,” said Mark Bohr, a senior fellow and director of process architecture and integration at Intel. “It’s now becoming clear that what other companies call a ‘10nm’ technology will not be as dense as Intel’s 10nm technology. We expect that what others call ‘7nm’ will be close to Intel’s 10nm technology for density.”
>It wasn’t always like that. Traditionally, chipmakers scaled the key transistor specs by 0.7X at each node. This, in turn, roughly doubles the transistor density at each node.
>Intel continues to follow this formula. At 16nm/14nm, though, others deviated from the equation from a density standpoint. For example, foundry vendors introduced finFETs at 16nm/14nm, but it incorporated a 20nm interconnect scheme.
Yes, that's what a lay reading gets you - and why I was disappointed with the article.
Before SSD it was common to see a CPU wasting >50% of its time waiting for IO. Now you rarely see iowait being a problem.
I don't understand: 60% of what? What embedded software?
Why is it cheaper to design a 7nm chip for a given performance than to design a 10nm chip? The article mentions the yields on newer processes are lower, so I would've expected them to be the performance option not the cost option.
It will take a really long time to actually squeeze out CPU optimization. We are still contending with most software not even properly scaling to four cores, let alone well implemented job queue systems for arbitrarily scaled cores. And once the chips start dropping in price as the node standardizes and companies can lay off the intense R&D investments, we can start introducing more heterogeneous compute clusters to drive performance cheaper even when the silicon itself is stuck until graphene takes the market.