5nm vs. 3nm
semiengineering.com
semiengineering.com
Since Dennard scaling ended about 15 years ago, these new devices will probably run hotter, adding to the dark-silicon, eh, let's call it a situation. It's a problem from the traditional point of view where you expect to be able to use all your hardware all the time, but maybe it's an opportunity if you see it as a chance to handle burstier computational loads or to pack a greater diversity of specialized cores onto a chip. But of course that increases both design costs and the complexity of programming the device once it's been fabbed.
The impending collapse of Moore's Law has thus been delayed for two or three years, or softened anyway, but the appetite for computation due to deep learning continues unabated. Since scaling Jack Kilby's planar process down is becoming increasingly uneconomic, this would be a good time for a non-planar process to emerge — a trillion squares occupies a one-million by one-million area, while a trillion voxels is only ten thousand by ten thousand by ten thousand, a scale a hundred times larger and therefore less demanding on your fabrication processes. You'll probably need some plumbing in there for coolant. I don't know of anybody working on this, surprisingly.
No commercial projects that I know of, though.
Source: I work on 3D chips for deep learning
Note that exponential term for probability of failure - 10000 layers means 10000N etch steps but P^10000 turns a 99.9% yield for 1 layer device into a .004% yield for a 3d 10k layer device .....
note: by 'layer' here I'm including all the metal/poly/passivization layers in a current device - also after each layer you'd have to leave space to plane the finished layer flat as a base for the next - which means insulation between them deep enough to fill in all the gaps in the previous layer plus vias - and as others mention some way to pull heat out of the core of what has historically been called a "hairy smoking golf-ball" CPU
If you have a certain probability of failure in each transistor (I know, I know, ICs don't consist entirely of transistors, but let's simplify) then making more transistors by stacking layers needn't lead to more failures than making more transistors by making larger-area 2D chips.
Existing fabrication methods already produce some failures. In at least some cases, the way that's dealt with is to make the hardware able to cope with some bits not working, and then e.g. you can turn an 8-core chip with a defect in one core into a 6-core chip that you sell for less money, or a memory device with one block of memory not working into ... exactly the same as all your other devices because you already budgeted for a couple of blocks not working.
So I think the p^n issue is illusory: n doesn't get larger just because you're making your device in 3D, it gets larger because your device has more stuff in it, so the issue is large devices not 3D ones, and there are already known ways to deal with the fact that large devices often have defects. 3D devices might well have a higher failure rate per unit of stuff on them because the processes would be more complicated, because there are connections in more directions, because of thermal issues, etc., but that's a matter of increasing p, not of increasing n.
Not all defects are local. Sometimes you have to throw out the entire wafer. The rate and kind of defects is specific to each step in the process and there are many steps.
To make a wafer of modern semiconductor devices is a process that takes up to 4 months! If you double the number of layers you're already looking at 8 months and without some improvement in the reliability of each step you definitely are squaring the probability of producing a defect free device.
Making a device with ten thousand layers would require completely different technology, since we can't wait 40,000 months for the chips to arrive.
However we're talking here about building 3D stacked devices instead, to get speed we're going to build cubic CPUs (to reduce speed of light delays), not piles of single layer square CPUs, maybe 100 layers deep - certainly we can disable bad CPUs but that still means we're getting an error rate at P^100
I have a random question. Imagine this hypothetical scenario: semiconductor manufacturing comes to a halt completely, and new physics and technology is still decades ahead, can we further improve the practical performance of general computation by innovation alone?
I imagine, instead of riding on the train of Moore's Law, more resources would be invested to design optimization and R&D of new architectures, e.g. faster FPU, faster pipeline, etc.
Also, previously discarded and ignored ideas may be implemented again, and delivers some real advancement. While exponential growth is not possible, at least linear scaling should be the case. For example, computers without clocks. Based on this comment on HN, https://news.ycombinator.com/item?id=19554248, one major obstacle of clockless chip is the entire VLSI toolchain is designed and optimized for synchronous logic, and this can be changed if the industry invest some serious resources. Another example comes to my mind is high-level programming based on hardware, e.g. Lisp machine.
Does my imagination make any sense?
we can improve the performance of average programs by a factor of ~100x simply by more careful software engineering. Don't use python or javascript, optimize code for cpu cache hits, use more cores effectively, etc.
However games, maybe only 2x if anything.
It is clearly a rant to me! Don't get me wrong, I'm not a huge fan of JavaScript, but discussing about useless programs is simply not what I'm asking here. I'm not interested in how Google Chrome can be 1000% faster.
> People expect to write whatever they want and have it run faster. 1% of the transistors in a CPU are actually running instructions, the other 99% are trying to keep it busy.
I agree, now, your perspective makes the talk on software interesting.
Recently I've seen an article C Is Not a Low-level Language (https://queue.acm.org/detail.cfm?id=3212479), in the article the author argued that the in-order, synchronous, sequential execution model from the PDP-11 heyday is outdated. However vast majority of programs (i.e. C programs) are still written based on this model, so a CPU must use a lot of resources to dispatch these sequential code and introduce countless transparent optimizations (e.g. ILP), to make existing sequential code faster, on the other hand, the compilers are becoming monsters because they must be as intelligent as possible to understand the algorithms in a program and rewrite them automatically for optimum performance on a modern CPU. As a result of this disparity, the capabilities of what the hardware can actually provide is often underutilized.
The author purposes that we should try discarding the PDP-11's classical "in-order, synchronous" view of a program, and try developing new programming languages and models that designed with the capabilities of modern hardware in mind, such as low-level parallelism to eliminates this disparity. So the CPUs can focus on what they are good at with less overhead of dispatching the instructions.
Your basic point is correct. We have actually got 20%+ of speed ups in the last 30 years from better algorithms; not just better hardware.
Many programs are I/O limited and wouldn't see 100x from more effective CPU use (eg., Word).
Others need complete re-writes (eg., browsers) and are typically heavily optimized already.
I'd say we could squeeze another generation of Moore's Law just from software design; and another still from algorithm research.
I don't think it's really an exaggeration to say that optimization of the programs can lead to 100x performance improvements. Sorting algorithms, for instance, improved dramatically from the first computer algorithms in the early 50s (yay Bubblesort) to QuickSort in the 1960s to newer algorithms in the 1980s/90s like Melsort, and even algorithms being written today, like Neatsort: https://arxiv.org/abs/1407.6183
And the thing about algorithms is that for certain problems and large n, they can quickly blow up. Something with just a slightly better big-O could easily mean a 100x better performance at large n.
I agree that many problems have experienced super-Moorean algorithmic speedups, but the examples I'd point to might be solving large linear systems (successive over-relaxation), integer factorization, suffix-array computation, and all kinds of search problems, especially including SAT and SMT, rather than sorting of records.
I have a long list of encountering software doing things like reading a billions of rows with getline() in C++. That can be a problem [1]. It doesn't matter how good a heapsort is if it's blocked on some architectural primitive and we'd do better to teach to that optimization in CS 101 than sorting.
1. https://lemire.me/blog/2019/06/18/how-fast-is-getline-in-c/
As with all things, it depends on what your needs are and what you are doing. Modern hardware is pretty damned good, and most apps don't need more than what we already have. Or have had for a decade or more at this point.
Eventually we'll have nanotube transistor or photonic or spintronic or nano-electro-mechanical or quantum computers. There are firm physical laws saying what the limits of efficient computation are and we haven't reached them yet.
Why don't we move to analog computing for DL? This application seems well suited for it.
Also, Moore's Law is dead by years now. Not only process improvement became subexponential and much slower paced, but transistor density isn't growing much and the core of the Law that is cost/transistor has increased with the last 2 or 3 iterations, instead of going down.
Now I know SMIC since its creation, but afaik it has always been several nodes behind leading edge, and never scored any major contract. Without any large production contract, how does it get enough experience to even get to 7nm? UMC, their partner, pretty much gave up on 7nm already. TSMC has resisted all attempts at espionage. Is it from Samsung?
a) China is turning inward, maybe they're getting some state support, monetary, political, or otherwise; or at least benefiting from a lucky coincidence.
b) SMIC has been credibly accused of misappropriating TSMC secrets in the past, and settled.
c) I think a number of major Chinese manufacturing companies invest resources and money in them, possibly as some form of insurance policy.
d) The Chinese government does procure armaments, maybe SMIC does manufacturing they've been asked not to talk about.
This seems to be the most likely factor.
As we see, SMIC has already withdrawn from New York Stock Exchange entirely [1], and purchased a 7nm EUV lithography machine from ASML for $120 million [2]!
[1] https://www.scmp.com/business/article/3011737/chinas-biggest...
[2] https://www.anandtech.com/show/13941/smics-14-nm-mass-produc...
It's already increased a bit in the microprocessor world with the rise of GPUs, but that "just" puts supercomputer-style vector processing in SBCs and laptops. It isn't as new of a concept as systolic arrays, for example; those can be fast even on slow hardware if you pick your problem right, just like how GPUs are only speed demons on certain tasks. Well, if we can't beat problems to death with increasing scalar speed, it might give us more incentives to design ever-more-specialized chips which solve specific problems extremely effectively even if they're comically useless for general-purpose computing.
For those not immediately familiar with the terminology: "scalar" is Single Instructions, Single Data.
A big problem is heat. But stacked designs produce less heat overall, just in a smaller area. That can actually make heat dissipation more centralized and more efficient (due to higher thermal gradient) so I think it's a solvable problem.
I do agree, the longer term will lead to stacking, I'm just not sure how far off that is in realistic terms, and if the complexities may outweigh the benefits in a lot of use cases.
The node shrinks are definitely slowing and have for a several years.
Wikipedia provides this rough estimate of feature size - year table: 10 µm – 1971 6 µm – 1974 3 µm – 1977 1.5 µm – 1982 1 µm – 1985 800 nm – 1989 600 nm – 1994 350 nm – 1995
What about only a 100K USD?
[1] capital costs only. Not labor.
You couldn't even buy the cleanroom building for $1m
Edit: Page 64 of the pdf https://dokumente.unibw.de/pub/bscw.cgi/d9262701/01_History.... suggests in 1980 the total investment cost of a fab was $100million. I would expect that same tech could be replicated a lot cheaper today. A 100x improvement doesn't seem outrageous.
I'm no expert, but a 100x savings seems pretty outrageous to me and 300x even more so.
So, that's one data point.
To the effect, only two. It is surprising just how fast it has turned into a duopoly.
Samsung and TSMC are now the Airbus and Boeing of semiconductor industry
At these scales, very few players can afford to play, even if a breakthrough could net you a very comfortable position for many years.
Thinking about it naively, I can guess at a few possibilities:
- it's highly labor-intensive
- it requires highly specialized skills that command incredibly high prices on the labor market
- it requires lots of prototyping iterations that require expensive materials
- the prototypes are produced on machinery with high opportunity costs
- IP licensing costs
I'm wondering if one of these in particular is the dominant cause, or if it's all of them in conjunction, or if there are other major factors I'm not thinking of.
Also, figuring out how to maximize yields to an acceptable level must be a lot of experimentation.
I think those design costs are nonsensical. There are major projects done by groups of 100 or fewer engineers. A really unique design can cost that much, but a licensed design built.on standard cells is less. That is how rocketchip, RiscV, and others get by. As an approximation, every M$ is 2 engineer years. A 100M$ design is...200 engineer years. With an ARM license and other IP, that is likely manageable.
There is a large infrastructure of design tools that keep design costs constrained per design. Fab costs are not so easily limited....
- EUV comes from a plasma of metal created in a vacuum chamber.
- the plasma is created by concentrating a beam of light on a tiny pellet of molten metal
- that beam of light is a 25kw laser, which is borderline weapons-grade. To the point that countries hesitate to even allow it to be imported.
- the laser must be pulsed precisely to vaporize the metal as its flying through the chamber
- residue from the metal vapor quickly builds up and deteriorates the process
getting this stuff to work requires armies of engineers and scientists. the machines themselves are the size of a tour bus.
Compare min metal pitch for Intel’s “10nm” (36nm) and TSMC’s “7nm” (40nm) [1]
Then there's the question of why they aren't skipping 10nm if 7nm is so ahead.
And theres the question of why they can't 'scale up' their 7nm to do 10nm.
So manufacturing 10n, with UV is much harder than manufacturing 7nm with EUV (although it is more expensive).
I've been out of that loop for a while, so I don't know how the design of the EUV machines turned out.
Apparently, tiny droplets of tin are created by spinning some disk(s) and then those droplets are zapped with with a laser before they emit EUV. Those machines are crazy complex.
[0] https://www.anandtech.com/show/14175/tsmcs-5nm-euv-process-t...
[1] https://www.anandtech.com/show/14312/intel-process-technolog...
And although "2020" and "2021" look like actual dates, they aren't; all they are is marketing labels, because saying you're going to do something in 2020 is cheap.
So what we know is that TSMC say they are going to begin volume manufacturing of "5nm" chips in "2020" (reality: at some unknown date they will be making large quantities of chips on a process to which they have found it convenient to attach the label "5nm"), and that Intel say they are going to launch "7nm" chips in "2021" (reality: at some unknown date they will be making unknown quantities of chips on a process to which they have found it convenient to attach the label "7nm").
Whether TSMC's "2020" is really earlier than Intel's "2021", and whether TSCM's "5nm" is really denser than Intel's "7nm", who knows?
(I'd guess TSMC are in fact ahead, but this isn't a field where you want to be trusting manufacturers' announcements much.)
Then what is Intel?
IBM fabricated single-atom transistors that worked with adequate reliability back in the 1990s, IIRC. I don't know how bad the noise problem you allude to really is.
Wikipedia has a nice picture of the lattice structure:
https://en.wikipedia.org/wiki/Diamond_cubic#/media/File:Visu...
Part 3 of the picture is a 3x3x3 block of elementary cells, for silicon this would be 1.63nm along each edge, so for a 2nm element you get a few more atoms in each direction.
That being said, direct S/D tunneling is expected to become an issue at channel lengths of around 1nm, although anisotropic carrier mass can be used to delay this further.
On top of all that, the design rules become crazy at small dimensions. You can't just make a sharp bend in a wire because the high frequency component of the bend causes ringing in the interference pattern of the EUV laser.
* - I'm just a software guy with a BSEE... I'm curious what a semiconductor designer has to say about all this.
Sharp bends don't happen because they cause large fields which in turn cause dielectric breakdown. Most critical metal layers are oriented in a single direction. 2D printing with EUV isn't really an issue.
There are three components which make the area of a cell that are used to infer the scaling. The fin pitch, the metal pith, and the cell height (track count). An older technology (22nm) might have a 9track, 40nm by 60nm size. 14nm would be 9T2842, 10 would be 7.5T2638, 7 would be 6.5T2638 with SDB, and 5 might be 5.5T2230
Numbers very approximate, but the key is that design compaction (which requires major physical integration changes) coupled with reduced key pitch shrink.
The nice part is that you tend to get more design innovation, because you are no longer competing with shrink. And the relative cost of additional masks and layers is low.
Is the comparison here a device with the same number of transistors, a device with the same area, or something else entirely?
Because if that's the per-area cost, then the design cost per transistor has barely budged.
It seems to me that going from 7nm to 5nm (a factor 1.4) is a smaller step than going from 5nm to 3nm (a factor 1.6666). Not just a smaller step in engineering effort, but also in effect size.
In future you can get current flagship under a watt too. But then the new flagship with it’s 5W consumption will be so much better.
That would be amazing but probably not feasible with the kind of display people have come to expect on their smartphone.
This seems on the face of it possible. I do not know enough about the specifics if it is reasonable to do.
I could see it though. Just leave your phone face up on your desk during the day to have it charge.