Intel's 10nm Is Broken, Delayed Until 2019
tomshardware.com
tomshardware.com
I think the cost and complexity of advancing lithography processes has increased so much that the technology risks have increased dramatically. Some foundries may blunder relative to others.
Intel has abandoned their tick–tock model for Prosess-Archictecture-Optimization. In retrospect it seems clear that they knew that 10nm process is risky step and it's impossible to predict when it's ready.
Now it's more like Process-Archictecture-Optimization-Increased Clock Speed And Power Consumption-We Don't Know What We're Doing Anymore
Also, you're right about the complexity. Intel's 10nm process is likely far more complex (more steps) than Samsung's 7nm EUV process, or even TSMC and GloFlo's 7nm DUV processes, which is why it's taking them so long.
In retrospect, Intel becoming a "manufacturer for ARM chip companies" seems laughable now, doesn't it?
To significantly reduce power consumption you need improved process and this is where smaller nodes are so important, because they are currently the only really viable way to significantly reduce power consumption.
Ouch. A day late and a dollar short.
They claim that this new node will still be better than TSMC’s new node, but we are now in leapfrog mode aren’t we? Where you have a year advantage on your competitor and then they will have the best tech but you’ll be halfway to unseating them again?
> The company will switch to EUV at 7nm.
Knowing nothing about chip design I'm probably thinking about this the wrong way but socket backwards compatability aside is it not feasible to simply increase the chip size? Is a higher density more rewarding?
Thermals and also power delivery are huge problems with large chips, just compare the massive 471 mm2 die of the GP102 (1080 Ti/Titan X) to the 150 mm2 die of the Coffee Lake hexacore chips. GP102 can draw 250-300W depending on boost clock, the Core i7-8700K can also draw upwards of 200W depending on how high you push the clocks and vCore (to keep said clocks stable).
There's a reason why board-partner GPU's always have huge coolers attached to them, and why people pushing CPU clocks are often using at least a giant air cooler like the Hyper 212 EVO or an AIO liquid cooler with a 240mm+ radiator.
Hell, let's skip thermals and just talk electricity - getting 200W+ of stable power to the cores on these dies is no easy task as-is, that's why you have people like buildzoid doing reviews of power delivery on motherboards and GPU boards to see if VRM's are going to blow up trying to power your expensive hardware if you're overclocking (or sometimes even if you aren't).
All in all, we have thermal and power scaling issues at current chip sizes - making them bigger isn't particularly feasible unless everybody is going to start installing 360mm radiators in their system and even that might not be enough depending on clock speeds and the vCore required to maintain them.
I guess it's like trying to get from LA to SF. You could get to SF of you 10x'd the size of earth, but it would still more than 10x as long despite having the same connections because you'd need to stop for some extra connections to even make it.
Another reason is yield. Defects are inevitable. Their probability per unit of area is roughly constant, i.e. it doesn’t depend on the area of a single chip. Therefore, the probability of one or more defects on a single chip is proportional to the exponent of the chip area. With larger chips, that exponent grows very quickly.
Cores only make about half of the area. If the error is not in a core but in e.g. RAM controller or IO controller, you have to throw away the complete chip.
If you compare semiconductors to crops, this determines how much bucks you get from an acre of silicon wafer.
Power use increases quadratically with voltage. You want small transistors to keep the voltage and power use from getting out of hand. You also need to increase voltage if you want to increase clock speed.
Electric signal travels in a conductor roughly 15 cm/nsec. With 3 GHz clock speed the electric signal travels travels roughly 50 mm in one clock cycle. Largest microchips are 30 mm across. You can't double the dimensions without dealing with the signal lag. Delivering the clock signal to every part of the chip in sync is already a problem. Modern microchips use lots of extra circuitry just to deliver the clock signal properly.
Or use lower VT cells (they turn "on" quicker), at the expense of increased leakage power. But at these geometries, faster clock speeds is getting less feasible and you need to find increased performance in other ways.
No logic signal needs to cross the entire die in one clock cycle, there is always an alternate design. For that reason, only registers 'talking' to each other need to see a clock at the same time, and even then there is a window. Clock routing is a consideration that takes resources but it's not a problem. Realistically, a logical signal won't be going anywhere remotely near 50mm at 3GHz in 10nm, so the clock doesn't need to either.
Max die size is also limited by the vendor's tooling i.e. what their machines can literally handle. And also physical issues such as warpage. If you make a massive die and it heats up in a non-uniform manner (different bits of it get hot at different times), it expands in a non-uniform manner. This can lead to all kinds of problems.
Chips stuffed full of memory will yield better than a logic-heavy chip since large SRAMs always now include redundancy. So this too has an impact on how big you can go for a given cost. You can however get registers that are built of multiple storage elements, the output value of which is the consensus. Don't know how much these get used.
In reality its the signal speed relative to jitter, time interval errors and data setup times that complicate the design and signal integrity.
If the clock is 3GHz, the margin of error is the fraction of the time of the clock rate. You need to divide the chip into clock regions and add local cache for each core because fetching data far away is too slow.
I've always wondered why you can't generate clocks locally, but in a synchronized way. So basically like clock regions, but without having to add extra logic for data that goes between regions.
CPU Cores use a singular clock. But when cores communicate or the L3 cache communicates (cache coherency is needed if you want that mutex / spinlock to actually work), then you need some kind of communication mechanism between the CPU Cores. Those clocks are likely "locally generated", but there needs to be a translation mechanism between Bus -> Core.
With buses, you have different clocks and data moves between them. Like you said: CPU core 1 has its own clock, the bus between them has its own and different clock, and then CPU core 2 has its own clock which is yet again different. And in those cases you actually want different clocks, because you want to be able to boost CPUs independently from each other.
What I meant goes in another direction: instead of having a single powerful clock source for e.g. a CPU core, you have multiple smaller clock sources distributed throughout the core, but synchronized to each other so they run at the same frequency and phase. So data can move freely like it does today, but clock signals don't have to be distributed as far, which would hopefully make clock distribution easier and less power hungry.
It seems like such a thing should be possible, but perhaps there are good reasons why it isn't done?
1. Clocks don't use a lot of power. Think of a pendulum: there's a lot of movement but the energy constantly swings between gravitational potential energy and kinetic energy. Although there's lots of movement, the device uses very little energy. Similarly, a clock circuit (called an oscillator) barely uses any electricity: it mostly "Swings" energy back and forth between an inverter and a capacitor.
2. Distributing a clock over a long distance similarly uses very little power (!!) due to transmission line theory. You can effectively use the parasitic capacitance in wires themselves to effectively do this pendulum effect for efficient long-distance transmission of clocks. See: https://en.wikipedia.org/wiki/Transmission_line
This gif shows an animation of the pendulum effect in a longer-transmission line: https://upload.wikimedia.org/wikipedia/commons/8/89/Transmis...
----------------
I guess things could be de-sync'd for more efficiency. But your question is kind of like "Well, can't we get rid of V-Tables in C++ to make branch-prediction more efficient??"
I mean, we can. But V-Tables / Polymorphism really doesn't take a lot of time. We only do that if the performance gain really matters.
I do have one follow-up question though: I was under the impression that clock trees contain repeaters in the form of CMOS inverters. Wouldn't those have dynamic leakage which the transmission line stuff doesn't account for?
From my understanding: yes, the CMOS inverters will certainly use power. But you can minimize the use of them through some passive techniques.
Looking into the issue more, it does seem like a naive implementation of synchronized clocks can become costly. But at the same time, I'm seeing a number of research papers suggesting that people have been applying transmission-line techniques to the clock distribution problem.
I've always assumed that it was something that was commonly done at the chip level, but apparently not. These papers were published ~2010 or so.
* Yields: silicon wafers have regular manufacturing errors. A bigger die means more failed CPUs, grossly increasing prices. Smaller dies isolate those errors better, leading to better yields. Lets say there are around 20-errors per wafer. 100-chips per wafer would result in ~80is to 85ish successful chips per batch.
If you shrunk the die so that you had 500-chips per wafer, then you have 480-chips after manufacturing (20-defects).
Wafers are a constant size. Errors are relatively constant as well. You can't change those numbers.
* Power: Smaller feature sizes use less power. Smaller capacitance, so the signals travel faster and generally speaking the design can be clocked higher.
The research project is http://invasic.de
I really like this change of course.
Of course I'd be loading a lot of data from disk into RAM that would have simply been encoded as immediate values in the executable before. But my rough estimate still puts this at (at least) an order of magnitude less RAM than abstraction layers like WPF use.
And that's when I realized: we jumped from severely resource constrained designed straight to completely wasteful designs without adequately exploring the design space in between. To paraphrase Richard Feynman's quote, "There's plenty of room at the bottom": "There's plenty of room in the middle."
The focus should be on developer productivity first, then product iteration, then performance.
For example, try to compute CRC32 checksums of 100 byte arrays at bus speed in C/C++ and then in the higher level language of your choice.
In all seriousness though, I wonder if once Moore's law really stops we'll see a surge of innovation in compiler and language paradigms, born out of necessity. It seems to me that the revolution in higher-level languages and paradigms came because we could (computationally) afford it, so it makes sense that once the incentives change we'll see innovation in other areas.
Big 5 want to hire like crazy is a myth.
I get why it's done, but part of me wishes that CPU growth would stall off for a bit so the software industry would have no choice but to treat optimization as a selling point again.
It is not even the case of using GC and such.
Our phones are way more powerful than any Xerox PARC or Lisp Machine ever was.
From what I understand LPDDR4 as part of Cannon Lake is directly affected by this delay.
Do we have confirmation of this? It seems startling to release yet another generation of laptop CPUs without this... but I admit I can’t find anything to to the contrary.
Pretty much anyone building mid-to-high-end laptops must be livid.
There is no other rumoured "lake" in between, so yes, another generation without 32GB RAM Laptop.
These information are widely available everywhere, not sure what you want as confirmation.
Well those which put a premium on sleep/suspend autonomy, which for DDR4 is the big edge of LP. IIRC the "active" energy consumption of DDR4 was lowered to LP-level already.
What does that mean?
Do other manufacturers, offering 32gb, use desktop RAM? What kind of compromises would that bring to a MBP?
I'm not _fine_ with it, but right now I need 32gb of ram more than I need those things.
They use regular DDR4, talking about "desktop" RAM may not be the best for comprehension, DDR4 is available as SODIMM module and lots of work went into making DDR4 significantly less power-hungry than DDR3.
> What kind of compromises would that bring to a MBP?
IIRC it would burn battery ~30% faster when sleeping (e.g. when you close the lid, unless you have changed the configuration to be strictly "suspend to disk", which requires going through pmset and the command-line).
while Moore's original paper: http://www.monolithic3d.com/uploads/6/0/5/5/6055488/gordon_m... was always about the trend of putting _more_ components in a chip being more cost effective than keeping the number of components constant. Surely if this were still the case we'd have seen Skylake++++ with cores << n by now.
I suggest you actually read the article you linked.
Does it meltdown ?
Microtransactions were the next step for the gaming industry after traditional expansion packs and then DLC, let's just hope we don't get CPU loot boxes (though I guess we already have that simply due to the silicon lottery, eh?).
https://groups.google.com/forum/message/raw?msg=comp.lang.ad...
"He went on to point out that they had calculated the amount of memory the application would leak in the total possible flight time for the missile and then doubled that number. [...] Since the missile will explode when it hits it's target or at the end of it's flight, the ultimate in garbage collection is performed without programmer intervention."