Knowing nothing about chip design I'm probably thinking about this the wrong way but socket backwards compatability aside is it not feasible to simply increase the chip size? Is a higher density more rewarding?
Knowing nothing about chip design I'm probably thinking about this the wrong way but socket backwards compatability aside is it not feasible to simply increase the chip size? Is a higher density more rewarding?
Power use increases quadratically with voltage. You want small transistors to keep the voltage and power use from getting out of hand. You also need to increase voltage if you want to increase clock speed.
Electric signal travels in a conductor roughly 15 cm/nsec. With 3 GHz clock speed the electric signal travels travels roughly 50 mm in one clock cycle. Largest microchips are 30 mm across. You can't double the dimensions without dealing with the signal lag. Delivering the clock signal to every part of the chip in sync is already a problem. Modern microchips use lots of extra circuitry just to deliver the clock signal properly.
Or use lower VT cells (they turn "on" quicker), at the expense of increased leakage power. But at these geometries, faster clock speeds is getting less feasible and you need to find increased performance in other ways.
No logic signal needs to cross the entire die in one clock cycle, there is always an alternate design. For that reason, only registers 'talking' to each other need to see a clock at the same time, and even then there is a window. Clock routing is a consideration that takes resources but it's not a problem. Realistically, a logical signal won't be going anywhere remotely near 50mm at 3GHz in 10nm, so the clock doesn't need to either.
Max die size is also limited by the vendor's tooling i.e. what their machines can literally handle. And also physical issues such as warpage. If you make a massive die and it heats up in a non-uniform manner (different bits of it get hot at different times), it expands in a non-uniform manner. This can lead to all kinds of problems.
Chips stuffed full of memory will yield better than a logic-heavy chip since large SRAMs always now include redundancy. So this too has an impact on how big you can go for a given cost. You can however get registers that are built of multiple storage elements, the output value of which is the consensus. Don't know how much these get used.
In reality its the signal speed relative to jitter, time interval errors and data setup times that complicate the design and signal integrity.
If the clock is 3GHz, the margin of error is the fraction of the time of the clock rate. You need to divide the chip into clock regions and add local cache for each core because fetching data far away is too slow.
I've always wondered why you can't generate clocks locally, but in a synchronized way. So basically like clock regions, but without having to add extra logic for data that goes between regions.
CPU Cores use a singular clock. But when cores communicate or the L3 cache communicates (cache coherency is needed if you want that mutex / spinlock to actually work), then you need some kind of communication mechanism between the CPU Cores. Those clocks are likely "locally generated", but there needs to be a translation mechanism between Bus -> Core.
With buses, you have different clocks and data moves between them. Like you said: CPU core 1 has its own clock, the bus between them has its own and different clock, and then CPU core 2 has its own clock which is yet again different. And in those cases you actually want different clocks, because you want to be able to boost CPUs independently from each other.
What I meant goes in another direction: instead of having a single powerful clock source for e.g. a CPU core, you have multiple smaller clock sources distributed throughout the core, but synchronized to each other so they run at the same frequency and phase. So data can move freely like it does today, but clock signals don't have to be distributed as far, which would hopefully make clock distribution easier and less power hungry.
It seems like such a thing should be possible, but perhaps there are good reasons why it isn't done?
1. Clocks don't use a lot of power. Think of a pendulum: there's a lot of movement but the energy constantly swings between gravitational potential energy and kinetic energy. Although there's lots of movement, the device uses very little energy. Similarly, a clock circuit (called an oscillator) barely uses any electricity: it mostly "Swings" energy back and forth between an inverter and a capacitor.
2. Distributing a clock over a long distance similarly uses very little power (!!) due to transmission line theory. You can effectively use the parasitic capacitance in wires themselves to effectively do this pendulum effect for efficient long-distance transmission of clocks. See: https://en.wikipedia.org/wiki/Transmission_line
This gif shows an animation of the pendulum effect in a longer-transmission line: https://upload.wikimedia.org/wikipedia/commons/8/89/Transmis...
----------------
I guess things could be de-sync'd for more efficiency. But your question is kind of like "Well, can't we get rid of V-Tables in C++ to make branch-prediction more efficient??"
I mean, we can. But V-Tables / Polymorphism really doesn't take a lot of time. We only do that if the performance gain really matters.
I do have one follow-up question though: I was under the impression that clock trees contain repeaters in the form of CMOS inverters. Wouldn't those have dynamic leakage which the transmission line stuff doesn't account for?
From my understanding: yes, the CMOS inverters will certainly use power. But you can minimize the use of them through some passive techniques.
Looking into the issue more, it does seem like a naive implementation of synchronized clocks can become costly. But at the same time, I'm seeing a number of research papers suggesting that people have been applying transmission-line techniques to the clock distribution problem.
I've always assumed that it was something that was commonly done at the chip level, but apparently not. These papers were published ~2010 or so.
Another reason is yield. Defects are inevitable. Their probability per unit of area is roughly constant, i.e. it doesn’t depend on the area of a single chip. Therefore, the probability of one or more defects on a single chip is proportional to the exponent of the chip area. With larger chips, that exponent grows very quickly.
Cores only make about half of the area. If the error is not in a core but in e.g. RAM controller or IO controller, you have to throw away the complete chip.
I guess it's like trying to get from LA to SF. You could get to SF of you 10x'd the size of earth, but it would still more than 10x as long despite having the same connections because you'd need to stop for some extra connections to even make it.
* Yields: silicon wafers have regular manufacturing errors. A bigger die means more failed CPUs, grossly increasing prices. Smaller dies isolate those errors better, leading to better yields. Lets say there are around 20-errors per wafer. 100-chips per wafer would result in ~80is to 85ish successful chips per batch.
If you shrunk the die so that you had 500-chips per wafer, then you have 480-chips after manufacturing (20-defects).
Wafers are a constant size. Errors are relatively constant as well. You can't change those numbers.
* Power: Smaller feature sizes use less power. Smaller capacitance, so the signals travel faster and generally speaking the design can be clocked higher.
Thermals and also power delivery are huge problems with large chips, just compare the massive 471 mm2 die of the GP102 (1080 Ti/Titan X) to the 150 mm2 die of the Coffee Lake hexacore chips. GP102 can draw 250-300W depending on boost clock, the Core i7-8700K can also draw upwards of 200W depending on how high you push the clocks and vCore (to keep said clocks stable).
There's a reason why board-partner GPU's always have huge coolers attached to them, and why people pushing CPU clocks are often using at least a giant air cooler like the Hyper 212 EVO or an AIO liquid cooler with a 240mm+ radiator.
Hell, let's skip thermals and just talk electricity - getting 200W+ of stable power to the cores on these dies is no easy task as-is, that's why you have people like buildzoid doing reviews of power delivery on motherboards and GPU boards to see if VRM's are going to blow up trying to power your expensive hardware if you're overclocking (or sometimes even if you aren't).
All in all, we have thermal and power scaling issues at current chip sizes - making them bigger isn't particularly feasible unless everybody is going to start installing 360mm radiators in their system and even that might not be enough depending on clock speeds and the vCore required to maintain them.
If you compare semiconductors to crops, this determines how much bucks you get from an acre of silicon wafer.