Because the amount of power leakage (heat) is proportional to the size of the transistors. So you cannot improve the efficiency of the chip without making the transistors smaller or lowering the frequency.
> Obviously there would be downsides in cost of materials and power consumption,
Huge costs. Data centers aren't only worried about the power consumption of chips, because for every watt that's generated as heat, they have to use >1W to cool that (because cooling systems aren't 100% efficient themselves).
As you may know from other branches of science, the resistance of a conductor rises with heat, meaning that electrons running through the conductor are more likely to hit a vibrating atom and dissipate as heat. This is why so-called "super conductors" are usually super cooled. Very little atom movement = very low chance of an electron hitting a moving atom. Silicon is a semi-conductor, but the principal is the same. If the temperature of the chip rises, so does the heat generated. This is why world-record setting overclocks are done using liquid nitrogen to cool the chip.
To prevent things from getting too toasty, data centers would have to reduce the density of the servers, which means they would need larger buildings and more land to house the same number of servers.
> but wouldn't that be offset by avoiding the need for nano-manufacturing advances in every single round?
In short, no.
Small correction: leakage actually increases as transistors shrink. This is why high-k dialectric and fin-fets were such important developments. They pushed back the point at which leakage power overtakes switching power as the dominant source of waste. Even with these technologies, we have to do a lot of design work to reduce leakage. I can't even guess how many power domains are on modern cpus -- certainly dozens if not hundreds. Most of those domains can be switched off to eliminate leakage in those domains altogether.
I'm nearly certain you meant switching power reduces as transistors shrink, which has been true so far. Things are getting weird with these new processes, and a lot of things that we've held as fact are looking less and less reliable.
Yup. Thanks for the correction.
http://synisma.neocities.org/perf_scale_cheatsheet.pdfAfaik temperature dependence of materials is a more complicated relation than this. The graph of their coefficient does not even need to be monotone and can change depending on what properties dominate.
E.g. a semiconductor can have NTC because charge concentration increases with the temperature, however when it reaches saturation this effect diminishes and it will behave similar to a regular PTC conductor.
https://en.wikipedia.org/wiki/Temperature_coefficient#Negati...
Not at all. A typical cooling system might use 1 watt to move 4 watts of heat outside.
And that's not really how superconductors work. That's how normal conductors work.
Can you name any super conductors which function at room temperature? I'm not aware of any. [1]
[1] https://en.wikipedia.org/wiki/High-temperature_superconducti...
Five wafers stacked on top of each other would take 5 times as much capacity and materials to produce and cost about 5 times as much. The interconnect would be very difficult. Lastly, while you could cool the top wafer, the bottom layers would have to go through a lot of silicon to remove heat. Five times the number of layers on each wafer would also be prohibitively expensive, you don't just cut it thicker, you vapor deposit each layer under a mask and often also need to etch off parts of previous layers or ensure that higher layers are still planar. That gets difficult when you have a few steps of logic layering and some metal layers for interconnect, and would be much more expensive (and hot) with many layers.
This isn't so much of a problem with eg. NAND flash chips which are low power and the address and data pins can just be shared with a couple wafer select wires to separate them, but processors are neither low power nor trivially paralleled.
But true, heat is a huge problem.
Imagine you have a 4in by 4in square. If your chips are 1in^2 you can fit sixteen chips but if your chip is 4in^2, you can only fit four. Now imagine you have a thick scratch going diagonally right down the middle. With the bigger chips, you might lose all four to that single error, wiping out your yield. With the smaller chips however, you'll only lose part of your wafer.
Specialized processors like those made for mainframes or RAD hardened ones can be much bigger since the set up costs will vastly outnumber the fabrication cost anyway. Companies like IBM that aren't in as cost competitive a market as Intel make big chips all of the time.
Intel has an advantage here: they use 12 inch wafers while other fabs use 10 inch wafers. This improves their relative yield.
That calculus (so-called Dennard scaling) has now broken down, which is why Intel has abandoned their tick-tock development model in favor of a model that doesn't solely rely on process shrinks in order to achieve better performance.
All this feeds into a competitive market between chip manufacturers. If you are a memory manufacturer and can find the right balance that makes the product 5% cheaper to manufacture, they can make a ton of money and gain marketshare. There are a few companies like Apple, Intel that can differentiate on brand name but the rest are commodity products that compete on cost and features.
Just my best guess from a single processor architecture class so definitely not positive that's the answer.
Correct. The chip has to be small enough that the clock can propagate everywhere within the chip within a single cycle, or problems will occur.
> If a chip is physically bigger, it takes longer to move bits inside of it.
Yes, so either you would need to delay for some cycles to ensure that the information has propagated (which will basically nullify the performance gains from cranking up your clock) or clock parts of the chip differently, but you're pretty much always limited by the slowest part of the chip (which is why every modern chip has a cache, because otherwise it would stall waiting for data).
Problems like what? These chips are already chopped up into different clock domains, and it's easy to install some PLLs so that perfectly-synchronized clock signals can blanket a chip even if it's inches across.
Moving data around is also not a big deal. The Xeons in the article already have multi-nanosecond ring busses running around between cores[1]. They don't slow the chip down because the design simply lets long-distance data transfers take multiple cycles. L3 and I/O don't have to be blazingly fast in terms of latency.
[1] http://images.anandtech.com/doci/8423/HaswellEP_DieConfig.pn...
The chip won't work.
> These chips are already chopped up into different clock domains, and it's easy to install some PLLs so that perfectly-synchronized clock signals can blanket a chip even if it's inches across.
Sorry, I didn't explain myself well enough. Of course chips have different clock domains, but these also come at a cost. The more synchronization you need to do between domains, the less die space you have for computationally useful stuff.
> L3 and I/O don't have to be blazingly fast in terms of latency.
I would argue differently, the impact of latency is highly dependent on the type of computation you're doing. If you're doing something with a lot of data (say, encoding video) then you need to be moving data as quickly as possible between the processor and memory. Any additional latency in cache or I/O will cause the performance to suffer.
Ideally, you want the latency of L3 and I/O to be as low as possible.
And encoding video is nearly the platonic ideal of not caring about memory latency. You could easily make memory requests ten thousand cycles before you need the results. You just need throughput.
Overall these are very very strong factors which push towards shrinking chip size at every opportunity.