The Dark Silicon Problem and What It Means for CPU Designers (2013)
informit.com
informit.com
I can't tell if this is hyperbole or not, it amazes me but no amount of googling is coming up with a useful answer. Is anyone able to confirm or deny it for me?
A 815 mm^2, 250 W GPU will be 250 W / 8.15 cm^2 ≈ 31 W / cm^2.
I'm getting 6300 W/cm^2 (see: https://news.ycombinator.com/item?id=17589371)
http://www.wolframalpha.com/input/?i=Sun+%7C+luminosity+%2F+...
The sun is: 6300 W/cm^2
> inside of a nuclear reactor
Inside means volume? That's not the same unit, so I'm not really sure how to calculate that.
Interestingly, on a per-volume bases, the sun's output is quite low. The sun is just very very large.
See: http://www.wolframalpha.com/input/?i=Sun+%7C+luminosity+%2F+...
And you find the sun is only: 1.383×10^-6 horsepowers per gallon :) Or if you insist: 0.2725 W/m^3, which as you can see is really really low.
I doubt it comes close to the surface of the average star but certainly close to some of the cooler red dwarf stars and probably most brown dwarf stars.
(Sun's Radius) = 695 x 10^6 m
(Sun's Power) = 4 x 10^26 W
(CPU Output) = 75 W
(Die Size) = 37 mm x 37 mm
(Sun Intensity) = (4 x 10^26 W ) / (4 x pi x (695 x 10^6 m)^2) ~ 66 x 10^6
(CPU Intensity) = (75 W) / (0.0014 m^2) ~ 53 x 10^3
I'm getting a several orders of magnitude higher W/m output for the sun. Perhaps I made an algebra mistake?
I did some research into what can be archived with suitable dielectric fluids and nucleate boiling to allow much higher heat flux on the surface of the silicon without requiring any energy input to actually separate hot and cold, but allowing the use of pumps and such to provide forced flow on the silicon. It seems that if you have a nozzle to provide sufficient flow speeds across the die you can get higher critical heat flux than can be reasonably handled through the power pins on an AMD EPYC socket, which has some limitations due to the LGA technology used for the contacts. A zero-insertion-force PGA socket should not be bound by this limit of temperature rise at the contact of the LGA spring/pin and the chip contact pad due to resistive heating resulting in the contact pressure declining (due to the spring weakening), which yields a feedback loop that can jump over to a domino effect on the other nearby power pins.
I remember the flow rate being proportional to the critical heat flux as well as the distance the flow from the nozzle has to cross and provide cooling for. There were speeds of iirc. about 20m/s if one were to cool a delidded AMD EPYC at maximum power draw @4GHz, or rather, extrapolating from what can be archived with a die surface temperature of 50 degree Celsius.
So, yeah, the heat flux is actually a problem, but I think one could integrate some sensors and sufficiently fast switches that shut the section of the die off if the temperature reaches dangerous levels (and do so fast enough to not damage anything, i.e. microseconds or what time there is), and use such a test chip to engineer a forced-flow direct-die (or maybe even with a heat spreader on top, but that costs you due to the nucleate boiling not going below 20 kelvin temperature loss, and additional system losses in the radiator/heat exchanger as well as pipes likely resulting in costs of about 5 additional kelvin). This could open the door to operating CPU dies at much, much higher power densities than currently normal. It's like a heat pipe on steroids. One would want circuity in the CPU to rapidly shut down in case the temperature rises quickly, as regardless of why this happens, not doing so will literally blow a hole in the chip before you can drain the inductors providing smooth power to the socket.
Temperature on the surface of the sun: 5,600 degrees Celsius. [2]
Temperature of a CPU: ~75 degrees Celsius. [3]
[1] http://academic.brooklyn.cuny.edu/physics/sobel/Nucphys/pile...
[2] http://coolcosmos.ipac.caltech.edu/ask/7-How-hot-is-the-Sun-
A CPU is a lot smaller than both a nuclear reactor and the sun.
that's not true at all. Any time you destroy information (for example an and gate can destroy information) you use energy.
If you're willing to explain, I wonder how an AND gate destroys information. Certainly the output of a 2-input AND gate carries less information than both inputs together. But unless the gate is destroying the input signals, which it isn't, I don't see how infomation is being destroyed.
Related topic: reversible computing. You might wonder if a computation need take any energy at all: it turns out that an irreversible computation (e.g. NOR, XOR, AND etc) is physically guaranteed to waste energy, but if you make sure your compute steps are always reversible (i.e. each input maps to one and only one output) then you can theoretically compute for free (the Feynmann lectures on computation cover this well).
But these energies are negligible compared to the wattage that goes through a standard CPU.
2. Power delivery is now ~n times worse where n is your transistor layer count. *
3. Connections between chips are very slow, power hungry and expensive. Fabrication of "monolithic" 3D is temperature wise painful and usually results in crummier transistors.
With that being said, innovative 3D integration methods in specific applications can help a lot. Shameless plug: we at Vathys do this for deep learning chips.
* to a first order of approximation
Having more dies lets you dissipate more heat, but then it's kinda hard to build low-latency / high-bandwidth interconnects between the dies. Inter-die buses go over a PCB or interposer, which impose higher parasitic capacitance and make it difficult/expensive to run wide interfaces. That's why techniques like "dark silicon" allocation are important - it allows us to get more perf in a single die.
All you'd end up doing is heating that inside up to the temperature of the dies and after that there would be no more cooling effect (and this would happen in a few seconds after starting the whole thing up). You could do an 'inverse' of this by cooling the dies from the outside and having the interconnects in the space in between. This would still require a lot of cooling and there would be an issue with connecting the resulting assembly to the underlying PCB.
Not directly an answer to this specific question, but a (somewhat) colleague who writes his PhD thesis about 3-dimensional chip design made a popular scientific lecture about this topic. As I understood it, the central problem is that it is very hard to produce chips with multiple (lots of) layers (where you want to have interconnects inbetween). In particular producing the interconnects between the layers is really hard if they can lie "everywhere" instead of only at the border. There exist multiple ideas how this might be done (e.g. drill holes with high-precission lasers into the substrate and try to fill them with something conductive), but none of them as of today "really works" (at least if we are talking about chips with somewhat more than 4 layers and interconnects everywhere inbetween).
Another goal we have is low latency (and high clock rate, which is depenend on low latency), which suggests putting everything in a qube or even a shere. So we compromise somewhere in the middle with a square with a few layers.
Made me think of this:
https://en.wikipedia.org/wiki/Menger_sponge
Say we have roughly 300 sq mm, that's about 17.32 mm square, which is 8.24e+7 silicon atoms (0.21 nm) across.
Then we have a surface area of roughly 1.36e+16 atoms - the area x 2.
If we make a fractal sponge down to the limit of single atoms, then that's about 16.5 cycles of removing cubes from a 17.32 mm cube. Let's ignore the difficulty of doing it half a time. According to the formulas from wikipedia, the result has a volume of about 0.7% of the original, with 2.4e28 "sides" of atoms exposed.
So the third dimension gets you about 1.8 million times the surface area. I suppose this isn't nearly as good as 4e10 flat sheets with 1 atom separation between each, but you could argue it's more practical because everything is connected...
>"You can emulate floating-point arithmetic by using integer instructions—but taking 10–100 times as long."
Exactly how is/was floating point arithmetic emulated using only integers?
Why is that range given an order of magnitude? Is this dependent on the precision I'm guessing?
Another source is the Handbook of Floating Point Arithmetic, which covers both hardware and software methods in detail.
An implementation in source code may be found in libgcc. IIRC their ARM assembly code versions are some of the fastest routines available anywhere for software floating point emulation. While they are in assembly, ARM assembly is fairly intelligible.
Finally, you can probably come up with 90% or more of the correct solutions yourself if you start with the definition of a floating-point number as (-1)^s * 1.m * 2^(e - bias) and grind through the algebra with the bias as a constant. The theorems that prove the minimum number and nature of the guard bits to get correct rounding are another matter, but you can start off by computing the equivalent infinite-precision terms and then rounding.
Cheers.
>"The most obvious is the instruction decoder, which is near the start of the pipeline, and is responsible (in the loosest possible terms) for passing the inputs to each of the execution units."
Why would it "in the loosest possible terms"? Isn't this "precisely" the job of the decoder?