2. Power delivery is now ~n times worse where n is your transistor layer count. *
3. Connections between chips are very slow, power hungry and expensive. Fabrication of "monolithic" 3D is temperature wise painful and usually results in crummier transistors.
With that being said, innovative 3D integration methods in specific applications can help a lot. Shameless plug: we at Vathys do this for deep learning chips.
* to a first order of approximation
Having more dies lets you dissipate more heat, but then it's kinda hard to build low-latency / high-bandwidth interconnects between the dies. Inter-die buses go over a PCB or interposer, which impose higher parasitic capacitance and make it difficult/expensive to run wide interfaces. That's why techniques like "dark silicon" allocation are important - it allows us to get more perf in a single die.
All you'd end up doing is heating that inside up to the temperature of the dies and after that there would be no more cooling effect (and this would happen in a few seconds after starting the whole thing up). You could do an 'inverse' of this by cooling the dies from the outside and having the interconnects in the space in between. This would still require a lot of cooling and there would be an issue with connecting the resulting assembly to the underlying PCB.
Not directly an answer to this specific question, but a (somewhat) colleague who writes his PhD thesis about 3-dimensional chip design made a popular scientific lecture about this topic. As I understood it, the central problem is that it is very hard to produce chips with multiple (lots of) layers (where you want to have interconnects inbetween). In particular producing the interconnects between the layers is really hard if they can lie "everywhere" instead of only at the border. There exist multiple ideas how this might be done (e.g. drill holes with high-precission lasers into the substrate and try to fill them with something conductive), but none of them as of today "really works" (at least if we are talking about chips with somewhat more than 4 layers and interconnects everywhere inbetween).
Another goal we have is low latency (and high clock rate, which is depenend on low latency), which suggests putting everything in a qube or even a shere. So we compromise somewhere in the middle with a square with a few layers.
Made me think of this:
https://en.wikipedia.org/wiki/Menger_sponge
Say we have roughly 300 sq mm, that's about 17.32 mm square, which is 8.24e+7 silicon atoms (0.21 nm) across.
Then we have a surface area of roughly 1.36e+16 atoms - the area x 2.
If we make a fractal sponge down to the limit of single atoms, then that's about 16.5 cycles of removing cubes from a 17.32 mm cube. Let's ignore the difficulty of doing it half a time. According to the formulas from wikipedia, the result has a volume of about 0.7% of the original, with 2.4e28 "sides" of atoms exposed.
So the third dimension gets you about 1.8 million times the surface area. I suppose this isn't nearly as good as 4e10 flat sheets with 1 atom separation between each, but you could argue it's more practical because everything is connected...