The trick is to add 3D features that don't involve making an entirely new device on top of another. This is the case for RAM, MRAM, and flash memory: they share devices in the stack, the only parts that are being scaled vertically are charge/spin carrying parts, but not amplifiers, backend, data lanes, or other devices.
There was a lot of talk about making 3D standard cells (stuff from which normal, non-memory, devices are made of.) The amount of work is immense. Every year there are a dozen cookie cutter PhD work like "3D NAND/XOR/INVERT device that is N percents smaller than before," but it will take years to cover and unify the whole cornucopia of devices in cell libraries. And only once it's done, will major fabs think of switching to that. No fab will try to add much more litho layers just to reduce footprint of only one device or macrocell.
Or a tube-like cylindrical thing?
I suppose a long tube might work, but that'd be a weird shape to work with.
https://en.wikipedia.org/wiki/3D_XPoint https://en.wikipedia.org/wiki/Flash_memory#Vertical_NAND
(2) How the heck would you cool it?
Now imagine, thanks to 3D cells, you now magically get double the number of registers. How much efforts do you think you need to spend to double the amount of loops running in parallel?
2) I seem to remember IBM was looking at microfluidics for both power and heat distribution, like our blood.
I would love to see the microfluidics cooling. I wonder what would happen if you reached boiling though--it'd probably explode the chip.