This video from Cerebras perfectly explain how they solve the interconnect problem, and why their approach greatly reduces the risk of Blackwell-type hardware design challenges.
https://www.anandtech.com/show/16626/cerebras-unveils-wafer-...
The cooling is significantly better than what you'd see on a server platform with water cooling channels going to each row of the wafer.
I had guessed that Cerebras had made some trade-offs in process in order to make it work at scale, but then they aren't actually building these devices at scale (yet).
https://cerebras.ai/press-release/cerebras-systems-smashes-t...
And actually there have been attempts to do it, I mentioned in an earlier version of my comment that Gene Amdahl had attemped to make WSE work something like 20 years ago, without success - but also without the clear profitability story of AI to attract the same mountains of cash being thrown around today.
What's surprising is not that it is hard, or that it's hard as fuck, but that given the potentially stratospheric rewards for success there have not been more attempts in this direction.
IIRC cerebras' design was originally for HPC workloads, so even it may not necessarily be optimized for LLMs