Intels upcomming Saphire Rapid server CPUs are extremly similar, with wide connections between two close dies. Crossectional bandwith is in the same order of magnitude there.
It sounds like this operates as if it was one giant physical chip, not two separate processors that can talk very fast.
I can’t wait to see benchmarks.
It's probably best to think of this chip as an extremely fast double socket SMP where the two sockets have much lower latency than normal. Software written with that in mind or multiple programs operating fully independent of each other will be able to take massive advantage of this, but most parallel code written for single socket systems will experience reduced gains or even potential losses depending on their parallelism model.
Whereas AMD's solution is focused on increasing the cache size (hence the 3D stacking), Apple here seems to be connecting the 2 M1 Max chips more tightly. It's actually more reminiscent of AMD's Infinity Fabric interconnect architecture. https://en.wikichip.org/wiki/amd/infinity_fabric
The interesting part for this M1 Ultra is that Apple opted to connect 2 existing chips, rather than design a new one altogether. Very likely the reason is cost - this M1 Ultra will be a low volume part, as will be future iterations of it. The other approach would've been to design a motherboard that sockets 2 chips, which seems would've been cheaper/faster than this - albeit at expense of performance. But they've designed a new "socket" anyway due to this new chip's much bigger footprint.