Here are a couple of articles with more detailed info:
https://semiwiki.com/semiconductor-manufacturers/tsmc/306329...
https://www.anandtech.com/show/16051/3dfabric-the-home-for-t...
https://www.eetimes.com/amd-tsmc-imec-show-their-chiplet-pla...
https://www.techpowerup.com/292256/amd-details-its-3d-v-cach...
In addition to splitting the CPU or GPU into distinct units, you can also take other functions and use different processes for them. For example, in package I/O or L2 cache don't really see the same advantages for newer processes, so you can make these using more established (and cheaper/more available) processes.
I'm not an expert at how they combine chips. Like I said for AMD, they also wanted their units to be composable w/ small number of chips, so they basically have a die w/ a few cores and different die w/ memory, and I believe they have a proprietary communication mesh for connecting them. I think there is some considerable signal/energy overhead to communicating between chips. The cost of masks and interconnects is probably high enough to make a high number of diverse chiplets inviable, but I wouldn't say it's impossible in the future.
>I think there is some considerable signal/energy overhead to communicating between chips.
I tried looking for some info on this but couldn't find any. Do you have any source I could read on this?
Another way of looking at it is bit error rate [2].
I think you'll be hard-pressed to find concrete numbers since these designs are closely-guarded trade secrets. You might find some examples by searching for wireline transceiver / PHY papers on IEEE xplore, especially at ISSCC (conference) or in JSSCC (journal).
[1] https://en.wikipedia.org/wiki/Shannon%E2%80%93Hartley_theore...