We currently have 128KByte dual-ported SRAM per core (which is physically part of the core, and not a giant array somewhere else on die). It has single cycle latency to the core's registers and to the Network on Chip router.
The on chip mesh network is custom 128 bit wide going core to core. The router can do a read or write to SRAM per cycle AND allows a passthrough to another core in the same cycle.
Our chip-to-chip interconnect is a custom 64 bit (72 lane) parallel interface allowing 48GB/s. There are two of these (unidirectional) interfaces per side, giving you a total of 8 of these interfaces per chip.