There are kernel patches to enable P2P communication via PCIE on 3090 which is almost as fast as NVLink for vllm
* 3090 memory bandwidth: 936 GB/s
* Maximum NVLink bandwidth between two 3090s: ~56 GB/s one way
* Maximum PCIe v4 bandwidth between two 3090s: 31.5 GB/s one way
* M5 Ultra memory bandwidth: 1.2 TB/s
I know that memory bandwidth is only one factor influencing LLM performance, but this seems like a major problem if your goal is to run larger models that won't fit on one 3090 - and even those fitting on two 3090s with NVLink will be pretty restricted.