CXL Is Dead in the AI Era
semianalysis.com
semianalysis.com
I searched HN for the acronym and there are two posts with more than a few comments, with the most popular one posted five months ago. “Promised as the messiah” my behind.
It's an interesting & valuable comparison: low bandwidth low latency PCIe serdes vs high band high latency Ethernet-like serdes.
But I'm not sure I'm on board with the half-conclusion we get from this half-paywalled analysis. AI is seemingly insatiable sure & there's a relentless push to higher bandwidth, yes. NVLink seems to be kicking ass & PCIe is super struggling to keep any kind of pace absolutely, but it still seems wild to me to write off CXL at such an early stage
CXL has some really real things going for it. It's lower level than PCIe and reuses the same serdes that many systems already have, so it's a natural low cost extension of what we already have, for little intrinsic cost. There's 512GBps of PCIe on a modern Epyc (128 lanes of PCIe 5.0); the proposal of CXL is that simplifying the protocol can let us do more with those links easily. GPUs could use more efficienct protocols to talk to the cpu, with better coherency, that might let cpu and main memory work much better in tandem; not unified memory but no longer so very far apart memories. Large memory or storage pools that nodes can attach to on demand offers some interesting capabilities to disaggregate storage, which mirrors the huge success we've seen with modern disaggregated databases being able to do successfully scale.
CXL 3.1 also introduces host to host capabilities. This makes CXL a potential competitor to Ethernet at some levels of scale. Ethernet is getting around to 1.6Tbps/200GBps but via enormous nics and switches. For sure the CXL switches will also be huge but there's no nic. Imagine a Epyc with PCIe 6.0 speed (8Gbps/1TBps) or 7.0 (2TBps) that doesn't need a nic to talk to fabric, that has much much lower latency than Ethernet, that is an rdma beast. It's insufficient as a full speed AI fabric that can satisfy AI needs, but its a huge speed bump, built right into the core. It's hard for me to reject our ability to find good uses for this in the future, just because it's now how AI works today.
The article very briefly mentions Infinity Fabric XGMI, which is what AMD is floating as what they'll use to compete with NVLink. The article talks about this like it's the future, like it's what will crush CXL. But the same PCIe switch chips that run Infinity Fabric XGMI will also run CXL. Maybe perhaps CXL goes nowhere, IF-XGMI rises, but I tend to think there'll be more sympathetic growth & development among these standards. https://www.servethehome.com/next-gen-broadcom-pcie-switches...
Current PCIe devices work without cache coherency to host memory. Drivers follow a handful of patterns for communicating with the device via PCI BARs and DMA. Cache coherency is not required.
If we're not talking about devices but disaggregated memory, then an application running on a single host doesn't need cache coherency on the fabric because the host already ensures cache coherency between the CPUs. (Is that correct?)
The only case I see where cache coherency might be interesting is for multiple nodes running applications on disaggregated memory. Kind of like a single system image cluster, but it's not a software architecture that has been successful in recent years.
This is something that has puzzled me about CXL. Would be great to understand what use case cache coherency serves. Thanks!
To understand how expensive that is, consider the fact that NVMe durable writes to an enterprise SSD are generally faster than a CXL write to remote ram, but applications still insist on using asynchronous I/O to talk to the SSD. CPU level access to CXL is fundamentally synchronous (because of the X86, ARM and RISC-V instruction sets).
Actual measurements seem to support this: https://arxiv.org/pdf/2303.15375.pdf (see Figure 3)
They don’t provide absolute numbers, so it’s hard to compare with other technologies, but note that you can (in theory; and in prototypes) dispatch and complete multiple NVMe I/Os per cache miss, thanks to amortization. CXL is blocking the CPU for the equivalent of multiple cache misses per I/O.
You could also imagine a DMA engine within the CXL ASIC to assist with paging things in and out of main memory in an asynchronous fashion.
You mention amortization for NVMe, but when operating at high queue depths the per-request latency is larger than at low queue depths.
Fast local NVMe drives aren't an alternative to disaggregated memory. There is no fabric with local NVMe drives. It would be necessary to compare against NVMe over Fabrics, which is quite a bit slower (look at flash SAN appliance datasheets, probably around 100 us latency).
Both the comparison itself and the numbers you mentioned are not what I expected. Can you elaborate and share links to numbers?
You also seem completely unaware of the super scalar and out of order nature of modern processors or why hyperthreading/SMT is even a thing. The way a processor works nowadays is more akin to treating the instruction sequence as a graph for which it is possible to find subgraphs, which can be processed in parallel and that includes memory accesses.
Then finally you didn't consider how frequent a remote memory access or how expensive a remote memory access is without CXL.
Since CXL is cache coherent, you only need to load from remote memory when the data is not in local DRAM or your local cache. This means that remote memory accesses are relatively rare. In the case that you do nothing but remote memory accesses, you have to consider your competition as a baseline. The competition used to be RDMA over infiniband or Ethernet and the total latency for those is significantly higher. This means that if remote memory access is important to you, CXL is faster than the alternatives. You seem to be thinking of a hypothetical situation in which no one uses remote memory accesses at all. That is not a fair comparison.
The reason why CXL is unsuitable for AI is already mentioned in the article. AI doesn't care about cache coherence, only bandwidth but NVLink is even worse at what you are concerned about.
Here is a relevant paper: https://dl.acm.org/doi/abs/10.1145/3529336.3530817