Manufacturers do segment the market, but also a 4090 is a totally different, cheaper piece of silicon than an H100. Not really just a price increase. HBM and the requisite packaging is very very expensive.
Why does CXL not make sense? Coherency is great!
NVLink is great for say 4x + gpus. With two 4090s , the effective speedup is the same as with two NVlinked A100s from my testing. So it’s really not a big deal it was pulled out.
Consumers really really don’t build dual GPU systems anymore. Gaming stopped using SLI/crossfire over ten years ago. So AM5 and Intel’s 1700 socket are designed to that idea, and price point. Why is socket SP3 so insanely massive? All those PCIE lanes and extra memory bandwidth take pins. More costly boards, more costly socket, etc.
Also, how are you going to cool 4x 4090s on air, when they are next to each other? That’s 1800w of heat!! How are you going to power it? That’s more than you average American house can power with a 15A 120v circuit!
I just built a dual 4090 system, while I was waiting for my waterblocks to come in, I was testing everything on air. The second card getting fed hot air by the first eventually had to downclock to half the clock of the first card, in order to keep the GPU under 90c at 100% fans. The waterblocks are single slot and go into a server class SP3 epyc board with 4x 1 slot spaced pcie. I have a 1600w psu that has a special plug, and it sits on the special single plug circuit that’s for an AC.
Nvidia is hugely innovative, and a 4090 core is really not simple at all. The cores were simple back in the 8800GTX days, kinda, now, you have pretty significant scheduling resources and tensor cores and different kinds of ALUs and a bonkers massive register file. CUDA is their secret sauce, way easier to use than openCL.