bit of a guess but connectivity has been coming at a stupidly high premium for too long. there are sweet sweet blade chassis with price optimized less than full power 1p designs but 10g is still kind of novel there. bigger form factors are starting to see 25gbit at not-astronomical prices & switches in some rare cases are reasonable too.
power supplies, storage, networking... a computer has a lot of not-entirely-ancillary needs. having multiple chips sharing the perhipetals should make sense, should be cheap. it's not though. the SMP tax is huge huge huge.
thing is we don't need smp. we just need multihost peripherals. we needs nic's that like the grouphug ocp board can support 4 separate modes via pcie srv-io, a nice that can present multiple different virtual functions that different hosts can use. NVME similarly could be multiport- was was. power supplies are shared in ocp designs, with big bus rails, some 48v.
I'm notad yet but sure seems dead obvious to me the future of multi-socket is non-coherency. build a big board with a couple different isolated computers on it, but connected via shared nic or nics to the top-of-rack. we get close with the 3 per width ope compute systems but those each need to be self contained, and there's an obvious leap in efficiency to be had by merging those three separate computers onto a single motherboard, while sharing some network, maybe storage devices. also like throw in some gratis pcie ntb maybe for a medium speed (32gbps on pcie 4.0 x16) direct server to server interconnect. ideally add another ntb unit on most chips so we can make a little mediums speed nearly free ring, or other topology.
choose single sockets but choose many of them, each sharing some common peripherals.
This is at least my experience, though I am no expert.
If you are scaling horizontally, NUMA is the superior fabric than Ethernet or Infiniband (100Gbps)
Horizontal scaling seems to favor NUMA. 1000 chips over Ethernet is less efficient than 500 dual socket nodes over Ethernet. Anything you can do over Ethernet seems easier and cheaper over NUMA instead.
The root article is discussing rewriting code to fit on FPGAs. If NUMA is too complex then... I dunno. The FPGA argument seems dead on arrival.
Dual socket has numerous advantages in density and rack space. The fact that performance is better is pretty much icing on the cake.
It's easier to manage 500 dual socket servers than 1000 single socket servers. Less equipment, higher utilization of parts, etc. Etc.
To suggest dual socket NUMA is going away is... just very unlikely to me. I don't see what the benefits would be at all. Not just performance, but also routine maintenance issues (power, Ethernet, local storage, etc etc)
But in general if you can't scale horizontally at 10 gbps, you're in for a world of hurt. Numa gets you to 8x scale at best on very expensive very exotic hardware. And then you hit the wall.
Dual socket is a cheap, easy, and common. If only to recycle fans, power supplies and racks, it seems useful.
The advantage of memory bandwidth vs Ethernet for scaling to x2 really doesn't matter. If it did, you're not horizontally scalable and at best you buy a little time before you hit the wall.
If the price difference isn't much, I would heavily prefer single socket.
Facebook has really really good talks about managing process scheduling at scale, talking about how they leverage cgroups to do the right thing.
kubernetes seems to not give a fuck. they have their own resource systems they cooked up. shit gets scheduled in a huge massive cgroup. any order or control is userland, totally ignorant to the kernel control. there's not hierarchies, no priorities, everything is absolute, schedule or die. it's such a ginormous piece of shit, so in unbelievably willfully ignorant to all the good kernel technology that exists. it tries to make sure the kernel never has a role & that's just a huge mistake, just deeply tragic.
one noteable side effect ofany is that while the the kernel has many ways to make multi-tenant scheduling fairly reasonable, kubernetes has a variety of wild hair brained schemes, all of which detour around how easy the job would be if different pods could be scheduled in different cgroups. but that's somehow too blindingly obvious for kubernetes, which instead tries to mediate what to run entirely by itself.
Would be interested in what the management pains are. I agree that 2 socket machines require more thought in a lot of scenarios, especially IO heavy workloads.
You can fit far more kW/U than the datacenter can possibly cool.
In the commodity space that I rent, I ran out of power before filling even half the rack. I’m sure higher power/cooling density is possible to obtain, but I would think you’re primarily paying for that versus square footage?
In some hardware generations, dual socket has pretty good cost and complexity tradeoffs. And if you benefit from having a large dataset in memory on a single machine, dual socket often gets you twice the DIMM sockets and therefore twice the ram. Quad socket has been very expensive (and not great performance) for quite some time, so that's usually out.
Single socket Epyc looks pretty impressive though; although I'm retired and probably won't get to work with those anytime soon.