122 karma · joined August 30, 2025
https://docs.nvidia.com/cuda/gpudirect-rdma/index.html
The "R" in RDMA means there are multiple DMA controllers who can "transparently" share address spaces. You can certainly share address spaces across nodes with RoCE or Infiniband, but thats a layer on top
Our racks are provisioned so that there are two independent rails, which each can support 7kW. Up until the last few years, this was more than enough power. As CPU TDPs increased, we started to need to do things like not connect some nodes to both redundant rails or mix disk servers into compute racks to keep under 7kW/rack.
A single HGX B300 box has 6x6kW power supplies. Even before we get to paying the (high) power bills, it's going to cost a small fortune to just update the racks, power distribution units, UPS, etc... to even be able to support more than a handful of those things
I read it originally when it came out, didnt have a chance to respond initially then came back to look through the comments. If there was anything was changed, it's not obviously apparent...
I hope the "founders" go back to the drawing board and retry
Absolutely not! Every major FOSS license has copyright as its enforcement method -- "if you don't do X (share code with customers, etc depending on license) you lose the right to copy the code"
Your wife is making a business, and you want to write some code to help.
Then suddenly your requirements balloon to multiple concurrent users, needing to have a system tray icon and then also the ability to take this code and sell it to other people. Wow this project is suddenly complex!
This is just "I need to be able to scale infinitely" written in different words. The complexity comes from wanting a ton of things before they're actually needed (with the wrinkle of wanting to use some previously written scheduler for this project.