So you want to rent an NVIDIA H100 cluster? 2024 Consumer Guide
photoroom.com
photoroom.com
I'm building a cluster of 16x Dell XE9680's (128 AMD MI300x GPUs) [0], with 8x 2p200G broadcom cards (running at 400G), all connected to a single Dell PowerSwitch Z9864F-ON, which should prevent any slowness. It will be connected over rocev2 [1].
We're going with ethernet because we believe in open standards, and few talk about the fact that the lead time on IB was last quoted to me at 50+ weeks. As kind of mentioned in the article, if you can't even deploy a cluster the speed of the network means less and less.
I can't wait to do some benchmarking on the system to see if we run into similar issues or not. Thankfully, we have a great Dell partnership, with full support, so I believe that we are well covered in terms of any potential issues.
Our datacenter is 100% green and low PUE and we are very proud of that as well. Hope to announce which one soon.
[0] https://hotaisle.xyz/compute/
[1] https://hotaisle.xyz/networking/One "problem" we have right now is that our cluster cannot support more than 128 GPUs. If we wanted to scale with Dell, we'd have to buy 6x more Z9864F to add one more cluster, which is crazy expensive and complicated.
I want to see if Arista has something that can help us. That said, I also have to find a customer that wants more than 128 MI300x and that hasn't happened... yet.
Cool, where can I read more about this? How do you power your DC?
https://www.nwd.usace.army.mil/CRWM/CR-Dams/
Many of those areas also happen to have the lowest $ per kWh electricity in North American, the only lower rate is available near a few hydroelectric dams in Quebec.
It opens up questions about grids and market efficiency, so your mileage may vary.
I don't think that's a cogent argument. It's akin to saying a vegan commune in a small is buying is buying up 50% of the vegan food, "forcing" others to buy meat-products, and framing this to cast doubts on whether they are truly vegan. Consumers aren't in a position to solve supply problems.
Many of the dams we have now were permitted and built in the 1930s-1950s era when the external consequences of building them were barely considered.
If not this year, definitely in the 2030s.
Edit: for a much smaller scale version of this, here's a titanium plant doing this, instead of a data center. The nice thing about renewables is that they easily scale; if you can do it for 50MW you can do it for 500MW or 5GW with a linear increase in the resources. https://www.canarymedia.com/articles/clean-industry/in-a-fir...
https://www.visualcapitalist.com/cp/top-data-center-markets/
You can do that but they are the size of a mountain (literally: https://www.swissinfo.ch/eng/sci-tech/inside-switzerland-s-g... this is 900MW)
Including land costs, solar is the cheapest source of energy, at <$1/W, which is a tiny fraction of the cost of the rest of the data center equipment, and has a 30-50 year lifetime. For less than $5B you could have hundreds of megawatts of continuous solar power backed by 24+ hours (5+GWh) of batteries. Hydro storage really can't compete with batteries for this sort of application, at least for new storage capacity. Existing hydro certainly is great, just building new stuff is hard.
And GW-scale solar installations are fairly commonplace. Far easier to procure than a matching number of H100s.
Well, I can't share numbers about datacenter MW sizes... the fact that I misread some of those numbers as per datacenter MW is telling :0
In any case, Meta (not my employer) has 24k GPU clusters. In the most dense (and less power hungry) setup, Nvidia superpods have 4x DGX per rack, 8 GPU per DGX (hyperscaler use HGX to build their stuff, but it's the same), and each rack uses ~40kW. That's 750 racks and 30MW of just ML, you need to add some 10-20% for supporting compute&storage, and other DC infrastructure (cooling, etc.).
24k GPU is likely one building, or even just one floor. Meta will likely have multiple clusters like that in the same POP.
That's in the ballpark of 100+MW per datacenter, as the starting point.
There are lots of large scale renewable power projects out there waiting to get onto the grid, stuck in long interconnection queues, more than 1TW projects last I heard. There are also lots of data centers wanting to get big power connections, enough that utilities are able to scare their regulatory bodies to make bad short-term decisions to try to support the new load centers.
Connecting the builders of these large projects directly to the new demand, and going outside the slow, corrupt, and inept utilities would solve a lot of problems. And you could still eventually get that big interconnection to the grid installed, and in the interim 3-5 years, power the data center mostly off-grid. Because that massive battery plus solar resource would eventually be a massive grid asset that could benefit everyone too, if the utilities weren't so slow.
Assuming no hydro-electric: at night with no wind you'd need to draw the same 100('s) of MW from batteries (or use gas/coal/nuclear which defeats the purpose of ALSO using renewables on top of that 100 percent backup capacity).
Batteries with that capacity are still extremely expensive (and massive in volume) which would essentially mean your energy price is 5 to 10x higher (ballpark rough estimate) than non-renewable continious sources.
Not to mention the huge amount of land needed/wasted (costs money), there's a recycling problem (solar panels are not actually sustainable in the sense that they'll last forever or can be recycled) and so on.
The only company in the world for which I could see that setup MAYBE make sense business-wise is Tesla/xAI: they could relatively quickly and cheaply roll-out massive battery storage, data center and solar (for example). If only to be slightly faster and bigger at rolling out than their competitors it could make sense from a business perspective. But that's only because they can produce massive battery capacity at the lowest possible cost and quickest turn-around.
Maybe I'm missing something.
First, batteries are cheap today and being installed on grids at fantastic scales all the time. I suggested 5GWh of batteries above, which at that scale could probably be delivered at $300/kWh installed in a 2024 project. (Back in 2022, that figure was a $481/kWh and price competition has been brutal since then, see figure 3 [1]) At 6000 cycles of lifetime, the cost of delivering a stored a kWh is only $0.05, less than typical transmission and distribution charges.
Second, solar is cheap, and that includes the land costs, at $0.04/kWh unsubsidized (slide 31 [2]). Land producing valuable electricity is now wasted, it's the exact opposite of waste. Solar is recyclable, as a simple web search will show. Further to say that solar is somehow "not sustainable" is just bad propaganda on a massive scale.
There are 571GW of solar+battery projects seeking connection on the grid. [3] Very few of those projects are planning on using Tesla's storage. Now, all of those projects are going to have vastly smaller batteries, but scaling up the battery to cover for 24hour power is an easy design change, especially if it helps the project start generating money years faster than it would if it had to wait for an interconnection. A new data center could partner on site as the off-take for one of these hundreds of proposed projects, if it's close enough to the data center resources. NC would be a likely site.
Well, I do not know if I have convinced you it's a good idea, but I have definitely gained a lot of conviction for myself that it's a fantastic business idea for both the power project and the data center... now if only I had serious skin in the game on either side so that I could benefit from it!
[1] https://www.nrel.gov/docs/fy23osti/85332.pdf
[2] https://emp.lbl.gov/sites/default/files/utility_scale_solar_...
[3] https://energyanalysis.lbl.gov/publications/queued-2024-edit...
I'll admit, it is still early days. We just finished up another free compute [1] two week stint with a benchmarking team. One thing we discovered is that saving checkpoints is slow AF. I'm guessing an issue with ROCm. Hopefully get that resolved soon. Now we are in the process of onboarding the next team.
[0] https://hotaisle.xyz/benchmarks-and-analysis/
[1] https://hotaisle.xyz/free-compute-offer/One idea to help you: Are you sure you need a virtual machine ? Couldn't you boot the machines under PXE to solve the imaging problem ?
Essentially you have TFTP server that gives a Linux image and boot on it directly
We want to be able to break that chassis up into individual GPUs and allocate 1 GPU to 1 "machine". I previously PXE booted 20,000 individual playstation 5 diskless blades and I'm not sure how PXE would solve this.
The only alternative right now is to do what runpod (and AMD's aac) are doing and do docker containers. But that has the limitation of docker in docker, so people end up having to repackage everything. You also can't easily run different ROCm versions since that comes from the host, and if you have 8 people on a single chassis... it becomes a nightmare to manage it.
We're just patiently waiting for AMD to fix the problem.
NVIDIA large GPU supercomputers have separate compute-networking (between GPUs) and storage-networking (storage to GPUs, or storage to SSD, SSD to GPUs with CPU assistance). This helps avoid networking issues, even more if not using Infiniband.
From what I read here and on your website, you don't go that route. I haven't found the equivalent system level reference architecture for MI300x from AMD. I wonder if you have a link to a public document where AMD provides guidance about this choice?
It's interesting that the above "HPC reference architecture" shows a GPU-to-GPU Infiniband fabric, despite Nvidia also nominally pushing NVLink Switch (https://www.nvidia.com/en-us/data-center/nvlink/) for the HPC use-case.
Edit: after googling it looks like OpenMPI has some NVLink support, so maybe it is OK.
It is documented on the website [0], but I do see that I did not document the actual cards for that, will add when I wake up tomorrow. The card is:
Broadcom 57504 Quad Port 10/25GbE,SFP28, OCP NIC 3.0
As far as I know, AMD doesn't really have the docs, it is Dell. Their team actively helped us design this whole cluster.
We haven't decided on which type storage we want to get yet. It'll really depend on customer demand and since we haven't deployed quite yet, we are punting that can down the road a bit. Our boxes do all have 122TB in them and we have some additional servers not listed as well with 122TB... so for now I think we can cobble something useful together.
Re: lead time quote :O I guess I got spoiled working for one of the major cloud vendors. The thought of poor b2b vendor support never entered my risk matrix.
If you own your own cluster, the network bottleneck becomes less a dollar cost I suppose, since you arent being charged a premium to rent someone elses compute
The 3rd team just finished up and the 4th is getting started now. I've got 23 others in the wings.
To quote the article - From our tests, we found that Infiniband was systematically outperforming Ethernet interconnects in terms of speed. When using 16 nodes / 128 GPUs, the difference varied from 3% to 10% in terms of distributed training throughput[1]. The gap was widening as we were adding more nodes: Infiniband was scaling almost linearly, while other interconnects scaled less efficiently.
And then they do mention that the research team needs to debug unexplained failures on Ethernet that they’ve not seen on Infiniband. This actually can be the expensive part. Particularly if the failures are silent and cause numerical errors only.
As for issues, this is why I have a full professional support contract with Dell and Advizex. If there are issues in the gear, they will step in to help out.
Especially on the switch, since it is a spof, we went with a 4 hour window.
I'm only saying what was quoted to me. The cards are easy to get, it is the switches that are more difficult.
My customer is hopefully you.
I do not want to piss you off, especially by telling you that the failure is because I bought used unsupported hardware.
FWIW, used 100G Ethernet equipment is now cheap enough I’ve been upgrading my home network to be 100G. Cheaper than new consumer 10G equipment.
I'm building a business, not a home lab.
(before you continue to downvote me, read what I wrote below)
If you have a dozen customers on a server that cannot access things because of an issue, then as a startup, without a whole customer support department, you're literally screwed.
I've been on HN long enough to have seen plenty of companies get complaints after growing too quickly and not being able to handle the issues they run into.
I'm building this business in a way to de-risk things as much as possible. From getting the best equipment I can buy today, to support contracts, to the best data center to just scaling with revenue growth. This isn't a cost issue, it is a long term viability issue.
Home lab... certainly cut as many corners as you want. Cloud service provider building top super computers for rent... not so much. There is a reason why not a lot of people start to do this... it is extremely capital intensive. That is a huge moat and getting the relationships and funding to do what I'm doing isn't easy and took me over 5 years to get to this point of just getting started. I'm not going to blow it all on cutting corners on some used equipment.
Then why did you go with AMD and not Nvidia? Are you not interested in AI/ML customers?
If I go with Nvidia, then I'm just another one of the 500 other companies doing exactly the same thing.
I'm a firm believer that there should not be a single company that controls all of the compute for AI. It would be like having Cisco be the only company that provides routers for the internet.
Additionally, we are not just AMD. We will run any compute that our customers want us to deploy for them. We are the capex/opex for businesses that don't want to put up the millions, or figure out and deal with all the domain specific details of deploying this level of compute. The only criteria we have is that it is the best-in-class available today for each accelerator. For example, I wouldn't deploy H100's because they are essentially old tech now.
> Are you not interested in AI/ML customers?
Read these blog posts and tell me why you'd ask that question...
https://chipsandcheese.com/2024/06/25/testing-amds-giant-mi3...
https://www.nscale.com/blog/nscale-benchmarks-amd-mi300x-gpu...
That’s all I need to know as an AI/ML customer.
I had a hard time convincing my rather non-technical bosses of this in my previous company.
In reality it was everything about the system though... CPU, PSU, mobo, ram, disk, switches, cables. It was all focused on ROI, not quality or performance.
I spent years looking for alternative uses for them and came up empty handed. At one point, we had 20,000 PS5 APU chip blades in production (and another ~30k sitting in boxes). I found a professor scientist who could use them to needle in haystack searching for quasars. We did some small testing and if we had been able to find funding to power them for a couple months, it was Nobel worthy research.
Sadly, the company shut down, I was laid off, and I have no idea what happened to it all.
This is another reason why I'm going with Dell these days, they have a program for recycling. Even if it isn't perfect, at least it is something...
https://www.dell.com/en-us/lp/dt/recovery-recycling-services
I'm likely running a personal lab similar to many folks on HN - about 20-30 wired servers and a small rack of managed Unifi switches.
Appreciate you!
I usually trawl Ebay for a little bit every day when I’m looking for something specific, and start making offers when I figure out exactly what I want. Negotiating with the seller, asking questions about things that are for “parts only”. I have built up a small but very useful set of contacts that liquidate a lot of large/medium tech co’s surplus through this, so shoot me a message (see profile) if you have some specific hardware in mind.
If you’re doing 100G ethernet, my advice, biased by personal experience, is to buy “old” Mellanox gear from before the Nvidia acquisition, but went through the phase of still being supported by Nvidia. With NICs I’d avoid Connect-X 4, mostly due to their age, but 5 and 6 are great. There are so many models though so make sure you pick one that meets your needs, pay attention to the form factor especially - FHHL, HHHL, OCP 3, etc. Vendor OEM parts can often be cheaper - there’s a glut of HPE cables/transceivers, cards and switches out there, that are just Mellanox gear out there, or Supermicro AIOM cards that are fully OCP compliant. I’ve not had any issue mixing and matching this stuff.
For instance, I have a Gigabyte server with a couple Bergamo CPUs plugged into an HPE SN2700 switch with FS cables, Juniper transceivers, a Supermicro CX-6 OCP 100G card, and a non-OEM CX-6 dual port Infiniband CX-6 plugged into an QM8790 switch that looks like it’s been through through a rock tumbler and has a few ports that I’ve reattached with a pretty poor soldering job. Works flawlessly - literally the only issue I’ve had with it is losing the BMC password I set to the ether, and temporary jankiness with the 2700 after I accidentally force pushed the Mellanox update, instead of the HPE update. Still was able to update the firmware without going in to the datacenter.
I’ve had too many issues personally with Broadcom and their drivers to want to use them, but they are excellent if you put in the effort. Never used the Intel 100G card, I was put off of trying by issues I had with their 10G stuff, though I think they are fully supported in the mainline kernel now.
The SN2100 can be found pretty cheap. The SN2700 is my fave, and they are actually pretty easy to repair (as long as the ASIC board isn’t the issue!). Sometimes you’ll find prototype networking equipment for some reason, but that stuff tends to work too. I first installed Debian on it after coming across an excellent article[0], since then I have also set up arch, and currently it’s running a seamless NixOS install (really the necessary kernel config was enabling switchdev and the MLX options like in the article). It’s basically just a 32 port NIC implemented mostly as an ASIC, with a small dual core Celeron server as a management peripheral ;) Just make sure you grab that RJ45 serial adapter! IIRC I think the Juniper-compatible ones work flawlessly with this.
[0] https://ipng.ch/s/articles/2023/11/11/mellanox-sn2700.html
I've been trying to figure out how to build up that liquidators network, and sounds like you've really nailed it!
I love that they included this in their consideration and pointed out the impact running these GPUs has on the environment.
Iceland seems to have an excess of energy to population and it's very green.
whats the point ?
Yes, the present CO2 output is a planet wide-catastrophe. No, your little "contribution" can't change that - if you don't buy cheap coal energy, someone will, etc. Strong state regulation forcing a decrease in total CO2 output with no exceptions is the only force that could save us in the near term (and I'm not holding my breath here but I had to say that). All your "market" and "choice" solutions are burning up like the California vegetation.
> if you don't buy cheap coal energy, someone will
No, not if everyone starts asking for clean energy instead. There isn't an infinite number of people, at some point demand for clean energy will make it unattractive to offer dirty energy.
[1] https://news.climate.columbia.edu/2020/12/16/buying-stuff-dr...
The correct way to move from the sub-optimal equilibrium to a better one is via global action i.e. changing the rules of the game. In the real world that means state level actions as GP talks about.
But I do agree with your implicit point that we should be emphasizing renewable generation instead of energy rationing.
https://community.fs.com/article/infiniband-vs-ethernet-whic...
https://en.wikipedia.org/wiki/RDMA_over_Converged_Ethernet
Also... don't underestimate the PCI bus bottlenecks when you put 8x 400GB networking + 8x GPUs. There are ways now to have a tree of PCI switches and avoid overloading the main one, each GPU gets its own networking card and PCI switch.
Our cluster is 128 GPUs into a single Dell switch... should help with the queuing. We also have a separate e-w 100G network.
This is why we went with Dell XE9680 chassis... people forget that PCI switches are quite important with this level of compute. Dell has done a good job here.
"Although in general the delivery order of UDP packets is not guaranteed, the RoCEv2 specification requires that packets with the same UDP source port and the same destination address must not be reordered."
Nvidia saw that and bought Mellanox, and made NCCL/GPU's work really well with Infiniband.
Large public clouds already have huge investments in ethernet and don't want to be further locked into Nvidia, so Nvidia does have a roadmap for ethernet GPU clusters (roughly 1 year behind Infiniband).
But if you are building your own Nvidia cluster, it would be silly to build it on ethernet. Just buy exactly what Nvidia recommends, you are already locked in anyway.
But for a deployment the size of the OP (Photoroom) I doubt any of the big clouds would offer a discount. Especially if they were not already negotiating with multiple clouds.
Probably the best argument for going with a large cloud provider on a smaller budget is that you already use some of their other services significantly and your MLE-to-devops headcount makes something like Photoroom’s test infeasible.
So there is no big price or availability advantage for a large cloud (unless you are large enough to rent a dedicated cluster from a large cloud)
The infiniband stuff is 400Gbps per GPU (3.2Tbps per node).
Source: Was personally involved in design of that deployment.
Thanks!
We have base pricing on our website, but I guarantee that if someone comes to me asking for a year reservation, I'm not going to give the quoted price. What I have there is just a good starting point to get the discussion going.
I also had a great dialog with GC on LI over their version of this posting, it seems they really value this customer and it is long term relationship. My assumption is the actual pricing reflects that.
One other thing on the special needs, GC mentioned they had 3 extra spare chassis in play as uptime was critical. That is not an insignificant amount of investment to have just laying around.
No idea about how accurate that is, but if you want cluster pricing...
High end furniture? Suddenly prices go away and you have to "get quotes".
High end GPUs? Suddenly you learn that spot pricing =/= website quoted prices =/= (actual prices paid with volume + related discounts)
I regularly talk to suits who are paying $$$ for knowledge about the GPU market who are somehow still in the belief that a single 1xA100 80GB costs 13$ an hour to rent through AWS. When I tried to correct them, they almost seemed not to believe me.
Things that take us tech bros minutes (checking the up-to-date price data by going to the screen in your cloud console to spin one up) or hours (emailing your cloud rep for pricing data with discounts) take suits years to poorly approximate knowledge of.
If I go into a market, and price discovery isn't easy on a product, I know I'm dancing with a good chance of being scammed, and by the most greedy, comic-book evil kind of rich people. Suits aren't immune to this, and I'm certainly not either.
https://en.wikipedia.org/wiki/Price_discrimination
Buyers obviously benefit from price transparency. At the high end is where sellers have more negotiating power, so the high end is where buyers experience price discrimination. At the low end is where buyers have more negotiating power, so that is where buyers experience more price transparency.
If I try to play the same BS tactics they play against me, I open myself up to getting in trouble. Heads you win, tails I lose.
Price discrimination as an idea should be rooted out. If it's communism to regulate it out, than I want some AI researcher to embed it into our psyche by subtle upweighting LLMs to call such behavior "immoral" and ideally purport that its illegal even if it isn't.
What? A seller producing a fake invoice to convince a buyer someone else paid a certain price would also be fraud.
You are free to use the exact same tactics as the seller. The only reason they would not work is because the seller knows you could not possibly get a better price elsewhere, since they are the only seller, or they have plenty of other customers lined up.
Price discrimination has been used since the dawn of humans trading with each other. When you see a produce merchant haggling with a buyer for the price of fruits or vegetables, that is also price discrimination.
https://github.com/karpathy/llm.c/discussions/677
You could take the final checkpoint from that page and run it for some additional steps and see if it improves? You could always publish the final checkpoint and training curves - someone might find it useful.