Nvidia L40S is a Nvidia H100 AI alternative
servethehome.com
servethehome.com
'For those OEMs to win larger H100 allocation, Nvidia is pushing the L40S. Those OEMs face pressure to buy more L40S, and in turn receive better allocations of H100.'
Lots of articles and opinions popping out about the L40S as a serious H100 alternative at the same time. The people buying these kinds of cards have already been looking seriously (10-12k easy to source against 35k impossible to get?) and rejected them because for training memory bandwidth, intranode (nvlink) and internode (pcie5 and... Nvlink too?) interconnect, and tensor core performance are critical.
Now if your application is fp32 and compute bound (or mostly compute bound) or if you can make it fly with cuda cores and bit of tensor cores, a bit of work yes, and you're still on Eth200G (so pcie 4 is OK) the L40 is an incredible beast and you can build incredibly dense compute nodes with up to 700 TFLOPS which is just mad.
In inherits the... quirks of the 4090, namely the loss of NVLink (which the 3090 had) and basically the same memory bandwidth as the 3090. Also, apparently the crippled tensor core performance for the L40. And the price hikes.
I'm not really trying to sound bitter, but its hard to talk about Nvidia's lineup now while ignoring the artificial segmentation or price inflation and some small corner cutting. They would make Intel at their anti-competitive peak blush.
https://fortune.com/2023/11/01/nvidia-shares-fall-report-us-...
>The restrictions were supposed were supposed to only come into play on Nov. 17, 30-days after the US first announced it. But in a filing on Oct. 24, Nvidia said it was informed that the rules were effective immediately and that it would affect shipments of Nvidia’s A100, A800, H100, H800, and L40S products. The 800 series chips were designed specifically for the Chinese market to circumvent the earlier iterations of the export control rules.
Nvidia repurposes their gaming silicon every generation, and this one is no different.
If you are talking about the L40 vs L40S, I'm not sure. Does the L40 just barely sneak by the export restrictions?
> exceeding certain performance thresholds (including but not limited to the A100, A800, H100, H800, L40, L40S, and RTX 4090).
Seems like it's included in the SEC filing[0]
[0]: https://www.sec.gov/ix?doc=/Archives/edgar/data/1045810/0001...
EDIT: Ah interesting, so it is included.
It has been a pretty painful experience speaking to the companies due to shortages and because of all the limitations of non-H100 chips. For getting a machine with 4 H100s and NVLINK, I was told I'd have to wait one year. With that long of a wait, I might as well wait for the next generation. Without NVLINK, I could potentially get this machine sooner. I struggled to find any reports of the additional complexities and how much slower the machine would be.
I'm basically looking to buy two machines. One for under $40K and one for up to $100K. The $100K machine is intended for training LLMs, and the other would have multiple GPUs where the other would just have a bunch of GPUs for my PhD students to use. For the $40K machine, I'm being told by companies I could only put in 2-3 L40S cards, and they are really pushing these as the only viable option.
In contrast, the last time I had $40K for hardware I got two machines with 4 RTX A5000s and companies were a lot more responsive and seemed more helpful. It is unlikely that I'll have $100K for hardware again, so I'm very reluctant to go with cloud computing since that decreases my hardware budget and I need these to last 3+ years.
I wish I could go with 4090s, but the limited VRAM and that NVIDIA has disabled functionality needed for multi-GPU training makes that pretty much a non-starter.
The FP8 and higher cuda level is nice I guess, if your toolchain supports it.
Lambdalabs if you're lucky will grant 1x H100 at $2/hr :)
Every other serverless provider charges 3-5X the typical hourly cost.
$13,744.95
I think AMD could theoretically make a 96GB inference card this gen with their modular memory controller...
This is 2xGPU, and two very expensive GPUs with HBM.
I was thinking they use the 7900 XTX die, and swap out the tiny memory controller dies or somehow double them up (which is infinitely cheaper than redoing the whole big die). They could massively undercut the L40 in cost while still doubling up the vram to 96GB, and doubling the bandwidth as a happy side effect.