Remember retail GPUs are much more powerful, much cheaper and much more available than cloud GPUs.
Just go to the store and buy one.
Remember retail GPUs are much more powerful, much cheaper and much more available than cloud GPUs.
Just go to the store and buy one.
I agree with your basic point about building vs renting, but retail GPUs aren't simply more powerful. Retail GPUs are just as fast but have a lot less VRAM and memory bandwidth. The RTX4090 has 24 GB of VRAM with 1 TB/s of bandwidth. The A100 has 80 GB at 2 TB/s. For some tasks, the memory and bandwidth are the strongest constraints.
Acquiring the latest technology, such as an A100 GPU, individually from a store can indeed be quite expensive, with prices around $10,000 per unit. Additionally, setting up and maintaining a home lab to scale from one GPU to thousands can be a significant investment in terms of both cost and resources.
In contrast, our platform provides access to the latest GPU technology without the need for users to individually purchase and manage the hardware. We offer the scalability to deploy multiple GPUs as needed and scale back down when required, making it a more cost-effective and flexible solution compared to setting up and maintaining a personal lab.
Furthermore, it's important to note that our GPUs are directly dedicated to each virtual machine, ensuring that the power and performance are not compromised by sharing resources. This ensures that our GPUs provide the full capability and performance expected, making them just as powerful as individual units.
Now, "running" may also mean different things. For example, you may want to do things s.a. performance diagnostics (in order to understand if your code uses resources efficiently), and then you'd need stuff like NVML, DCGM and co. Consumer-grade hardware might not be supported, or might not in principle support diagnostics collection / instrumentation. Or, if you have multiple GPU-dependent workloads that cannot saturate your resources -- you might think of MIG as being a way to address that... and, again, consumer-grade GPUs won't help you here...
I'm not saying you shouldn't try self-hosting. I'm actually all for it. But, you also need to be mindful of the pros and cons. NVidia must have some reason to pitch the datacenter family of GPUs to, well, datacenters. They aren't just blowing up prices.
Also, to give some sense of comparison:
Titan X ~= 3.5K CUDA cores.
V100 ~= 5K CUDA cores.
A100 ~= 7K CUDA cores.
H100 ~= 18.5K CUDA cores.
It probably doesn't translate directly into H100 being six times as fast as Titan X, but, I hear that these GPU workloads might be lengthy...
For some of these, you could obviously pay with your time. Sometimes that time is very valuable, and sometimes you have a lot to spare.
Also, the amount of VRAM (3x)... Sometimes having too little of it means having to re-write the program, or it could dramatically impact the speed. Similarly for bandwidth (8x). But, again, if the kind of workload you have isn't constrained by either, then you could probably win by running on more smaller / older GPUs.
So the consumer would care because the service is lower cost.
It's more expensive to rent a GPU than to buy it.
This whole comment thread started because GP wrote:
>Remember retail GPUs are much more powerful, much cheaper and much more available than cloud GPUs.
So, still, why would a consumer care about performance per watt?
(ofc assuming you're not going to use it for like a day)
And GPU-GPU interconnects are important for this type of training so putting as many GPUs as possible close together is necessary.
If you have conducted your tests on shared instances and have observed such differences, then it's understandable. However, I want to emphasize once again that we provide dedicated virtual machines, which means that the resources are exclusively allocated to each user and not shared.
At least in the past there was a tendency by NVIDIA to run their professional GPUs at a lower frequency, offering reliability and correctness guarantees as tradeoff. Meanwhile most of the retail/gaming GPUs came overclocked and almost outright guaranteeing that they would start to crash and fail if you tried to run them 24/7.
Our intention is not to discourage anyone from purchasing their own GPUs. We offer our GPU cloud services as an alternative for those who may not have the resources or prefer the convenience and flexibility of renting GPUs on-demand. We apologize if the discussion took a different direction, and we appreciate your understanding.
With the 4090 being a fair amount more powerful, it might give the a100 a challenge for about 10% of the cost. That’s still about 10k hours of renting one with this service though.
[1]: https://www.nvidia.com/content/dam/en-zz/Solutions/Data-Cent...
[2]: https://images.nvidia.com/aem-dam/Solutions/Data-Center/l4/n...
[1] https://timdettmers.com/2023/01/16/which-gpu-for-deep-learni...
This is a very silly question... the answer is the same as with virtually anything else: real estate, transportation, wedding dresses...
I mean, you need to compare the price that will ultimately depend on the kind of thing you are doing... that's all. Sometimes it makes sense to rent, other times it makes sense to buy. There's no one right answer.