New Compute Engine A2 VMs–First Nvidia Ampere A100 GPUs in the Cloud
cloud.google.com
cloud.google.com
1. https://developer.nvidia.com/blog/nvidia-ampere-architecture...
I've also read that the A100 can deliver 4x the performance on some DL workloads.
Intel Xeon Gold 5120 14Core @ £1200 used
renting 1Core @ £1.20/mo.
would take 6 years to fully pay for that single core alone, and that's excluding the 512MB RAM, 10GB SSD and unlimited 400Mb/s bandwidth
it might throw the comparison off if there was variation in performance at different times, but there isn't
cores scale with RAM, you cannot rent 32 with 512MB
Hetzner dedicated cores come with 4gb of memory per vCPU for reference.
1.2 GBP/mo for a dedicated core is wildly off market price. There's no way they'd be able to make a profit, so either they're going out of business tomorrow or it's not dedicated.
Still comes out cheaper than cloud offerings.
I think the right way to think about the economics here is either “I would pay $X/hr for this short-lived job” or “I want to compare with buying it” (3-yr committed use discount in our case, RIs / Instance Savings Plan for AWS). Unless you are an ML research lab (Google Brain, FAIR, OpenAI, etc.) or an HPC style site sharing these, you won’t get 100% utilization out of your “I just bought it” purchase. Worse, in ML land, accounting math about N-year depreciation is pretty bogus: if the A100 is 2.5x faster, you’d have been better off with a 1-yr CUD on GCP and refreshing, rather than buying Voltas last year.
One amusing thing that’s not clear about “just buy a DGX” is that many people can’t even rack one of these. At 400 watts per A100, our 16x variant is 6.4 kW of GPUs. That’s before the rest of the system, etc. but there are (sadly) a lot of racks in the world that just can’t handle that.
But also, you can rent a single slice of one for tens of seconds and then walk away. I personally probably do about 20-ish GPU hours a month while dabbling (and only because I’m lazy and have my GPUs attached while I’m working, even if it’s just debugging the Python bits). For a V100 that’s $50/month, which is in the noise compared to dealing with owning infrastructure (and would be even less if I’d get by with a T4).
Makes you wonder what kind of power costs you would incur running one 100% utilisation. Certainly even with best prices, be looking at several thousand a year I would of thought, that's not even factoring in provision. Which would mean 3 phase power for that type of load and then you have to balance out the phases. So many little details that become more an issue when you start getting to datacenter level power usage. Then UPS load/capacity costs/planning, networking. So whilst the costs of these units are high, the other costs that add up, sure do add up fast.
I’d say keeping a spare system in case of failure is probably a bigger deal than the $10k/yr to house it :).
Obviously there are factors going in to that -- I could live without paying the Tesla tax (didn't need to virtualize, didn't need the vram, did need the fp64), I bought used, I didn't have a problem keeping it fed, I didn't need to burst, etc, but my point is that for some GPU workloads the cloud GPUs are really expensive and the break-even utilization is far south of 100%, more like 5%.
It's really expensive, and I think I should lean into buying hardware at this point.
I want to build a high end GPU rig, but was wondering how easy the setup was. I've only built "consumer" systems before (2x 1080Ti). Is there any appreciable difference?
Do you have a single card? Multiple? What motherboard do you use?
Do you have any takeaways or resources you can share?
I'm not using this for machine learning, so you might want to talk to someone who is before pulling the trigger. In particular, my need for fp64 made the choice of Titan V completely trivial, whereas in ML you might have to engage brain cells to pick a card or make a wait/buy determination.
The “consumer” parts are certainly popular ML workstations, and rightly so.
We're in Colovore, which has a fantastic power density and is running roughly 1000k DGXs in their datacenter. It really wasn't all that difficult to get up and running. For us it made total sense, but we utilize physical hardware completely and have to scale into the cloud fairly regularly.
Did you move all your other gear into Colovore? (That’s one of the challenges, you often need to be close to the rest of your systems / data)
We also run at 100% utilization 24/7 and frequently use the cloud for burst before we go out and buy more cards/servers.
Colovore was super easy to get up and running and we will save at a minimum hundreds of thousands of dollars with this setup over exclusively using cloud instances.
Oh, so cloud providers are just giving away free resources? How generous!
Seriously, let's do the math:
If a V100 instance cost $9000 new and you bought it a year ago, you could still sell it today for over $3000. On AWS, an instance on a 1-yr CUD costs more than $1400/month, for a total of over $16000. You don't even need 50% utilization to break even. It doesn't matter how fast the A100 is.
The same thing happens with these large instances. Because these instances are so much bigger that you won't use them for the entire month. You'll use them to get results within hours or days.
Because of that, seeing this A100 announcement is just a bummer, as I fear it'll be just another "resource unavailable" GPU...
Sorry to hear you have experienced this.
Customers can experience stock outs sometimes based on a variety of factors but we can surely help you out as we have T4 GPU capacity available and like you said, we may direct you to a different zone or region. Open up a support case on the issue and we can help you out.
Lots of architectural changes like MIG, new floating point formats, etc. Great to see GCP getting VMs out pretty soon after launch so people can start kicking the tires.
It does seem a little fishy to me that NVIDIA often boasts with figures like 10x performance upgrade whilst in practice those are only possible if you use one of their non-default float types which are hardly supported in most deep learning libraries :(
the "A" series on AWS = AMD instances
The "A" series on GCP = Nvidia instances.
I know - probably on no ones radar at all :)
Even worse is that for GCE, A was for AMD originally (and N was for iNtel). In any case, this A is for Accelerator.
However, I’ll note that the 16 A100s here are way more expensive than the cpu cores (and we can just run vanilla VMs on those left over cores if really needed).
Of course second generation uses names like m6g to keep us on our toes..
Meanwhile M5A, as you note, is AMD.
Still not sure why competitors don't have their GPU's offered in cloud services. At the time, it seemed the alternatives offered a more economical alternative. I was building a GPU version of Hashcash at the time fwiw.
AMD has many GPGPUs, but while they have kind of ok-ish raw numbers, their software stack makes quite poor use of them.
Users care about perf/$, and nvidia is king there. For deep-learning and other apps, hardware without software is useless.
I think this is also part of the problem. AMD has a nearly function-for-function compatible GPGPU stack with CUDA in ROCm/HIP that fixes many of the pitfalls in OpenCL, and even provides automated translation tools for CUDA developers switching over.
However, this stack is so poorly documented and marketed that the mass majority of developers still associate AMD == OpenCL. Add onto that the lack of Windows, MacOS and Navi GPU support, and you've removed a ton of opportunities/incentives for folks to experiment with the platform.
This is really frustrating because in theory AMD gpgpus have a better architecture for ml.
My guess is that it's also tough to get deals with AMD for their workstation class products which are competitive with Nvidia's rtx line, and the price/performance delta (both in terms of unit cost and power cost) is not huge.
There are some consumer class amd gpus that have attractive price/performance, but they are being sunsetted and the future driver support is questionable. I looked at the Radeon 7, but it's bazel is a few millimeters out of pci spec so it literally doesn't fit in a 8u GPU chassis, and it's an architecture with an uncertain future.
I'd got the impression that Vulkan would bridge that gap, but it's quite a low-level API and verbose. Would be interesting to hear comments on that, perhaps just for the case of the hosted GPU market being more competitive.
Vulkan also suffers from being too late to the game.
NVidia moved CUDA into a polyglot runtime early in the game (around version 3 I think), and took a 10 year effort to redesign the hardware for optimal mapping to C++ memory model.
Vulkan might eventually get it via SYCL, now that SYSCL has turned into backend agnostic mappings for C++ heterogeneous computing.
But even if SYSCL turns out to be adopted across the industry, Vulkan will just be another backend alongside CUDA, Rom, FPGAs,....
The cloud it's a bit scary from learners perspective because it's not exactly clear how much power one would need to practice the core concepts and see actual results.
In the other hand a pc build, it's an upfront investment, less scary because it's a fix cost, but also feels risky in case things progress quickly and hardware gets outdate soon.
Honestly, I’d start with Colab until you decide you need “dedicated hardware”. It’s better to focus on learning before you decide “Okay, I’m serious about this now”.
Only approach these types of instances once you actually have a handle on how much power you need.