Are GPUs Worth It for ML?
exafunction.com
exafunction.com
Perhaps I'm a bit biased towards all kinds of self-supervised or human-in-the-loop or semi-supervised models, but the notion of discarding large amounts of good domain-specific data that get processed only for inference and not used for training afterward feels a bit foreign to me, because you usually can extract an advantage from it. But perhaps that's the difference between data-starved domains and overwhelming-data domains?
The other is that for various reasons the customer doesnt want to share their data (or at least have sharing built into the inference system) so even if you'd like to have everything they record, it's just not available. Obviously something to discourage but it seems common
If I play chess on my computer, the games I play locally won't hit the Stockfish models. When I use the feature on my phone that allows me to copy text from a picture, it won't phone home with all the frames.
We have a 7 figure GPU setup that is running 24/7 at 100% utilization just to handle inference.
Forgive my ignorance.
That isn't because we aren't training that often - we are almost always training many new models. It is just that inference is so computationally expensive!
Not that anyone should think any aspect (training nor inference) is cheap.
Marketing blogspam like this is always targeting big(not Google, but big) companies hoping to divert their big IT budgets to their coffers: "You have X million queries to your model every day. Imagine if we billed you per-request, but scaled the price so in aggregate it's slightly cheaper than your current spending."
People who are training-constrained are early-stage(i.e. correlate with not having money), and then they need to buy an entirely separate set of GPUs to support you(e.g. T4s are good for inference, but they need V100s for training). So they choose to ignore you entirely.
Wouldn't that depend on the size of your customer base? Or at least, requests per second?
That's what I've seen in my experience, but I concur that there might be cases where the ML is a more-or-less solved problem for a very large customer base where inference is more. I've rarely seen it happen, but other people are sharing scenarios where it happens frequently. So I guess it massively depends on the domain.
The split definitely depends on what you're doing past developing/deploying.
(Source: https://kstatic.googleusercontent.com/files/2f51b2a749a284c2...)
Training was always GPUs (for speed), non-spot-instance (for reliability), and cloud based (for infinite parallelism). Training work tended to be chunky, never made sense to build servers in house that would be idle some of the time, and queued at other times.
And if you are solo dev its even easier choice as you can reuse your rig for other stuff when you dont train anything (for example gaming :D).
Only possibility is if you get free 100k from AWS and then 100k from GCP you can live with that for a year or even two if u stack both providers but it is special case and im not sure how easy it is to get 100k right now.
I’d rather be able to spin up 200 gpus in parallel when needed (yes, at a premium), but ramp to 0 when not. Data scientists waiting around are more expensive than GPUs. Replacing/maintaining servers is more work/money than you expect. And for us the training data was cloud native, so transfer/privacy/security is easier; nothing on prem, data scientists can design models without having access to raw data, etc.
Market for local small and efficient models running on device is pretty big maybe even biggest that exist right now [ios, android and macos are pretty easy to monetize with low cost models that are useful]. I can assure you of that and you can do it on even 4x RTX 3090 [ it wont be fast but you'll get there :) ]
Ah yes, my code can't be useful to people unless it takes a long time to compile...
What the above post was probably trying to get at is that the ML specific hardware is far more efficient these days than consumer GPUs.
There is tons of value to be had from smaller models. Even some state of the art results can be obtained on a relatively small set of commodity GPUs. Not everything is GPT-scale.
That's a great point. We'll be addressing this in an upcoming post as well.
We've served workloads that run entirely on spot GPUs where it makes sense since a small number of spot GPUs can make up for a large amount of spot CPU capacity. The best of all worlds is if you can manage both spot and on-demand instances (with a preference towards spot instances). Also, for latency sensitive workloads, running on spot instances or CPUs sometimes is not an option.
I could definitely see cases where it makes sense to run on spot CPUs though.
Probably one of HNs most common mistakes in comments
Times change though, we’re about to conduct the same analysis over again, with latest models better architected for accelerators.
Similarly, for small-scale convolutional CPU inference, where you only need to do maybe 20 ResNet-50 (batch size 1) per second per CPU (cloud CPUs cost $0.015 per hour) you can use inference engines designed for this purpose, e.g., https://NN-512.com
You can expect about 2x the performance of TensorFlow or PyTorch.
Throw some 3090s in a rack and you’ll break even in 3 months
I empathize a bit with the cloud providers as they have to upgrade their data centers every few years with new GPU instances and it's hard for them to anticipate demand.
But if you can easily use every trick in the book (CPU version of the model, autoscaling to zero, model compilation, keeping inference in your own VPC, using spot instances, etc.) then it's usually still worth it.
reading again - it seems this paper calls HPC with GPUs a slightly different name "GPGPU" and lists the research activity separately.. so I didn't see it as HPC; basically what I wrote is not accurate. got it
We're using GPU(some contains a TPU block inside) due to 'historical reasons'. With vector unit(x86 AVX, ARM SVE, RISC-V RVV) that is part of the host cpu, either put a TPU on a separate die of the chiplet, or just put it into a PCIe card will do the heavy lift ML job fine. It shall be much cheaper than the GPU model for ML nowadays, unless you are both a PC game player and a ML engineer.
WRT putting a TPU on a separate die -- this has been done for several years in the mobile space: Apple Neural Engine for iPhones, TPU (not same as server TPU) on Pixel, SNPE on Qualcomm, etc.
[0] https://cloud.google.com/compute/gpus-pricing
[1] https://cloud.google.com/tpu/pricing#v4-pricing
[2] this is somewhat unfair, because the GPU pricing number is for just the GPU and not the host it runs on, whereas the TPU pricing number (for TPU VMs) includes the host it runs on. If you include the price GCP charges for the host, preemptible A100s are about $1.20/hr. Why does Google make GPUs look cheaper than TPUs when they're not? Your guess is as good as mine.
With Hopper 100 on the way, I wonder when TPUv5 will come out.
I also wonder how Intel's Gaudi2 vs Ponte Vecchio will work together, looks like duplicate efforts for me.
AMD has its MI300 on the way, but it seems still far behind Nvidia|TPU|Intel at this point.
I'm sure every little thing I've discovered (e.g. measuring cpu/gpu workloads, trying to multiplex access to the gpu, etc) was probably covered in somebody's grad school notes 12 years ago, but I haven't found a source of info on the topic.
Doesn't look like it. Consumer:
AMD ThreadRipper 3970X: ~3000 USD on NewEgg
https://www.newegg.com/amd-ryzen-threadripper-2990wx/p/N82E1...
NVIDIA RTX 3080 Ti Founders' Edition: ~2000 USD
https://www.newegg.com/nvidia-900-1g133-2518-000/p/1FT-0004-...
For servers, a comparison is even more complicated and it wouldn't be fair to just give two numbers, but I still don't think GPUs are more expensive.
... besides, none of that may matter if yours is a power budget.
I am so confused how there seems to be a startup around having a work queue that does batching...
What is expensive? Those 3090ti's are looking very tasteful at current prices.
Game-playing (e.g. AlphaGo) is computationally hard but the rules are immutable, target functions (e.g., heuristics) don’t change much, and you can generate arbitrarily sized clean data sets (play more games). On these problems, ML-scaling approaches work very well. For business problems where the value of data decays rapidly, though, you probably don’t need the power of a deer or complex neural net with millions of parameters, and expensive specialty hardware probably isn’t worth it.
If it doesn't work it has to be retrained on new data again and there are no efficient alternatives to this energy waste other than use more GPUs, TPUs, etc emitting more CO2 after years of Deep Learning existing.
A complete waste of resources and energy. Therefore it is not worth it at all.
As humans we have our own adversarial examples, we get tired, we get sloppy, we might be even more biased than a calibrated model and always much more expensive.
It is entirely true and it just takes an invalid input to trick them and it messes up easily and even worse when there are always biases involved. Thus the value is nullified.
And once that model breaks and doesn't work, what is the solution? More retraining on new data? Even with that like I said there are ZERO efficient alternatives, which the cost outweighs the benefits.
Therefore, it is not even worth it.