In fact we do offer GPUs from multiple ways
- Dedicated servers (baremetal) : hidden from the website today, we are refreshing the hardware parts. New ones with GPUs are planned in few weeks. But you will need to buy servers with GPUs for at least 1 one. Good but no so flexible. Not good for AI Training budget imho espect if it's 24/7 trainings.
- Public cloud / VMs : like AWS/GCP/AZure we provide VMs with NVIDIA 1 to 4 x Tesla V100 16GB The good thing is flexibility (pay go hourly) The "bad thing" is about perf when dealing with large dataset, and it also require sysadmin tasks (ssh, dist upgrade, you know it). You'll also need tools such as kubeflow to allow a team to work with pipelines, orchestration and debugging tools
- (NEW) Public Cloud AI Training : it's basically GPU as a service. like paperspace for example. You start a "job" via UI/CLI/API with 1 to 4 GPUs NVIDIA Tesla V100s, plugged to provided images (notebooks + tensorflow, notebooks / pytorch, fastAI, HuggingFace, ...), or you own images (hello docker public/private registries).
The good thing is : you have full flexibility (you can launch as many jobs as you want), pay per minute, start in 15 seconds. No SSH required, no drivers to install, ... And one last good thing is the hidden part : specific "cache" storage near the GPUs to play with large datasets (several TBs) without bottlenecks such as latency. It's as good a local NVMe storage the bad thing is : nothing :)
I'm not here for hidden advertisement but fuck yeah we have something to propose for GPU :p To find the links or price, everything is public in our website, for a free Voucher just DM me ! Have a good day.
It’s often a misconception that you need GPU for inference. It many cases, the overhead of data transfer to GPU makes it much slower than a well tuned CPU.
For example, here is ResNet50 for a particular Skylake-X CPU:
(It depends on the CPU, how many cores, and the structure of the neural net, but MKL will generally only achieve between 80% and 90% of NN-512's AVX-512 performance)
If you're still using EC2 you should just switch to us, your monthly bill will thank you.
We made BudgetML not to replace the above but to make it easy for scenarios where the need would be to just "get out there and deploy". Do you see value in that too?