Show HN: Deploy ML Models on a Budget
github.com
github.com
Disclaimer, not affiliated just enjoy the entire experience of using cortex. We are moving to KubeFlow because we need scale and more features but cortex is great for what it does.
However, this works without a Kubernetes cluster, while Cortex does not. A managed kubernetes cluster on GCP e.g. costs $100/month[1].AWS, Azure is similar. For bare-metal, self managed solutions, if you can manage a cluster on your own like that then yes there is no need for BudgetML :-)
As the pricing link notes, it's worth pointing out your first cluster is free (aside from VM costs), and people interested in this type of project probably don't have any other clusters.
https://lucidbeaming.com/blog/fine-tuning-a-gpt-2-language-m...
The extra monthly $40 for a novelty project has me searching for cheaper alternatives. Google preemptibles scared me off because they can scale dynamically and generate runaway costs if my project suddenly got popular on Reddit or something.
You can run preemptible instances without any kind of management or with only very simple management and you should be okay. I recommend a simple managed group with a policy of, "I want maximum one of these instances running at all times and, if it gets preempted, then please start a new one."
If you're using Google Cloud, don't deny yourself the cost savings of preemptible instances. When I was using the same infrastructure, preemptions were very rare!
- Setup time: Setting up GCP, setting up a certificate, adding a static IP, etc. is not seamless/adds friction
- Autoscaling and rolling updates (no downtime)
- Team management and collaborative environment with usage tracking, permissions, etc.
- Optional integration with a pipelining service for training, tuning, deploying models in a single tool
And a point of clarification: Practically speaking, neither tool is free. Both require a cloud instance so they will cost roughly the same for the end user (Gradient also supports preemptible instances).
We made BudgetML not to replace the above but to make it easy for scenarios where the need would be to just "get out there and deploy". Do you see value in that too?
It’s often a misconception that you need GPU for inference. It many cases, the overhead of data transfer to GPU makes it much slower than a well tuned CPU.
For example, here is ResNet50 for a particular Skylake-X CPU:
(It depends on the CPU, how many cores, and the structure of the neural net, but MKL will generally only achieve between 80% and 90% of NN-512's AVX-512 performance)
In fact we do offer GPUs from multiple ways
- Dedicated servers (baremetal) : hidden from the website today, we are refreshing the hardware parts. New ones with GPUs are planned in few weeks. But you will need to buy servers with GPUs for at least 1 one. Good but no so flexible. Not good for AI Training budget imho espect if it's 24/7 trainings.
- Public cloud / VMs : like AWS/GCP/AZure we provide VMs with NVIDIA 1 to 4 x Tesla V100 16GB The good thing is flexibility (pay go hourly) The "bad thing" is about perf when dealing with large dataset, and it also require sysadmin tasks (ssh, dist upgrade, you know it). You'll also need tools such as kubeflow to allow a team to work with pipelines, orchestration and debugging tools
- (NEW) Public Cloud AI Training : it's basically GPU as a service. like paperspace for example. You start a "job" via UI/CLI/API with 1 to 4 GPUs NVIDIA Tesla V100s, plugged to provided images (notebooks + tensorflow, notebooks / pytorch, fastAI, HuggingFace, ...), or you own images (hello docker public/private registries).
The good thing is : you have full flexibility (you can launch as many jobs as you want), pay per minute, start in 15 seconds. No SSH required, no drivers to install, ... And one last good thing is the hidden part : specific "cache" storage near the GPUs to play with large datasets (several TBs) without bottlenecks such as latency. It's as good a local NVMe storage the bad thing is : nothing :)
I'm not here for hidden advertisement but fuck yeah we have something to propose for GPU :p To find the links or price, everything is public in our website, for a free Voucher just DM me ! Have a good day.
If you're still using EC2 you should just switch to us, your monthly bill will thank you.
As far as I can tell, it cannot deploy models that use high dimensional input. It can log these models, and it can deploy them, but the "deployed" model can only accept two dimensional data.
This is an open issue[0] with open pull requests. We're experimenting with MLflow for our platform (https://iko.ai), but we automatically detect models, metrics, and parameters and then track that with MLflow instead of haivng people pollute their notebook code with MLflow code. We don't like leaking that detail (tracking code) into the notebook.
We circumvented these issues (https://iko.ai/docs/appbook/#deploying-a-model). The models can either be deployed, or you can build a Docker image and push it to a registry and do whatever you want with it (run a container and invoke the model's endpoint, for example).
This just works, and costs my team less for a year than one week of Sagemaker or kubernetes architectures.
https://www.corbettanalytics.com/deploy-machine-learning-mod...
Our on-demand instances are 50% the cost of GCP/AWS. Our 8x V100 instances are $12/hr and our 1x Quadro RTX 6000 instances are $1.25/hr. They come set up with TensorFlow/PyTorch/Jupyter notebook so you don't need to do much DevOps work to get started.