Show HN: SpotML – Managed ML Training on Cheap AWS/GCP Spot Instances
spotml.io
spotml.io
Congrats! Can I suggest charging more?
IIUC, your business plan is $10/month for the company, regardless of number of users?
You probably save a company that $10 in a day or less for one GPU: One A100 is ~$3/hr on demand and ~$.90/hr as Preemptible, saving over $2/hr.
Said another way, your pitch is to recover a lot of the 70% discount that they aren't going to do themselves. If you were a managed training service, you could pitch yourself as "half the price of AWS or GCP" and keep the 20%+ margin with both parties being happy. (The problem is that pass through billing makes that obvious, you need to support lots of bucket security and IAM controls, etc.).
Fwiw, I would also branch out into inference! Preemptible and Spot T4s are commonly used for heavy image models, but many people pay full price. Inference that takes X ms can easily be handled "without errors" in the shutdown time. The risk is handling all the capacity swings.
Also, interesting point about inference. I'm not sure though how common it is for companies to need GPUs for inference. Because if you can have a CPU based inference model, which I thought was most common, it's probably not a big usecase?
The need to reach for a T4 comes when someone is doing a big model on images or video and wants sub-second response time. (Think some of the stuff on Snapchat, etc.)
Spot Instances are 70% cheaper than On-Demand instances but are prone to interruptions. We mitigate the downside of these interruptions through the use of persistence features, including optional fallback to On-Demand instances. So you can optimize workflows according to your budget and time constraints.
History: We were working on a neural rendering startup that needed a lot of GAN training which was getting very expensive. We were blowing roughly $1000, to train a single category class. Training on Spot instances was cheaper, but still a mess. It needed lot of hand holding/devops stuff to make it usable. So we built SpotML to automate a lot of things.
Posting it here to see if the community finds this helpful, so that we can open it up to the larger community.
Congrats on the idea and godspeed. You’ll probably have a lot of interest if you execute well.
No one asked, but it's HN so I'll say it anyway. I think it's a questionable "vc business" but a great business for 1-2 people. The road from this to an enterprise sales motion, or even a 10K/year contract is hard for me to imagine. At some point, it becomes cost effective for my org to build this functionality in house.
However, as a hobbyist /single dev / small team, $120/Year is a no brainer after the first 2-3K I spend on GPU by mistake. As you know, setting up spots when I just want to get shit done is a pain and I'll gladly pay you (a little) to make that go away.
Still no one asked but... One thing that plays to your advantage is that the current price point is something anyone in your user group can buy on their own, and there are a lot of us / enough to make a nice business out of.
Good luck!
Apart from the fact that it could deploy to both GCP and AWS, what does it do differently than AWS Batch [0]?
When we had a similar problem, we ran jobs on spots with AWS Batch and it worked nicely enough.
Some suggestions (for a later date):
1a. Add built-in support for Ray [1] (you'd essentially be then competing with Anyscale, which is a VC funded startup, just to contrast it with another comment on this thread) and dbt [2].
1b. Or: Support deploying coin miners (might help widen the product's reach; and stand it up against the likes of consensys).
3. Get in front of the very many cost optimisation consultants out there, like the Duckbill Group.
If I may, where are you building this product from? And how many are on the team? Thanks.
[0] https://aws.amazon.com/batch/use-cases/
[1] https://ray.io/
Also i'm not sure how straightforward it is to detach/attach persistent volumes to retain data across different spot interruptions ? The latter can be done but it's just the same rote each time you wanna train something new.
Also thanks for the suggestions ! We're a team of 2 right now, I used to be in the bay area but in Mexico temporarily.
2. Checkpointing was a pain (we relied mostly on AWS Batch's JobState and S3, not ideal), but the current capability to mount EFS (Elastic Filesystem) looks like it would solve this?
3. No hot swapping on-demand with spot and vice versa. Interestingly, ALB (Application Load Balancer) supports such mixed EC2 configurations (AWS Batch doesn't).
So useful (and potentially lucrative) that AWS/GCP would likely take over this functionality eventually -- either by building on top of spot instances like you have, or underneath regular instances or re-use that capacity for some other managed service (thereby reducing the price differential). How do you plan to protect yourselves against that?
Nimbo docs: "In order to run this job on Nimbo, all you need is one tiny config file and a Conda environment file (to set the remote environment), and Nimbo does the following for you:"
SpotML docs: "In order to run this job on SpotML, all you need is one tiny config file and a Docker file (to set the remote environment), and SpotML does the following for you:"
Make of that what you will :).
Otherwise, this idea is interesting and probably generalizable to other applications. Maybe it's not crystal clear to me, but what are the advantages of your service over existing solutions such as Nimbo and Spotty? FWIW it might be worthwhile adding this to your website.
Good luck!
The biggest advantage which was missing in the Open source options was monitoring on the training job and auto recovery from spot interruptions which spotML does.
Also there's at least two open source free solutions for elastic training I know of. RaySGD, https://docs.ray.io/en/master/raysgd/raysgd.html and elastic horovod, https://horovod.readthedocs.io/en/stable/elastic_include.htm... For me to consider this I'd need a comparison table with native framework solutions and these solutions along with what does it add for me.
Why not just train on Spot Instances with a retry implemented?
I see that SpotML has a configurable fall back to On-Demand instances, and perhaps their value prop is that it saves the state of your run up to the interruption + resumes it on the On-Demand instance, but why not just set a retry on the Spot Instance if its interrupted?
I'm failing to see what is different about SpotML vs Metaflow's @retry decorator and using AWS Batch: https://docs.metaflow.org/metaflow/failures#retrying-tasks-w...
If you're in the comment still, Vishnu, would love to hear your thoughts
I've read through the docs, the one difference that comes to my mind is the automatic fallback to on-Demand and resume back to spot when available. I can't readily see a way to do this yet in Metaflow, but it's possible I've missed something.
Second, you can rely on Spot Fleets which handle both spot and on-demand instances seamlessly https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/spot-fle...
I used to have these same pains. My trick with spot instances has been to set my maximum price to the price of a regular instance of that class or higher and to sync weights to S3 in the background on every save. The former is a parameter when starting the instance in the console or terraform, the latter is basically a do...while loop. I've noticed that often one gets booted from a spot instance causing interruption because the market price increases a few cents above the "70% savings price." Increasing the maximum to on demand is basically free money because you don't get booted often, and your max price is the regular on demand price.
This has seemed to mitigate about all of the spot downsides (like interruptions or losing data) because you don't easily get kicked out of the spot unless there's a run on machines and the prices rarely fluctuate that much (at least for the higher end p3 instances). This has seemed to prevent data loss and protected the downside risk by setting the instance to a knowable max price. There are instances where the spot price goes higher than the on demand price, so you still get booted once in a while, but it's very infrequent as you can imagine most people don't want to spend more for spot than on demand.
Anecdotally I still average out to getting the vast majority of the spot savings with this method with very few interruptions. Looking at the SpotML it seems to be a lot of tooling that achieves these same goals (assuming one would be interrupted when a spot dies and moves to full-freight on-demand with SpotML), which makes SpotML's solution feel very over-engineered to me given that the majority of what SpotML provides can be had with a simple maximum cost parameter change when spinning up a spot instance.
I would be very interested in using anything that doesn't have great overhead and saves money. Our bill seems "big" to me (but I realize it may be small to many others), so even these small savings add up. Would you compare the potential benefits of SpotML to the method I described above?
why? sure, you don't want to pay more than the on demand price, but iirc spot prices often spike very momentarily. so the question becomes whether the sum over spike time cost at the higher spot price exceeds the time to boot/shutdown/migration overhead at the on demand price.
but i've never actually tried it so...
i also suppose this way of thinking about it comes from thinking around how to minimize cost from a purely mathematical standpoint. when you think about how and why the spot market is operated, and what those short term spikes may actually be, it may run counter to the intended purpose. (cheap capacity that they may recall at any time, because capacity is actually fixed)
funny how market based approaches can gamify things sufficiently that sometimes they obscure the underlying intention or purpose of having a market in the first place.
We spend about $10k/month on spot instances and I don't specify any max price. The way to avoid terminations is just to make sure you spread your workload over a large number of instance types and availability zones.
I hadn't heard of nimbo, maybe I can read how they're doing it since it's open sources. Does anyone have any idea how they're saving state so fast (NVME SSD disk?)
After looking around I thinking more about CRIU/docker suspend. The google stars aligned and I found this https://github.com/checkpoint-restore/criu-image-streamer + https://linuxplumbersconf.org/event/7/contributions/641/atta... which actually seems perfect. I wonder how fast it is
Edit: Also no GPU support AFAIK but https://github.com/twosigma/fastfreeze looks really nice, turnkey. I wonder if I write to a fast persistent disk if I can get higher maximum ram than over the NW
(or, hacking on a checkpoint idea, have a daemon periodically 'checkpoint' other programs so even if it's too slow over 60 seconds, revert to the last checkpoint. Even an rsync like application where only send the changes)
Makes sense, just save checkpoints to disk. What I'm doing is more CPU bound and not straight ML so less easily check-pointed, sadly. Cool though, it's worth jumping through hoops for 70% reduction
Source: https://cloudoptimizer.io
Seems like the services I've tried, have focused heavily on supporting notebooks (colab, paperspace).
Any ideas?
Some offer both notebook and ssh, some just one of the two. The cheapest are often the p2p ones, where you essentially rent someone else's consumer gpu
Could you give an approximate cost of fine-tuning a model like Bert or even GAN given this system?
Just want to get a sense of cost of using the system.
For detecting if the training process is still running or errored out it registers the training command pid when launching the task and then monitors for the Pid for completion. It also registers and monitors the instance state itself to check for interruptions and resuming.
All you need to do is schedule your jobs (just call the Kubernetes API and schedule your container to run). In the rare case the node gets preempted, your job will be restarted and restored by Kubernetes. Let your node pool scale to near zero when not in use and get billed by the _second_ of compute used.
However, there's at least a couple of things that matter here that aren't covered by "just use a preemptible node pool":
* SpotML configures checkpoints (yes this is easy, but next point)
* SpotML sends those checkpoints to a persistent volume (by default in GKE, you would not use a cluster-wide persistent volume claim, and instead only have a local ephemeral one, losing your checkpoint)
* SpotML seems to have logic around "retry on preemptible, and then switch to on-demand if needed" (you could do this on GKE by just having two pools, but it won't be as "directed")
This is a hustle to gauge interest (and collect emails) in a service that is a clone of nimbo.
unable to figure out how to construct the job itself and how to submit the dockerfile to be executed.
also - do u support distributed training ?
Have you used it btw ? and what has your experience been with Grid ?