Spot Instances are 70% cheaper than On-Demand instances but are prone to interruptions. We mitigate the downside of these interruptions through the use of persistence features, including optional fallback to On-Demand instances. So you can optimize workflows according to your budget and time constraints.
History: We were working on a neural rendering startup that needed a lot of GAN training which was getting very expensive. We were blowing roughly $1000, to train a single category class. Training on Spot instances was cheaper, but still a mess. It needed lot of hand holding/devops stuff to make it usable. So we built SpotML to automate a lot of things.
Posting it here to see if the community finds this helpful, so that we can open it up to the larger community.