Technically, you're gambling a bit because whenever a spot instance is terminated, you probably need to spend a bit extra to redo whatever work it was in the middle of. But spot instances have such a huge discount that it's almost always worth taking that risk, as long as you don't mind the extra management complexity.
I used to work for a company (since shut down) which provided big data processing systems in the cloud. This was one of the typical use-cases for our customers. Big data systems like Hadoop, Spark, etc are built to handle this kind of disruption where you lose a few nodes once in a while, and we had built in further optimizations to do it even better. This fact, combined with the much cheaper price of Spot Instances (upto 90% less than on-demand) make them a compelling alternative to on-demand instances.
In practice - at least on AWS - spot loss used to be quite rare. When it happened, it happened by the truckload, but we used to have spot instances run for several days without termination.
Similarly, Kubernetes works well with spot VMs. You can have Spot and Regular node pools in one cluster, and then your workloads can happily spread out to the extra capacity at a huge discount.
As for external storage: yes, you do generate more metadata, but that's often a tiny fraction of the cost saved by moving to spot machines.
In general, these tools are adopted more for increasing reliability of existing systems, but I predict they would be a neat fit to run them on spot machines.
I've also got another script that keeps track of how many spot instances are currently running and spins up some new ones (possibly in a different region) if they fall below a certain level.
https://aws.amazon.com/blogs/compute/new-amazon-ec2-spot-pri...
You always need some traditional instances to serve things that need to be available at all time, like your website. Now you could run things using spot instances and switch to traditional during peak hours but that requires confidence that enough capacity is going to be available when you need it (which is going to be at the same time as everyone else, so it's always a bet :))
"However, your cloud provider can pull these virtual machines out from under you at any time without notice (because someone else is willing to pay more). So how do you lower the probability of losing all of your instances to nearly zero? Cloud providers generally need to have a large buffer of capacity and they have many different virtual machine configurations with different CPU and memory sizes. So, don’t just spin up virtual machines of a single configuration type, spin up lots of them from many different configuration types!"
https://blog.comma.ai/scaling-for-10x-user-growth/
I'm not sure if they still have a fallback of traditional instances.
Or, more likely, for batch processing that's mostly time insensitive. Things that can be resumed without losing much work or lots of small short jobs. Retranscoding media with new settings comes to mind. If you've got deadlines, you'd probably need a mix of on-demand and spot nodes.
This is the corner of cloud pricing where if you can optimize use of this, you might be able to do better than traditional hosting. But only if your needs are variable, or you get a lot of high discount spot pricing.
cost savings.