edit: To clarify, I mean if you're running software that does something on a given set of physical machines and you cant get job throughput higher than some number, n, on a given machine it can be prohibitively expensive to scale machines ahead of time or during load. With Lambda we have the option to run a ton of independent processes on theoretically independent machines to vastly increase throughput while we slog through improving our design to improve per-machine throughput on our legacy EC2 infrastructure. Lambda allows you to isolate your scale problem to 1 job per container at a time instead of looking at machines as capable of only `n` jobs at a time.
But seeing as most of their work last less than an hour, I wonder how well they'd done to take advantage of the spot market.
Does anyone know of any tools that take advantage of the spot market? I know that Spark does [2].
[1] https://aws.amazon.com/lambda/pricing/ [2] http://spark.apache.org/docs/latest/ec2-scripts.html
If you're using Spot Fleets (a new feature) AWS will take a spec of the compute capacity you want and find it for you, across different instance types and availability zones, which makes it more likely you'll get the capacity you want.
In practice spot machines are pretty stable. We run hundreds across several different instance types and they're often up for weeks at a time.
So unless your batch job absolutely must not be delayed you might give spots a go. We used to run critical analytics EMR (Hadoop) jobs on spots and it saved us a a fortune - probably hundreds of thousands of dollars. Very occasionally we would switch them to on-demand (automatically) if we couldn't get spots at the price we wanted.
http://docs.aws.amazon.com/AWSEC2/latest/UserGuide/spot-requ...
I guess this is useful for longer-running batch jobs (longer than the 2 minute warning you get of termination on regular spots) of a relatively known duration.
https://aws.amazon.com/blogs/aws/new-ec2-spot-blocks-for-def...
For example, a fixed duration m3.medium is currently $.037/hr for 1 hour and $.047/hr at 6 hours. GCE's equivalent n1-standard-1 is $.05/hr all the time (and automatically triggers sustained use discounts). If you ask for a 4-hour block and only use 3.2 hours you pay the same 16 cents as you would have on GCE. As I tell my friends, if you want a cheap, guaranteed VM just use GCE on-demand; if you're locked into AWS though, check out Spot!
Disclaimer: I work on Compute Engine (and Preemptible VMs specifically), so I'm clearly biased.