Frugal Computing
muratbuffalo.blogspot.com
muratbuffalo.blogspot.com
I've always felt there were a lot of background tasks that just need to be done but have no time sensitivity that could fit this model, but in reality it hasn't ever been something you can generalize to a point that it's worth building it out. Still haven't given up on the idea though :).
It can really be great for experiments by a small startup or small, impoverished research group.
If you have customers expecting your results in a timely fashion that won’t work at all. But you can get them to pay for your computer time, so it’s no big deal.
For a small startup that hasn’t been funded yet it can be a lifesaver (let me set up this experiment and get some data — ok, now I can work on this other important thing until the response comes back, perhaps tomorrow). I used to work in the life sciences where experiments had long latency (spend all day preparing samples, stick them on the instrument and let it run all night, then spend a day or two on the data). It forced me to think differently about my strategy, which has been invaluable long term (I end up doing fewer, better experiments even when. It under such cost pressure). I suspect the folks who had to submit a card deck had the same benefits.
The context: providers need more instances than will sell at list price, so they sell access substantial discounts if you accept that a customer paying full fare can bump you out. (Note instance shutdowns might be correlated, not random.) I've definitely heard of folks running tools like Spark that way. From the software angle, being able to efficiently use various types of machine and being able to minimize work lost when a worker dies (i.e. to checkpoint) can help there.
I've heard of analogous setups for companies that own their boxes, e.g. Google's Borg paper talks about the tricks they do to make production and non-production jobs share the same box effectively.
I think sometimes cost optimization will lead you to counterintuitive answers. For example, obviously the post is right that memory is pricier than disk is. (Also, external memory algorithms are cool; more people should read about them.) But when box time itself isn't free, and more RAM or faster storage allows you to use less of it, the pricier computing resources can end up paying for themselves. Like a cost version of the race to idle in power management.
And the (good) discussion of the costs of going distributed may lead you to want to have big enough boxes available to allow you to distribute fewer jobs, even when those boxes cost (more-than-linearly) more than smaller ones.
I'm not saying this to advocate for a particular strategy; more want to say the right approach can depend in super tricky ways on the specific workloads and constraints you have.
Need more machines to handle unusual load? Just let the queue grow, if the thing are still bad next week, we'll order a few more nodes.
Machine went down? We can wait until tomorrow when sysadmins can troubleshoot.