Introducing Preemptible VMs
googlecloudplatform.blogspot.com
googlecloudplatform.blogspot.com
* you don't care too much about when a workload finishes
* it can be chunked into small units of interruptible work
* the cost of resuming interrupted work is zero or small
Examples of workloads that might look like this: * geocoding large batches of addresses
* map-reducing large event streams
* content scraping/crawling
Notice that you don't need to run all of your cluster as preemptible if you're using bdutil [0], too -- some of the workload can be preemptible, and some not. So you can guarantee a minimum processing throughput, but get extra throughput for very cheap.I think that's a great way to do it, though I wish it was generalized for all kinds of clusters, not just for Hadoop ones.
The guys on the go team are using it for their buildbots (http://build.golang.org) which was already prepared for the box to die.
In case anyone forgets, the pricing changes on Google App Engine caused many developers to abandon apps they had developed on the platform because of hundreds of percent increases in pricing:
https://groups.google.com/forum/#!topic/google-appengine/rSh...
During those 3 years no warning was given that the pricing changes were going to be of such a magnitude.
As a result people had sufficient time and not enough warning to build entire businesses on an infrastructure that they later had to abandon.
That's unforgivable.
Frankly, if your business is built entirely around a single 3rd party provider, and you are totally incapable of pivoting cloud providers at short notice... well then you are "doing it wrong".
Could you elaborate a little?
Netflix doesn't rely on a single provider for hardware, data centers, nor bandwidth. Sure they use EC2 for some things, but also have a great deal of hardware in data centers throughout the country, as well as custom hardware in ISP data centers, etc. I'd be shocked if there didn't use some of the other cloud provider offering too.
Netflix has no single point of failure.
Building your business around a single cloud provider creates a single point of failure.
App Engine has always been a unique PAAS because its APIs were designed to try to force developers into better distributed app architectures. The non-relational DB with entity groups and limited queries, the 30s request limit, originally not offering long-running instances, the task queue, memcache as a service, all made for scalable apps.
But the costs unfortunately weren't designed the same way. Charging by CPU time, and having datastore access free was especially bad for a market where apps typically use very little CPU, but access data a lot.
A lot of blog posts came shorty after the price change that said after making the recommended changes to datastore calls, and enabling multithreading that they got their bill down and performance up significantly. Many apps were doing absolutely no caching, or logging every request to the datastore, because datastore was ridiculously cheap for small records. Some were doing what should have been batch work per-request because they didn't use task queues.
Since 2011, I think all price changes with Google Cloud have been price drops, some pretty big. Last year, App Engine prices were dropped 30%.
The users didn't decide any of these things, and are not to blame for them. What kind of service blames its customers for its own mistakes?
Google earned this distrust fair and square by suddenly forcing huge numbers of customers to rewrite all their existing code. Little price cuts don't matter: other services are already more cost-effective and are cutting their prices all the time, but even if there were price parity the risk attached to the lock-in just isn't worth it.
https://plus.google.com/110401818717224273095/posts/AA3sBWG9...
As for the unique API, of course some of it's unique, because there weren't any standard APIs for the unique parts. Datastore is based on Google's internal NoSQL store. There's are/were standard APIs for NoSQL. Same with Task Queues.
The "proprietary API" bit is overblown though, IMO. The Python version shipped with a somewhat tweaked, Django support, and has grown to support WSGI and standard Django. The Java version uses Servlets with some constraints. Memcache is pretty standard. You can pretty easily abstract away the proprietary bits to run on a standard stack, or use an open source implementation of the APIs.
http://googleappengine.blogspot.com/2011/10/app-engine-155-s...
The price increase also came with an SLA. Just to be clear, you're saying that businesses were entirely built atop a product with no SLA, and that's not the bigger problem?
I run a $MM enterprise business almost entirely on GAE/python with a staff of ~40, public-facing site, etc and the monthly bill is under $1500/mon. Sure, I'd prefer lower $ and faster performance but truthfully, I'm no longer complaining: GAE/python saves me $$$ in IT staffing costs including security upgrades on dozens of packages that are either pre-integrated or I don't need at all (SQL & NoSQL databases incl multi-DC failover, memcache, reverse proxy, email hosting, auto-scaling, etc. etc.)
Sure, when the change from billing CPU time to instance hours came in some app's bills sky-rocketed. But that was because they were poorly coded such that instances were blocking and unable to serve incoming requests.
With a thread safe application and the proper configuration there is absolutely no reason why instance-hours pricing shouldn't be competitive.
Really? The free tier comes with 28 instance hours per day. That'd mean your app would have to serve hundreds of requests per second, meaning each request must take substantially less than 10 ms, on a 600 MHz, 128 MB RAM machine.
If your request do any work at all, I doubt you can handle them in <10ms on a 600 MHz CPU.
Correction, each request would need to have less than 10ms of CPU time - the instances support concurrency.
My web frontend, by design, does very little - any CPU heavy operations are done by other systems using the task queue. Writing it in golang has helped as well, wouldn't get that performance from python.
Write a simple Hello world example in golang and get it to do some mathematical calculations to simulate "work", I think you'll be surprised at how many requests a second you can squeeze out of a single instance.
But still, 10 ms really is not a lot on a 600 MHz machine. How long does your front end take to serve one request? How many qps do you serve from a single instance?
I have some go code with a trivial, completely unoptimized blog, rendering a couple of articles. Poking at appstats suggests I spend a bit more than 10 ms CPU time, App Engine reports ~30 ms CPU time.
That particular request is authenticated, writes a file to google storage, and then returns a response (with a few other things like logging etc.)
So as well as potential future price hikes, you should be prepared for the possibility that this service might not be around for the long haul. Therefore you should skip using any Google-specific functionality, and instead implement a design that allows you to easily migrate to another VM vendor.
I could easily beat those 4 with a multitude of examples from both Apple and Microsoft, that doesn't mean that any of them are untrustworthy, just that they evolve and continue to grow.
At least when google shuts down a service, they give a good amount of "heads up" to those using it, provide examples of trustworthy equivalent services from competitors, and ALWAYS provide an export function if it makes sense to have one.
Because their early marketing led people to not expect that of google, whereas with apple and microsoft it's just business as usual.
No, that's pretty much what it means. To varying degrees when you build on top of Apple, Microsoft or Google offerings, you're sharecropping rather than farming. Sometimes, they'll just find that it's in their interest to no longer lease out the farm to you -- or to change the rates/terms.
"Untrustworthy" might be an over strong term, if for no other reason than that invoking "trust" as a relevant concept in this context is probably itself incorrect.
The risk of writing apps on a platform like App Engine is far greater because people tend to spend years building software and a business tightly interwoven with the platform, only to be shut down by unpredictably massive rate hikes because you basically have to rewrite to get off the platform, not just the app itself but all your ops stuff too.
While it is possible to dig yourself into that hole using lots of Amazon services, Amazon doesn't require you to use platform-specific APIs, many of their services like EBS are very easy to replace, and Amazon's services have not been subjected to the same project-ending rate hikes. In practice, people move on and off EC2 all the time. So if you are going to trust a vendor, it's more reasonable to trust Amazon.
App Engine has a very idiosyncratic API; GCE is more or less a direct clone of EC2.
Or at the very least: never rely on one vendor's proprietary implementation(s). Always have at least two in active use (though you can do something like using one vendor 95% of the time), and if/when one goes down make it your top priority to find another.
When using the calculator this would make the price increase seem enormous to many customers. When concurrent request support was added (it was available to trusted testers at this time) all a user would have to do was add "threadsafe:true" to the app.yaml file to enable it (assuming their code wasn't doing anything silly).
Also when the bugs occured, I realised there was no way to get quick support.
Thankfully we moved our main app away to ec2 around when they increased the pricing a couple of years ago.
gcloud compute instances create my-vm --preemptible --zone=us-central1-c
or just set the preemptible bool to true via the API (it's under scheduling). When you want to test how your system behaves on preemption, you can just do: gcloud compute instances stop my-vm --zone=us-central1-c
which will give you the same 30 second timeout as when you're preempted. Most OSes have a fairly standard set of things they do on shutdown that will at least send all your running processes a signal (via kill), but if you need to add your own you can inject it via the new shutdown script support (https://cloud.google.com/compute/docs/shutdownscript).We tried to cover this in the docs (https://cloud.google.com/compute/docs/instances/preemptible) any feedback on that would be welcome!
When I worked at Google in 2013 part of my job was running very large calculations that were often pre-empted. I usually ran at the lowest priority and was in effect using spare capacity. No hassles, assuming that getting runs completed was not too time sensitive.
Why would someone go for this when cheap dedicated host providers like hetzner etc offer powerful dedicated servers with 64GB memory and multicore server grade CPUs? The comparison only gets worse taking into account that Google's offering is preemptible and can shutdown and come up as they wish.
[1]: https://cloud.google.com/compute/#pricing (0.12x24x30)
But time and time again I see infrastructure where people pay for these services for large amount of instances that are running continuously, blindly assuming that it's cheap because it's cloud. There's a bizarre level of price-blindness amongst certain subset of customers of Google Cloud and AWS that I've never seen anywhere else.
If you're going to use the cloud you have to do it right, and that means auto-scaling and variable resource usage. Then this option will save you money.
Similarly, if you're running hosted app that on most hours doesn't overload a single server, on peak hour fills three hosts, but on a large advertising event or accidental viral link takes fifty hosts for a day, and then goes back to normal - then you don't want to run it on VMs where you have to pay for them by month.
Nothing stops you from mixing and matching dedicated servers with handling batch jobs and peaks with cloud servers. In fact, most data centre providers can offer the full range from unfurnished colo space to cloud offerings out of the same data centre these days - either directly or via partners hosted in their buildings. At least that's my experience.
[Edit: Point taken though]
I use Google Compute for some personal project and the typical run time for a VM is about 10 - 20 minutes. If the vm has 90% chance to survive the first hour, it could be worth the trouble to make my process more fault tolerable.
The probability that Compute Engine will terminate a preemptible instance for a system event is generally low, but may vary from day to day and from zone to zone depending on current conditions.
Give it a shot in us-central1-a and let us know how it goes!
[Edit to undo my quote text (it wrapped poorly)]
gcloud compute instances start instance-name
Note of course that if you got preempted because we needed the capacity back for regular VMs (as opposed to say a maintenance event) you may not be able to start a new Preemptible VM in that zone.[1] http://storage.googleapis.com/pricingcomparator.appspot.com/...
[2] http://storage.googleapis.com/pricingcomparator.appspot.com/...
disclaimer: I work at Google on Google Cloud Platform
California tends to be more expensive than other AWS regions, though (I can't remember the reason - perhaps just availability?). If you're in us-east-1 the price sticks around $0.20 per hour, and eu-west-1 is rock-solid at $0.32 per hour.
If you have work that can be done on spots it can be a good idea to make it region-agnostic so that you can take advantage of better prices for different instances in different regions.
For example, with Google you don't need to get a specialized instance type to use it's monster-fast Local SSD - just use the same instance types. This alone should simplify use of Preemptible VMs.
I say "market" because no-one really knows how the spot market place actually works. We've had machines run for weeks, and other times the prices fluctuate in bizarre ways and we can't get out preferred instance types for hours or (in the worst case) days. There's an interesting analysis here: http://santtu.iki.fi/2014/03/20/ec2-spot-market/
I looked through the pricing table and played with the calculator, it seems something equivalent to our needs would cost around a third more on google but each cpu would have twice as much ram. Not worth it for us.
Note: we're very aware of this pain point, and maybe you'll see something soon ;).
But to be honest about the situation, the cost would have to be much lower to make it worthwhile for me to rewrite our scheduler on Google's API.
EC2's spot market fluctuates based on supply and demand, and there's no reason to think the same forces won't apply to GCE.
Preemptible VMs will prove to be very useful in fault tolerant Distributed Networks.
One more usecase that I can think off is Load Testing on a large scale ( in a distributed way ).
Tough to build a business model around a resource you can't even determine the availability of.
That's exactly why it's cheap: you're trading price for reliability.
Say you have a few hundred terabytes of images to process. You can prioritize images by pushing them to the head of the queue, but you don't really care how long the complete batch takes.
If you are happy to wait for your job to complete you pay less. Otherwise, pay more and guarantee completion.
You're not buying a dedicated instance for a month, you're renting it as long as you need it. If you only need a core for an hour, your bill will be $0.01.