GCP releases Spot VMs, the next generation of Pre-emptible VMs
cloud.google.com
cloud.google.com
AWS has had this kind of spot instance for years, but with a 2 minute grace period rather than the 30 seconds GCP is offering. Azure and GCP both originally went with the 24-hour cutoff (which can easily be replicated on a regular spot instance if needed), but now GCP are backing off on that restriction.
I used LP/spot both as scaling sets and the more recent single VMs
Also i wouldn't call the cutoff and grace period as "paths". There were much more substantial differences between the different clouds
(I work for Azure but don't know anything about the policies)
For services though, GCP pre-emptible instances are perfect combo for kubernetes.
This would be easy to work around, but nonetheless could lead to unexpectedly high charges if you were relying on this behavior only to have it silently change.
I edit application code on my (relatively underpowered) laptop, and automatically mirror it to the instance where the running services picks up any changes, and recompiles and relaunches as needed. It's a fairly chunky app code-wise so moving the CPU usage off of my laptop is very helpful.
When and if the instance shuts off early, I just relaunch it and reconnect. This amounts to one click of the mouse in the GCP UI, one locally run terminal command to connect, and one remote terminal command to (re)start my app. No work lost, and the cost savings are worth the minor inconvenience. Usually the instances last the full 24 hours anyway, and I usually shut it down when I'm done working, so interruptions are very rare.
I can understand that in the context of a larger company with more resources, a dev would be put off by this. But it works very well for my uses.
You can approach this by continuously deploying upgrades to the dev fleet, but it's simpler to simply set a lifetime bound (with an opt-out for special circumstances).
It did guard you from keeping things on by mistake but there are way more downsides to this kind of restrictions. Especially that you can simulate the daily shutoff like other comments here said
Instead they created a whole new API for basically the same feature.
Functionally, I see no difference.
So preëmptible VMs are still supported but may now cost as much as 40% of non-preëmptible VMs, "to reflect the supply-demand dynamics of Compute Engine’s excess capacity." That's quite a big change from the current pricing structure.
The doc says: > Preemptible VM instances are available at much lower price—a 60-91% discount—compared to the price of standard VMs.
Maybe you're using huge disks, those don't get discounts on preemptible instances from what I remember.
Technically, you're gambling a bit because whenever a spot instance is terminated, you probably need to spend a bit extra to redo whatever work it was in the middle of. But spot instances have such a huge discount that it's almost always worth taking that risk, as long as you don't mind the extra management complexity.
I used to work for a company (since shut down) which provided big data processing systems in the cloud. This was one of the typical use-cases for our customers. Big data systems like Hadoop, Spark, etc are built to handle this kind of disruption where you lose a few nodes once in a while, and we had built in further optimizations to do it even better. This fact, combined with the much cheaper price of Spot Instances (upto 90% less than on-demand) make them a compelling alternative to on-demand instances.
In practice - at least on AWS - spot loss used to be quite rare. When it happened, it happened by the truckload, but we used to have spot instances run for several days without termination.
Similarly, Kubernetes works well with spot VMs. You can have Spot and Regular node pools in one cluster, and then your workloads can happily spread out to the extra capacity at a huge discount.
As for external storage: yes, you do generate more metadata, but that's often a tiny fraction of the cost saved by moving to spot machines.
In general, these tools are adopted more for increasing reliability of existing systems, but I predict they would be a neat fit to run them on spot machines.
I've also got another script that keeps track of how many spot instances are currently running and spins up some new ones (possibly in a different region) if they fall below a certain level.
You always need some traditional instances to serve things that need to be available at all time, like your website. Now you could run things using spot instances and switch to traditional during peak hours but that requires confidence that enough capacity is going to be available when you need it (which is going to be at the same time as everyone else, so it's always a bet :))
"However, your cloud provider can pull these virtual machines out from under you at any time without notice (because someone else is willing to pay more). So how do you lower the probability of losing all of your instances to nearly zero? Cloud providers generally need to have a large buffer of capacity and they have many different virtual machine configurations with different CPU and memory sizes. So, don’t just spin up virtual machines of a single configuration type, spin up lots of them from many different configuration types!"
https://blog.comma.ai/scaling-for-10x-user-growth/
I'm not sure if they still have a fallback of traditional instances.
https://aws.amazon.com/blogs/compute/new-amazon-ec2-spot-pri...
Or, more likely, for batch processing that's mostly time insensitive. Things that can be resumed without losing much work or lots of small short jobs. Retranscoding media with new settings comes to mind. If you've got deadlines, you'd probably need a mix of on-demand and spot nodes.
This is the corner of cloud pricing where if you can optimize use of this, you might be able to do better than traditional hosting. But only if your needs are variable, or you get a lot of high discount spot pricing.
cost savings.
Say I am running a long batch process on one of these spot VMs and I get pre-empted ...
Does my job restart where it was stopped automatically and everything is transparent save how long the job takes to complete, or do I have do to checkpointing myself and deal with the fact that my jobs may be killed at anytime?
Also, if the restarts are indeed transparent to my job, what of stateful network connections?
Finally, what kind of guarantee do I have that my job will ever complete if pre-emption can happen at an arbitrary frequency for arbitrarily long?
So you'll need to adapt your process to be resumable (or partially resumable through checkpoints) and/or idempotent, so nothing goes wrong if you run the job (or parts of the job) twice.
Basically this is good for running short jobs and not long running ones. If you have a service that process chunks of data from a queue it's probably the ideal scenario.
In practice, with AWS your instances don't get killed that often. Some in the comments are claiming months but I was using mine with AWS Batch and they usually lived 1-2 weeks.
(google could run negative forever: the point is, some VP won't want to, and the KPI will morph to "kill it")
Somebody once told me a policy of random offswitch helps
This is just aligning the offering with Azure and AWS
Link to pricing page: https://cloud.google.com/skus/?currency=USD&filter=C054-7F72...
That means if you just need one VM, you won't pay for IP charges.
Obviously that might change anytime without notice.
A per minute granularity can make Google's offering more enticing for a lot of users.
https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/spot-int...
On GCP seems like you pay per seconds, but if you use a Premium OS and GCP stops your instance, you would pay for the premium OS pricing anyway:
https://cloud.google.com/compute/docs/instances/spot#spot-wi...
Don't have an authoritative answer, sorry, but granularity is an area where Google has historically been better than the competition.
I believe AWS has had spot instances for a very long time (more than 5 years at least)
Is my understanding correct? and if so any insights on why it took so long
https://developer.apple.com/documentation/contacts/cnlabelco...
dfs.datanode.available-space-volume-choosing-policy.balanced-space-preference-fraction dfs.datanode.asvcp.bspfI think this gives a hint though - when something is very-frequently-used, people need a short name so that it doesn't stand in the way of the discussion; and due to frequency, the fact that it's an acronym doesn't matter: I bet there are more people know what ASCII is than there are people who know what the acronym stands for.
For configuration, though? asvcp.bspf is never going to be frequently used by anybody. That's generally the point of configuration, for the vast majority of cases, you touch it at most once. So I argue you want long names.
Same thing can be said about build scripts. Those who strive for terseness (hey sbt) become unreadable mess 1 year after writing, when you need to read them again.
Looks like Java culture is alive, well and spreading.
https://kubernetes.io/docs/concepts/scheduling-eviction/assi...
Two guesses as to where the folks who implemented Kubernetes hail from (as in: which programming language)
Go, obviously, though from their rationale article they “love” C/C++, Java, and Python, too.
It's basically unusable in Firefox or Safari.