My take at making AWS EC2 cheaper by automating SPOT instances with AutoScaling
mcristi.wordpress.com
mcristi.wordpress.com
- Preemptible VMs are sold at a fixed 70% off discount, removing pricing volatility entirely
- Google Cloud's Compute Engine has far fewer VM types, thus making it much easier to construct the resources you want (exception being GPUs, you don't need a "network optimized" instance to get fast network, nor you need a "storage optimized" instance to get fast/large storage - these things are modular on GCE).
(disc: work on Google Cloud)
Second, what are you looking to run? A Java-based web app? Does App Engine's free tier not serve you better?
Finallty, and it's perhaps poor form to point this out, but we (and AWS) will end up blowing out your budget once you include networking while OVH/Hetzner/Digital Ocean won't. Compute Engine isn't a VPS, we're giving you a small slice of a machine but it's connected to a crazy awesome network that we bill for by the byte. When you compare that to a bundled, heavily-overcommitted across tons of customers VPS networking offering the dollars don't work out.
Disclosure: I work on Compute Engine and care a lot about our pricing.
I ask because I don't need a CDN level network like GCE provides but I do need a ton of compute and a ton of bandwidth for batch data processing on data hosted at another provider (and not a candidate for GCS). A pre-emptible network offering would open up the way for me to use compute resources, which is surely a good thing.
Can you describe more about the "data hosted at another provider"? Would one of our interconnect options (https://cloud.google.com/interconnect/) help? If it's at AWS would a pair of interconnect options help (DirectConnect to Equinix in VA, then Carrier Interconnect to us).
> Can you describe more about the "data hosted at another provider"?
It's all at OVH and Hetzner. They offer very cheap storage and bandwidth.
> interconnect
I'm not big enough for those options and regardless, all of your peering/interconnect options are CDN priced for inter-region and only slightly better for intra-region. They just don't make sense.
Say I want to store 100TB of data and pump it through some kind of processing pipeline, outputting the same amount of data. At OVH I pay ~$0.0075/GB on storage coming to ~$1.5k/month (since the processing results in double storage). Say my processing is light on memory and can run at 1MB/s/CPU and I want to be finished in a day. I need 1150 CPU cores for 24h and at OVH that would cost me ~$1300 and bandwidth is free and unlimited. At GCE I can run on pre-emptible instances and pay ~$284 for the compute but I have to pay for 100TB of premium egress bandwidth I don't need, which adds 100 * 1000 * 0.04 (generously assuming I can get peering/interconnect) which is $4000. GCE is absolutely unusable for jobs like this in its current state.
Sure I could move data to GCS but in that case my $1.5k bill for storage turns into a $5.2k bill per month.
it's just that I need at least one vm that serves the web, this will upload the images to the object storage then it will fire premeetive vms and converts them, this will put the images in the real object storage and calls the always on vm to insert something in the database (i.e. image path) actually the one vm needs to contain a small database (we are talking about mb so memory isn't an issue) and a java app. so as said micro is too small and small is too big for that always on vm.
my current planned setup is ovh instance -> gcloud object storage -> preemtive vm (if there is one) -> gcloud object storage - ovh instance
btw. I made some calculations and with the 3.4 € from ovh and somtimes preemtive vms and something like 30g/10g object storage (which hopefully will raise, but than again no problem on the costs) will cost us less than 7€ + Domain (which is billed per year) btw. the preemtive vms runs something like 48 hours per month (mostly less since the most images will started between november < - > february)
Edit: btw i wanted to use the datastore (database) aswell, but I couldn't activate it in a project without a app engine.
However after creating a new project everything in gcloud started to work. Does the App Engine free tier apply to 'App Engine Flexible Environment' ? That would be great since that would make the site really really cheap.
Amazon is still eating everyone's lunch on the GPU front.
One major use case for GPUs is ML. Google Cloud externalizes its ML through serverless APIs. At that point what matters is your ability to derive value from ML, rather than having exposure to building blocks.
I work on some analytics software that includes GPU ML algorithms & non-ML GPU algorithms, neither of which Google makes. Our peers do a lot of visual computing in AWS. Really weird thing to hear from a Google rep.
"Yes, we know! Sorry! If you just need some ML things and can use the Cloud ML or Cloud Vision services we've got something to tide you over.
If not, we're always excited to get direct 'I want to do X' feedback that we can translate into 'Customers are demanding Y, and willing to pay Z'".
Please don't take the comment to mean we're trying to box people out of this space. We're not. We do think (most? many?) people don't want to actually roll their own, but we love everyone.
Disclosure: I work on Compute Engine.
http://docs.aws.amazon.com/AWSEC2/latest/UserGuide/spot-flee...
"and it seems they now have a full fledged solution for the problem, based on pretty much a reimplementation of AutoScaling, using machine learning and with a beautiful UI and they are really successful with it. Funnily enough, they even contacted me to sell that solution to my company and we are seriously evaluating it"
The problem with the spot fleet 1) it's kind of awkward to use 2) it has statically defined capacity so you can't scale it 3) it has a static bid price, so if at some point your are outbid on all the group's bids, you end up with no capacity 4) among other things it lacks integration with the ELB so you can't really use it for so many use cases.
My solution is simpler, better integrated with the rest of AWS and more resilient and once I iron out the bugs and get it production-ready, it should be a better choice.
Ended up going between two regions to avoid them but it's some hassle, just got a gtx970 in the end.
Say you have an autoscaling group spanned across zones A and B, and say you have 1 machine in zone A and 1 in machine B.
Now, the price in zone A goes up, and your machine in zone A dies.
The issue (bug?) was that the autoscaling group was trying to re-instantiate a new VM in zone A. Of course, since the price was high, the new VM was basically immediately dying. And so on.
Edit: issue apart, it's a great way that can save money, especially if you have a group of VMs whose computation can be interrupted/restarted relatively cheaply.
All I do is I later attempt to replace them with whatever I can buy from the spot market.
On the other hand, the spot bidding implemented out of the box in AutoScaling will fail if you are outbid in all Availability Zones at the same time, since it doesn't fall back to on-demand instances. I've seen people often use a second on-demand AutoScaling group that would scale out when you get outbid on the spot one, but then you have a problem defining scaling policies so that they can scale nicely, and/or shifting the capacity between them. Someone had a nice talk at re:invent about how they do all that.
I'm curious how does that tool handle the scaling of the group, since the nodes are actually outside the group and can't contribute to group-wide metrics like average CPU usage, often used for scaling out.
If I could make the Python one work a bit better, I'd probably use it with Ansible - http://snappishproductions.com/2014/03/24/Spot-Instances-Wit...
Disclosure: I work on Compute Engine (and launched our preemptible VMs product).