AWS vs. GCP reliability is wildly different
freeman.vc
freeman.vc
I suspect the 409 conflicts are probably from the instance name not being unique in the test. It looks like the instance name used was:
instance_name = f"gpu-test-{int(time())}"
which has a 1-second precision. The test harness appears to do a `sleep(1)` between test creations, but this sort of thing can have weird boundary cases, particularly because (1) it does cleanup after creation, which will have variable latency, (2) `int()` will truncate the fractional part of the second from `time()`, and (3) `time.time()` is not monotonic.I would not ask the author to spend money to test it again, but I think the 409s would probably disappear if you replaced `int(time())` with `uuid.uuid4()`.
Disclosure: I work at Google - on Google Compute Engine. :-)
Why would any tenant supplied data affect anything whatsoever?
As a tenant, unless you are clashing with another resource under your own name, I don't see the point of failing.
aws S3 would be an exception, where they make that limitation on globally unique bucket name very clear.
You're inserting a VM with a specific name. If you try to create the same resource twice, the GCE control plane reports that as a conflict.
What they're doing here would be roughly equivalent to supplying the time to the AWS RunInstances API as an idempotency token.
(I work on GCE, and asked an industry friend at AWS about how they guarantee idempotency for RunInstances).
More practically, though, the instance name here is literally the name of the instance as it appears in the RESTful URL used for future queries about it. The 409 here is rejecting an attempt to create the same explicitly named resource twice.
When trying to create the same resource twice, all request should report the same status instead one failing, one succeeding.
In AWS, their APIs allow you to supply a client token if the API is not idempotent by default.
See https://docs.aws.amazon.com/AWSEC2/latest/APIReference/Run_I....
Before I quibble with the idempotency point: I agree with this, entirely, but it is what it is and a lot of software has been written against the current behavior. So I'll cite Hyrum's law here: https://www.hyrumslaw.com/
> GCP control plane is generally not idempotent.
The GCE API occupies an odd space here, imo. The resource being created is, in practice, an operation to cause the named VM to exist. The operation has its own name, but the name of the VM in the insert operation is the name of the ultimate resource.
Net, the API is idempotent at a macro level in terms of the end-to-end creation or deletion of uniquely named resources. Which is a long winded way of saying that you're right, but that from a practical perspective it accomplishes enough of the goals of a truly idempotent API to be _useful_ for avoiding the same things that the AWS mechanism avoids: creation of unexpected duplicate VMs.
The more "modern" way to do this would be to have a truly idempotent description of the target state of the actual resource with a separate resource for the current live state, but we live with the sum of our past choices.
You're right, we did it wrong.
// And paradoxically makes engineers like you.
Disclaimer: I work on GCE.
This is nicer!
Note that whether this creates collisions is entirely under the customer's control. There's no requirement for global uniqueness, just a requirement that you not try to create two VMs with the same name in the same project in the same zone.
You can also batch API calls[1], which also gives you a response for each VM in the batch while allowing for a single HTTP request/response.
That said, if you want to create a set of effectively identical VMs all matching a template (i.e., cattle not pets), though, or you want to issue a single API call, we'd generally point you to managed instance groups[2] (which can be manually or automatically scaled up or down) wherein you supply an instance template and an instance count. The MIG is named (like nearly all GCP resources), as are the instances, with a name derived from the MIG name. After creation you can also have the group abandon the instances and then delete the group if you really wanted a bunch of unmanaged VMs created through a single API call, although I'll admit I can't think of a use-case for this (the abandon API is generally intended for pulling VMs out of a group for debugging purposes or similar).
For cases where for whatever reason you don't want a MIG (e.g., because your VMs don't share a common template). You can still group those together for monitoring purposes[3], although it's an after-creation operation.
The MIG approach sets a _goal_ for the instance count and will attempt to achieve (and maintain) that goal even in the face of limited machine stock, hardware failures, etc. The top-level API will reject (stock-out) in the event that we're out of capacity, or in the batch/bulk case will start rejecting once we run out of capacity. I don't know how AWS's RunInstances behaves if it can only partially fulfill a request in a given zone.
[0]: https://cloud.google.com/compute/docs/instances/multiple/cre...
[1]: https://cloud.google.com/compute/docs/api/how-tos/batch
[2]: https://cloud.google.com/compute/docs/instance-groups
[3]: https://cloud.google.com/compute/docs/instance-groups/creati...
Is that not the conclusion? The tester was clashing with their own names?
The regions used are "us-east-1" for AWS [1] and "us-central1-b" for GCP [2].
1: https://github.com/piercefreeman/cloud-gpu-reliability/blob/...
2: https://github.com/piercefreeman/cloud-gpu-reliability/blob/...
Reminds of me this post on mtime which recently resurfaced on HN: https://apenwarr.ca/log/20181113
I just winced in pain thinking of the ways that can bite you. I guess in a cloud/virtualized environment with many short lived instances it isn't even that obscure an issue to run into.
A nice discussion on Stack Overflow:
https://stackoverflow.com/questions/64497035/is-time-from-ti...
Something similar caused my favorite bug so far to track down.
We were seeing odd spikes in our video playback analytics of some devices watching multiple years worth of video in < 1 hour.
System.currenTimeMillis() in Java isn't monotonic either is my short answer for what was causing it. Tracking down _what_ was causing it was even more fun though. Devices (phones) were updating their system time from the network and jumping between timezones.
Bosses actually came to us because our analytics team was trying to figure out who was causing it, because it had been caught by the team doing checks against the data. (a playback period should never have had > 30s of time)
Unfortunately, millisecond-precise timestamps proved to be a bit tricky in combination with sqlite.
503 would mean the IaaS API calls themselves are unavailable. Very different from the API working perfectly fine but the instances not being available.
Why would you think HTTP status codes are made for REST? They are made for HTTP to describe the response of the resource you are requesting, and the AWS API uses HTTP so it makes sense to use HTTP status codes.
”10.4.10 409 Conflict
The request could not be completed due to a conflict with the current state of the resource. This code is only allowed in situations where it is expected that the user might be able to resolve the conflict and resubmit the request. The response body SHOULD include enough information for the user to recognize the source of the conflict. Ideally, the response entity would include enough information for the user or user agent to fix the problem; however, that might not be possible and is not required.
Conflicts are most likely to occur in response to a PUT request. For example, if versioning were being used and the entity being PUT included changes to a resource which conflict with those made by an earlier (third-party) request, the server might use the 409 response to indicate that it can't complete the request. In this case, the response entity would likely contain a list of the differences between the two versions in a format defined by the response Content-Type.”
Then again, perhaps it is the service itself making that state change.
Amazon is pretty good about this, if their API says machine is ready, it usually is.
Although, it seems the author couldn't find out why they occurred, which points to poor error messages and/or lacking documentation.
If the instance takes too long to launch then it doesn't matter if it's "reliable" once it's running. It took too long to even get started.
What is your definition of reliability?
in other words, reliability is that it does what you expect it to. GCP does not have any particular guarantees around being able to spin up VMs fast, so its inability to do so wouldn't make it unreliable. it would be like me saying that you're unreliable for not doing something when you never said you were going to.
if this were comparing Lambda vs Cloud Functions, who both have stated SLAs around cold start times, and there were significant discrepancies, sure.
so that's why in engineering it's not really used as such. (as far as I understand at least.)
Calling this reliability is like saying a Ford is more reliable than a Chevy because the Ford has a better throttle response.
Like the article said, The promise of the cloud is that you can easily get machines when you need them the cloud that sometimes does not get you that machine(or does not get you that machine in time) is a less reliable cloud than the one that does.
The race car that finishes first is not “more reliable” than the one in 10th. They are equally as reliable, having both finished the race. The first place car is simply faster at the task.
If the article were measuring HTTP response times and found that AWS's average response time was 50ms and GCP's was 200ms, and both returned 200s for every single request in the test, would you say AWS is more reliable than GCP based on that? Of course not, it's asinine.
The promise of the cloud is that you can flexibly spin up machines if available, and easily spin down, no long term contracts or CapEx etc. They are all pretty clear that there are capacity limits under the hood (and your account likely has various limits on it as a result).
Consistently slow is still reliability.
Midnight - 6am is six hours. The on demand price for a G5 is $1/hr. That's over $2K/yr, or "an extra week of skiing paid for by your B2B side project that almost never has customers from ~9pm west coat to ~6am east coast". And I'm not even counting weekends.
But that's sort of a silly edge case (albeit probably a real one for lots of folks commenting here). The real savings are in predictable startup times for bursty work loads. Fast and low variance startup times unlock a huge amount of savings. Without both speed and predictability, you have to plan to fail and over-allocate. Which can get really expensive fast.
Another way to think about this is that zero isn't special. It's just a special case of the more general scenario where customer demand exceeds current allocation. The larger your customer base, and the burstier your demand, the more instances you need sitting on ice to meet customers' UX requirements. This is particularly true when you're growing fast and most of your customers are new; you really want a good customer experience every single time.
Maintaining a buffer pool is hard. You need to maintain state, have a prediction function, track usage through time, etc. just spinning up new nodes for new work is substantially easier.
And the author said he could spin up new nodes in 15 seconds, that’s pretty quick.
The errors might be considered a reliability issue, but then again, errors are a very common thing in large distributed systems, and any orchestrator/autoscaler would just re-try the instance creation and succeed. Again, a performance impact (since it takes longer until your target capacity is reached) but reliability? not really
In many cases, "guaranteed" just means "we'll give you a refund if we fuck up". SLAs are very much like this.
IN PRACTICE, unless you're launching tens of thousands of instances of an obscure image type, reasonable customers would be able to get capacity, and promptly from the cloud.
That's the entire cloud value proposition.
So no, you can't just hand-waive past these GCP results and say "Well, they never said these were guaranteed".
That said, while I agree that launch time and provisioning error rate are not sufficient to define reliability, they are definitely a part of it.
For this, I'd prefer a title that lets me draw my own conclusions. 84 errors out of 3000 doesn't sound awful to me...? But what do I know – maybe just give me the data:
"1 in 3000 GPUs fail to spawn on AWS. GCP: 84"
"Time to provision GPU with AWS: 11.4s. GCP: 42.6s"
"GCP >4x avg. time to provision GPU than AWS"
"Provisioning on GCP both slower and more error-prone than AWS"
yeah i guess it does make sense that one didn’t win the a/b test
84 times more launch errors seems like a valid definition for "less reliable".
If I depend on some performance metric, startup, speed, etc, my dependance on it equates to reliability. Not just on/off but the spectrum that it produces.
If a CPU doesn't operate at its 2GHz setting 60% of the time, I would say that's not reliable. When my bus shows up on time only 40% of the time - I can't rely on that bus to get me where I need to go consistently.
If the GPU took 1 hour to boot, but still booted, is it reliable? What about 1 year? At some point it tips over an "personal" metric of reliability.
The comparison to AWS which consistently out-performs GCP, while not explicitly, implicitly turns that into a reliability metric by setting the AWS boot time as "the standard".
Here it's the possibility to launch new VMs to satisfy dynamic projects' needs. Cloud provider should allow you to scale-up in a predictable way. When it doesn't - it can be called unreliable.
Also, "unreliable" is basically a synonym for "Google" these days.
The fundamental problem with cloud reliability is that it depends on a lot of stuff that's out of your control, that you have no visibility into. I have had services running happily on AWS with no errors, and the next month without changing anything they fail all the time.
Why? Well, we look into it and it turns out AWS changed something behind the scenes. There's a different underlying hardware behind the instance, or some resource started being in high demand because of some other customers.
So, I completely believe that at the time of this test, this particular API was performing a lot better on AWS than on GCP. But I wouldn't count on it still performing this way a month later. Cloud services aren't like a piece of dedicated hardware where you test it one month, and then the next month it behaves roughly the same. They are changing a lot of stuff that you can't see.
Some regions and hardware generations are just busier than others. It may not be the same across cloud providers (although I suspect it is similar given the underlying market forces).
At a surface level, the above (from the article) seems like a pretty straightforward explanation? GCP gives you more flexibility in configuring GPU instances at the trade off of increased startup time variability.
It’s neat…but like a lot of things in large scale operations, the devil is in the details. GPU-CPU communications is a low latency high bandwidth operation. Not something you can trivially do over standard TCP. GCP offering something like that without the ability to flawlessly migrate the VM or procure enough “local” GPUs means it’s just vaporware.
As a side note, I’m surprised the author didn’t note the amount of ICE’s (insufficient capacity errors) AWS throws whenever you spin up a G type instance. AWS is notorious for offering very few G’s and P’s is certain AZs and regions.
And NVIDIA's vGPU solutions do support live migration of GPUs to another host (in which case the vGPU gets moved too, to a GPU on that target).
I didn't understand how they were able to do this, I had thought volume types mapped to hardware clusters of some kind. And since I didn't understand, I wasn't able to distinguish it from magic.
Essential, memory state is copied to the new host, the VM is stunned for a millisecond and the cpu states is copied and resumed on the new host (you may see a dropped ping). All the networking and storage is virtual anyway so that is "moved" (it's not really moved) in the background.
Very cool.
This conjures up hilarious mental imagery, thanks
The disks are all on the network, so no need to move anything there.
This has memory performance characteristics - I ran a benchmark of memory read/write speed while this was happening once. It more than halved memory speed for the 30s or so it took from migration started to migration complete. The pause, too, was much longer.
https://cloudplatform.googleblog.com/2015/03/Google-Compute-...
And could it be phrased differently as "EC2 doesn't do live migration badly"?
AWS doesn't have live migration at all. You have to stop/start.
Azure technically does, but it doesn't always work(they say 90%). 30 seconds is a long time.
VMWare has live migration (and seems to be the closest to what GCP does) but it is still an inferior user experience.
This is the key thing you are missing – GCP not only has live migration, but it is completely transparent. We do not have to initiate migration. GCP does, transparently, 100% of the time. We have never even notice migrations even when we were actively watching those instances. We don't know or care what hypervisors are involved. They even preserve the network connections.
https://cloudplatform.googleblog.com/2015/03/Google-Compute-...
VMware is lightyears ahead of the big clouds, but unfortunately they "missed the boat" on the public cloud, despite having superior foundational technology.
For example:
- A typical vSphere cluster would use live migration to balance workloads dynamically. You don't notice this as an end user, but it allows them to bin-pack workloads up above 80% CPU utilisation in my experience with good results. (Especially if you allocate priorities, min/max limits, etc...)
- You can version-upgrade a vSphere cluster live. This includes rolling hypervisor kernel upgrades and live disk format changes. The upgrade wizard is a fantastic thing that asks only for the cluster controller name and login details! Click "OK" and watch the progress bar.
- Flexible keep-apart and keep-together rules that can updated at any time, and will take effect via live migration. This is sort-of like the Kubernetes "control loops", but the migrations are live and memory-preserving instead of stop-start like with containers.
- Online changes to virtual hardware, including adding not just NICs and disks, but also CPU and memory!
- Thin-provisioned disks, and memory deduplication for efficiencies approaching that of containerisation.
- Flexible snapshots, including the ability for "thin provisioned" virtual machines to share a base snapshot. This is often used for virtual desktops or terminal services, and again this approaches containerisation in terms of cloning speed and storage efficiency.
In other words, VMware had all of the pieces, and just... didn't... use it to make a public cloud. We could have had "cloud.vmware.com" or whatever 15 years ago, but they decided to slowly jack up the price on their enterprise customers instead.
For comparison, in Azure: You can't add a VM to an availability set (keep apart rule) or remove the VM from it without a stop-start cycle. You can't make most changes (SKU, etc...) to a VM in an availability set without turning off every machine in the same AS! This is just one example of many where the public cloud has a "checkbox" availability feature that actually decreases availability. For a long time, changing an IP address in AWS required the VM to be basically blown away and recreated. That brought back memories of the Windows NT 4 days in 1990s when an IP change required a reboot cycle.
No, VMware didn't miss the boat, vCloud Air was announced in 2009 and made generally available in 2013. Roughly same timelines as Azure and GCP, slightly trailing AWS, and those were the early days, where the public cloud was still exotic. And VMware had the massive advantage of brand recognition in that domain and existing footprint with enterprises which could be scaled out.
Problem was, vCloud Air, like vSphere, was shit. Yeah, it did some things well, and had some very nice features - vMotion, DRS (though it doesn't really use CPU ready contention for scheduling decisions which is stupid), vSAN, hot adding resources (but not RAM, because decades ago Linux had issues if you had less than 4GB RAM and you added more, so to this day you can't do that). When they worked, because when they didn't, good luck because error messages are useless, logs are weirdly structured and uselessly verbose, so a massive pain to deal with. Oh and many of those features were either behind a Flash UI(FFS), or an abomination of an API that is inconsistent ("this object might have been deleted or hasn't been created yet") and had weird limitations like when you have an async task you can't check it's status details. And many of those features were so complex, that a random consuming user basically had to rely on a dedicated team of vExperts, which often resulted in a nice silo slowing everyone down.
Their hardware compatibility list was a joke - the Intel X710 NIC stayed on it for more than a year with a widely known terribly broken driver.
But what made VMware fail the most, IMHO, was the wrong focus, technically - VM, instead of application. A developer/ops person couldn't care less about the object of a VM. Of course they tried some things like vApp and vCloud Director etc. which are just disgusting abominations designed with a PowerPoint in mind, not a user. And pricing. Opaque and expensive, with bad usability. No wonder everyone jumped on the pay as you go, usable alternatives.
My introduction to the industry. The memories.
But I think you underestimate the maturity and effectiveness of the underlying google compute and storage substrate.
(FWIW, I worked at both places)
Now how the Google's substrate maps onto GCP, that's another story. There is a non trivial amount of fluff to be added on top of your building blocks to build a manageable multitenant planet scale cloud service. Just the network infrastructure is mind boggling.
I wouldn't be surprised if your experience with a "VMware cloud" would surprise you if you naively compare it with your experience with a standalone vsphere cluster.
From the FAQ: https://aws.amazon.com/ec2/faqs/
Q: How does EC2 perform maintenance?
AWS regularly performs routine hardware, power, and network maintenance with minimal disruption across all EC2 instance types. To achieve this we employ a combination of tools and methods across the entire AWS Global infrastructure, such as redundant and concurrently maintainable systems, as well as live system updates and migration.
And yet, I keep getting almost every weeks emails like this:
"EC2 has detected degradation of the underlying hardware hosting your Amazon EC2 instance (instance-ID: i-xxxxxxx) associated with your AWS account (AWS Account ID: NNNNN) in the eu-west-1 region. Due to this degradation your instance could already be unreachable. We will stop your instance after 2022-09-21 16:00:00 UTC"
And we don't have tens of thousands of VMs in that region, just around 1k.
... then it seems like a device that limits bandwidth either on the storage cluster or between the node and storage cluster is present. 125MiB/s is right around the speed of a 1gbit link, I believe. That it was a networking setting changed in-switch doesn't seem to be surprising.
This is also how they are able to snapshot a volume at a certain point in time without having any downtime or data inconsistencies.
https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/modify-v...
** NOTE: If you're a low scale company this won't matter to you **
1. GKE
When you cross a certain scale certain GKE components won't scale with you and SLOs on those components are crazy, it takes 15+ mins for us to update a GKE ingress controller backed Ingress.
Cloud Logging hasn't been able to keep up with our scale, disabled since 2 years now. This last Q we got an email from them to enable it and try it again on our clusters, still have to confirm these claims as our scale is more higher now.
Konnectivity agent release was really bad for us, it affected some components internally, total dev time we lost was more than 3 months debugging this issue. They had to disable konnectivity agent on our clusters, I had to collect TCP dumps and other evidences just to prove nothing was wrong on our end, fight with our TAM to get a meeting with the product team. After 4 months they agreed and reverted our clusters to SSH tunnels. Initially GCP support said they said they can't do this. Next Q Ill be updating the clusters hopefully they have fixed this by then.
2. Support.
I think AWS support always were more pro active in debugging with us, GCP support agents most of the times lack the expertise or proactiveness to debug/solve things in simple cases. We pay for enterprise support and don't see getting much from them. At AWS we had reviews of the infra how we could better it every 2 Qs and we got new suggestion and was also the time when we shared what we would like to see in their roadmap.
3.Enterprisyness is missing with design
A simple thing as cloudbuild doesn't have access to static IPs. We have to maintain a forward proxy just cause of this.
L4 LBs were a mess you could only use specified ports in a (L4 LB) TCP proxy, For a tcp proxy based loadbalancer, the allowed set of ports are - [25, 43, 110, 143, 195, 443, 465, 587, 700, 993, 995, 1883, 3389, 5222, 5432, 5671, 5672, 5900, 5901, 6379, 8085, 8099, 9092, 9200, and 9300]. Today I see they have removed these restrictions. I don't know who came up with this idea to allow only a few ports on a L4 LB. I think such design decisions make it less Enterprisy.
However, there's one aspect where GCP is a clear winner on the reliability front. They auto-migrate instances transparently and with close to zero impact to workloads – I want to say zero impact but it's not technically zero.
In comparison, in AWS you need to stop/start your instance yourself so that it will move to another hypervisor(depending on the actual issue AWS may do it for you). That definitely has impact on your workloads. We can sometimes architect around it but there's still something to worry about. Given the number of instances we run, we have multiple machines to deal with weekly. We get all these 'scheduled maintenance' events (which sometimes aren't really all that scheduled), with some instance IDs(they don't even bother sending the name tag), and we have to deal with that.
I already thought stop/start was an improvement on tech at the time (Openstack, for example, or even VMWare) just because we don't have to think about hypervisors, we don't have to know, we don't care. We don't have to ask for migrations to be performed, hypervisors are pretty much stateless.
However, on GCP? We had to stop/start instances exactly zero times, out of the thousands we run and have been running for years. We can see auto-migration events when we bother checking the logs. Otherwise, we don't even notice the migration happened.
It's pretty old tech too:
https://cloudplatform.googleblog.com/2015/03/Google-Compute-...
https://learn.microsoft.com/en-us/previous-versions/windows/...
Huh... interesting, this has not been my experience with Azure VM launch times. I'm usually surprised how quickly they pop up.
Premium SSD allows 30 minutes of "burst" IOPS, which can bring down boot times to about 2-5 seconds for a typical Windows VM. The provisioning time is a further 60-180 seconds on top. (The fastest I could get it is about 40 seconds using a "smalldisk" image to ephemeral storage, but then it took a further 30 seconds or so for the VM to become available.)
Standard HDD was slow enough that the boot phase alone would take minutes, and then the VM provisioning time is almost irrelevant in comparison.
FWIW this article is saying the opposite--it's AWS that beats GCP in startup speed.
The reason, from what I understand, why GCP does live migration more is because ec2 focused on live updates instead of live migration. Whereas GCP migrates instances to update servers, ec2 live updates everything down to firmware while instances are running.
Curious, what instance types are you using on EC2 that you see so much maintenance?
We use a bunch of different types. M5 and R5 (different sizes) are the most commonly used types but we use many different families. I haven't done an analysis to figure out which types are hotspots.
This is across thousands of instances over many regions worldwide. The percentage is low, but that still translates to daily maintenance alerts.
X = Compute intances
Y = Launch
Z = Time to launch
T = LSL (N/A), USL (10s), Std Dev (2s)
Where LSL is lower spec limit, USL is upper spec limit. LSL is N/A since we don't care if the instance launches instantly (0 seconds).You can define T as per your requirements. Here we are ignoring the accuracy of the clock that measures time, assuming that the measurement device is infinitely accurate.
If your criteria is to, say for example, define reliability as how fast it shuts down, then this article isn't relevant. Article is pretty narrow in testing reliability, they only care about launch time.
If both AWS and GCP had the same SLA, and one did better than the other at starting up, you could say one is more performant than the other, but you couldn't say it's more reliable if they are both meeting the SLA. It's easy to look at something that never goes down and say "that is more reliable", but it might have been pure chance that it never went down. Always read the fine print, and don't expect anything better than what they guarantee.
> why I burned $150 on GPUs
How do you rent 3000 GPUs over a period of weeks for $150? Were they literally requisitioning it and releasing it immediately? Seems like this is quite a unrealistic type of usage pattern and would depend a lot on whether the cloud provider optimises to hand you back the same warm instance you just relinquished.
> GCP allows you to attach a GPU to an arbitrary VM as a hardware accelerator
it's quite fascinating that GCP can do this. GPUs are physical things (!) do they provision every single instance type in the data center with GPUs? That would seem very expensive.
However, live-migration can cause impact to HPC workloads.
...if there are any GPUs available in the AZ that is. I had a hell of a time last year moving back and forth between regions to grab just 1 GPU to test something. The web UI didn't have a "any region" option for launching VMs so if you don't use the API you'll have to sit there for 20 minutes trying each AZ/region until you managed to grab one.
Was asking myself the same question. From the pricing information on gcp it seems minimum billing time is 1 minute, making 3000 GPUs cost $50 minimum. If this is the case then $150 is reasonable for the kind of usage pattern you describe.
Spawning a single GPU at varying times is nothing. Try spawning more than one, or using spot instances, and you’ll get a very different picture. We often run into capacity issues with GPU and even the new m6i instances at all times of the day.
Very few realistic company size workloads need a single GPU. I would willingly wait 30 minutes for my instances to become available if it meant all of them where available at the same time.
I have always been feeling there is so little independent content on benchmarking the IaaS providers. There is so much you can measure in how they behave.
My personal opinion is that Google's resources are more tightly optimized than AWS and they may try to find the 99% best allocation versus the 95% best allocation on AWS.. and this leads to more rejected requests. Open to being wrong on this.
I spend a significant fraction of my 11+ years there clicking Reload on my job's borg page. I was able to (re-)start ~100K jobs globally in about 15 minutes.
The origin for the info that jobs take "minutes" likely involves jobs that were pending available resources. This is a valid state in Borg, but GCE has additional admission control mechanisms aimed at avoiding extended residency in pending.
As dekhn notes, there are many factors that contribute to VM startup time. GPUs are their own variety of special (and, yes, sometimes slow), with factors that mostly don't apply to more pedestrian VM shapes.
In my limited experience, persistent (on-demand) GCP instances always boot up much faster than AWS EC2 instances.
Definitely seems like interesting info, though.
bit of a stretch, right
This is a good point and should be part of the test: after launching, SSH into the machine and run a trivial task to confirm that the hardware works.
That would seem to indicate that asking for a VM on GCP gets you a minimally configured VM on basic hardware, and then it gets migrated to something bigger if you ask for more resources. Is that correct?
That could make sense if, much of the time, users get a VM and spend a lot of time loading and initializing stuff, then migrate to bigger hardware to crunch.
This is going to vary a lot based on the time of year. Why don't you try this same experiment at around some time when there's a lot of retail sales activity (Black Friday), and watch AWS suddenly have much less capacity to dole out on-demand.
To me reliability is a measure of what a cloud does compared to what it says it will do. GCP is not promissing you on-demand instances instantaneously is it? If you want that ... reserve capacity.
GCP on the other hand fills all machines with background jobs. When you want a machine, they need to terminate a background job to make room for you. That background job has a shutdown grace time. Usually thats 30 seconds.
Sometimes, to prevent fragmentation, they actually need to shuffle around many other users to give you the perfect slot - and some of those jobs have start-new-before-stop-old semantics - that's why sometimes the delay is far higher too.
I haven't used GCP much, but maybe they load the image onto the node prior to launch, accounting for some of the launch time difference?
We find reliability a diff story. Eg, our main source of downtime on Azure is they restart (live migrate?) our reserved T4s every few weeks, causing 2-10min outages per GPU per month.
Is the POW mining part true any more? Hasn't mining moved to dedicated hardware?
it would be like doing this in us-central1 when us-central1 is down for one provider, and not another, resulting in increased latency, and saying how much faster one is than the other.
unlike say a throughput test or similar, neither of these services promise particular cold-starts, and so the numbers here cannot be contexutalized against any metric given by either company and so are only useful in the sense that they can be compared, but since there are no guarantees the positions could switch anytime.
that's why I like comparisons between serverless functions where there are pretty explicit SLAs and what not given by each company for you to compare against, as well as one another.
That graph is a pain to see.
The word "Google" attached to anything is a strong indicator that you should look for an alternative.