For AWS and GCP, there's zone-specific VM scaling groups that ensure you're in multiple AZs so you'd know for sure you're not on a single physical server.
"You can easily get a box that has this many (or more) cores. I wouldn't be surprised if our cloud provider analyzed the topology of the connections, and decided to put the whole thing on the same server."
And the comment you replied to is referring to "the server"
It involves copying the whole VM image over, and rewiring the network connections virtually on the fly.
I say (nearly) because it's not 100% transparent, but my understanding is that it works properly the vast majority of time.
The question was how you reboot the posited singular hardware server with no downtime to any VMs running on it.
If the primary catches fire or crashes, the secondary boots the VM quickly without data loss.
If the primary reboots, it checkpoints RAM to the secondary pauses the VM and unpauses it on the secondary a few milliseconds later.
Note that this transparently handles storage replication, which is something that is notoriously difficult in Kubernetes.
If your cluster fits on one machine, you've paid for a lot of (currently) unnecessary complexity up front, both in dev time and in hardware cost.
If your app scales to need the cluster, congrats. Sometimes delaying time to market to allow for a smoother ramp makes sense. Sometimes it does not.
(Wikipedia is a good example of succeeding without ever needing to scale out the back end. I doubt they'd have won out over their competition if they took an extra 12-24 months to launch.)
So you don't have one VM's, but two. In case of Op's scenario you are paying for 196 CPUS instead of 98.
> If the primary reboots, it checkpoints RAM to the secondary pauses the VM and unpauses it on the secondary a few milliseconds later.
So a very small downtime, but not 0.
> You can easily get a box that has this many (or more) cores. I wouldn't be surprised if our cloud provider analyzed the topology of the connections, and decided to put the whole thing on the same server.
I wouldn't want to have paid for the engineering that went into that if you could actually do that.
Yes. GCP literally has a button that allows you to Replace/Restart all the instances in an Instance Group. Our deployment system just pushes a new GCP template and GCP does all the work rolling it out. No Kubernetes.
However, a permanent machine failure or unexpected reboot implies a few seconds or minutes of downtime.
These systems usually use a local disk, a synchronously updated disk across town (low latency, but on a different power grid, outside most natural disaster blast radii), and a far away asynchronously updated disaster recovery disk.
The disaster recovery disk provides a crash-consistent state that's less than a few seconds out of date.
These days, the "disks" tend to be deduped and compressed SSD's that use parity encoded raid. Their hardware cost is lower than a single copy on ext4, but they come with an enterprise tax / support contract, etc.
In practice, it's durable unless corporate HQ is wiped out. If that happens, it won't be the weak link in your business continuity plan (but a few customers' orders might've been dropped).
This stuff is tremendously unsexy, but it'll plug along just fine until you have some workload that can't be partitioned, and that needs more than roughly 40-100gbit of network/disk bandwidth, 128 cores, or 1TB of DRAM.
(Edit: Forgot to mention that the the apps run inside a VM infrastructure that supports live migration.)
https://docs.microsoft.com/en-us/azure/aks/availability-zone...
You might be able to get better utilisation of your VMs by running a smaller number with more cores on each - I'd try 16vcpu nodes.
And, yes, as a general recommendation, I would tend towards bigger, rather than smaller, nodes. There's a trade-off to be had, but you want each node to be at LEASt as large as the largest pod you intend to deploy, and you probably want at least some spare, on top of that.
I don't see what you win here. I don't want to have to write my own rolling updates, canary deployments, networking/traffic/load balancing scripts, configuration/secret management, dependency management, log aggregation, user management, etc.
I just spin up a Kubernetes cluster, create service accounts for my users, and let them use the k8s docs and declarative primitives to ship their code. If I'm lucky they are already familiar with the system from their last job, so they can bring expertise with them.
You can use inter-pod affinity and anti-affinity[1] to basically say "don't run the instances of this scaling group on the same physical nodes". Then, when the physical VMs are upgraded for whatever reason, the cluster knows not to bin all the Pods hosting your blog onto the same physical node.
Combined with how, e.g., Google handles node upgrades, means that I can keep services up and running pretty seamlessly without a lot of work. In fact, I haven't had downtime due to an upgrade of the underlying nodes or VMs in over 6 years, despite Google regularly updating the underlying nodes to newer host OS and kubelet versions.
You can also use node selectors[2] and related configurations to run specific pods on specific nodes or types of nodes (e.g., only run this pod on a node with a fast disk).
We do this with all of our Jobs. They are assigned to nodes in an autoscaling nodepool that scales down to zero. So, those nodes are only ever in use when we have jobs running, and jobs coming and going doesn't create any thrashing amongst our long running services.
Of course, your cluster operator could ignore all that somehow, but then when you have unexpected downtime during a machine or VM host upgrade, you'll be pretty peeved.
You do bring up a good point regarding node core counts. If you are running "30ish VMs ~= 98 cores", then you are probably running 2 and/or 4 core nodes. You may actually want to run "fatter" nodes with higher core counts. Keep in mind that each node is going to have dedicated resources for the kubelet, and any daemonset pods, which can add up pretty quickly sometimes. You may actually be wasting quite a bit of resources to overly redundant overhead. You may also suffer from slower scale-ups because you have to add a lot more individual nodes when things start to ramp up. There's a balance to find, not just with perf and overhead but with resiliency, but it's something I see a lot of folks overlook because they think they want to have super fine-grained node scaling.
[1] https://kubernetes.io/docs/concepts/scheduling-eviction/assi...
[2] https://kubernetes.io/docs/concepts/scheduling-eviction/assi...