Some quick details: It's a bit of a split between what GKE runs and what the customer runs. Alpha runs on vSphere 6.5 and we're packing up a Google-hardened OS in much the same way we package GKE for GCP. A lot of the integrations for things like networking and storage will be coming from partners. We'll also have remote mgmt capabilities so we can manage the cluster's control plane in much the same way our SREs do for GKE.
> GKE On-Prem has a fully integrated stack of hardened components, including OS, container runtime, Kubernetes, and the cloud to which it connects.
Which runtime are you shipping? CRI-O? What type of outgoing cloud connection is that? I have so many questions. I'm actually at the conference this week if you're willing to grab coffee.
* DNS (not the Kubernetes in-cluster one like kube-dns)
* DHCP
* LDAP or equivalent
* SSH and its keys
* on-prem security of the cloud identities (does it require TPM? SGX?)
* bootloaders
* base OS image
* drivers for attached storage, GPUs, etc.
* firewall rules
* routing
Does it ship with ceph out of the box? Some in-house block store? What happens when it breaks? Persistent volumes are IMO the very hardest thing to get right, and for me it's the big reason why I'd rather put my trust in a hosted solution in the first place.
If they solve this, and make it as seamless and easy as using cloud storage offerings, they've completely changed the game. Somehow I think they have a ways to go.
on bare-metal this is solved with rook.io. load balancing (not the api servers) is also solved with metallb.
Do you have a source for this? Is there documentation anywhere that says GKE On-Prem is using rook? Or are you just saying "people who use kubernetes on-prem often use rook.io"?
I have yet to hear about any big production deployments of rook, care to provide any to support your claims?
We also have great storage abstraction layers built into K8S - CSI, FlexVolumes, and a large suite of in-tree plugins - so adding additional ones is pretty easy.
I think you're talking about something else however - specifically scaled/distributed storage services.
We're investigating options here, though keep in mind you should have no problems running containerized storage services in GKE On-prem. They would be on top of the existing block support I mentioned above. I saw a couple comments in other threads from some vendors that sell solutions that do just that.
For storage systems that don't containerize, that is a different discussion. Happy to talk more.
In the Alpha, we are supporting vSphere 6.5. Which part of infra are you most curious about knowing?
Is this ever going to be a bare-metal thing? Like probably many others, I'm not really interested in doing on-prem virtualization... kubernetes is interesting to me because containers are a better abstraction than virtual machines in the first place. Why add a virtualization layer if you don't have to?
(I get that it makes your life easier as the developer of this product, but having to run a virtualization IaaS between your metal and your orchestration makes the whole thing rather uninteresting IMO.)
We are exploring additional options, such as bare metal support, based on customer demand.
Send me an email (karangoel [at] google) if you'd be interested in talking more about bare metal.
so what is the model for bare metal with GKE on-prem. The reason is for VSPHERE 6.5 is additional cost (license per cores to VMWARE and VCENTER license) which we want to avoid to use bare metal only.
Walk before you run.
Bare metal is a LOT harder to manage because, well, hardware fails. We hear the demand, for sure, but vSphere represents walking (and has a lot of customers, too :)
But, I don't buy the hardware failure argument, because the same is true of running a vsphere installation in the first place.
vsphere migrates VMs to other machines when the hardware fails, but, analogously, the kube scheduler moves pods to other machines when they fail as well. You have to worry about disk failures in both cases. You have to worry about keeping your vsphere's database up and in a high-availability mode (postgres in my experience), just as you have to worry about keeping k8s's etcd cluster up and in a high-availability mode.
For any problem k8s has on bare metal due to hardware unreliability, vsphere has an analogous problem, it's just pushed down one layer.
IMO the real reason why this is a pragmatic decision is because, people already have lots of experience in running vsphere and understand where the risks and challenges are, and vsphere has lots of tools for things like automating the installation of the hypervisor OS, base level network setup, expectations around NFS for VM storage, etc.
Vsphere represents a decent, known set of tools for getting an infrastructure up and running on bare metal, which is a prerequisite for getting kubernetes running, but what would be exciting to me would be a rethinking of those infrastructure components in a purely open source and industry standard fashion, in a no-frills way that only exists to get a basic k8s control plane up.