Does anyone know what the author is referring to with this claim? I don't see anything in the sheet to back this up. At least from a high level, all three options support network policies via CNI, and GKE and EKS use the same one, Calico.
Does anyone know what the author is referring to with this claim? I don't see anything in the sheet to back this up. At least from a high level, all three options support network policies via CNI, and GKE and EKS use the same one, Calico.
"Each cluster receiving an IP range for nodes and another for the containers inside, which are directly routable across your private network, other clusters, and regions."
AWS and Azure don't have a flat global network. You have to setup VPN's and complicated overlay networks.
AWS has cross-region VPC peering, which is simple to setup and makes it feel “flat”. There’s no need for VPNs or overlay networks.
Also, while setting up peering is a 10-20 minute task there are lot of constraints, such as the VPC's have to have different CIDR ranges. Oops, if you go with the AWS default, probably all regions using 172.31.0.0/20.
Google got networking right with a flat global any region setup. If you need segmentation then just create a new GCP project. If it turns out projects need to communicate after all, GCP has VPC network peering.
Two different design philosophies, there are been a global GCP networking outages due to having the single flat network. A few this year. All providers have had their issues, but, global blast radiuses concern me.
Cross-region networking. To this day, I'm flabbergasted this isn't something that AWS has. All the other networking bits seem more sensible to me as well.
Another, is 2Gbps per core up to 8 cores which is way more throughput than I've seen on any AWS or Azure instance.
On workloads where I care about Network IO:CPU ratio, I'm using 8 core nodes and seeing 16gbps throughput between them consistently.
AWS tries to avoid letting customers create multi-region or global failure modes, like:
* https://status.cloud.google.com/incident/compute/18005 22-hour incident with GLOBAL impact resulting in many VMs getting duplicate IPs (this no network connectivity) on 2018-06-15
* https://status.cloud.google.com/incident/cloud-networking/18... 3-hour (seemingly GLOBAL) impact to load-balancers (no updates or creations) on 2018-01-03
* https://status.cloud.google.com/incident/compute/16007 - 18-minute GLOBAL outage on 2016-04-11
* https://status.cloud.google.com/incident/compute/15055 - 5-minute GLOBAL outage on 2015-08-04
* https://groups.google.com/forum/#!topic/gce-operations/fynnX... 43-minute multi-region packet loss due to multi-region configuration deployment on 2015-03-07
* https://groups.google.com/forum/#!topic/gce-operations/1uw-q... unscheduled reboot of 28% of instances in multiple regions in ~1.5h on 2014-09-17
There have been recent changes in many AWS services (e.g. DynamoDB, S3, Aurora) to allow cross-region replication without introducing any multi-region dependency, precisely to make it easier to implement multi-region infrastrucure to tolerate a single-region failure (however rare that is on AWS compared to global outages on GCP).
> Another, is 2Gbps per core up to 8 cores which is way more throughput than I've seen on any AWS or Azure instance.
Which instance types did use on AWS? Many recent instance types (r4, i3, c5, r5, z1) support up to 10Gbps with once core (2 vCPU instances). However, you may need to use placement groups ( https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/placemen... ) in large regions in order to get the full throughput on low TCP connection numbers. The only reason I can think of that would explain why GCP doesn't have this problem is they don't have any regions anywhere near as large as large AWS regions ...
Edit: white-space only formatting changes
As for the instances I tested, they were m4 and r4. To get full 10GE out of either, I needed to use m4.10xl and r4.8xl, both roughly around $900 per month.
By comparison an n1-standard-8 is about $60/mo and the n1-highcpu-8 is $50/mo. I've tested the former and got something like 15.8 or 15.9gbps using iperf.
IMHO, that's pretty significant.
Also in general GCP networking is better than aws ime (ever tried to have multiple projects in aws share a network?). Don’t know anything about Azure but I suspect msft still sucks at networking things so it’s not great.
I'm curious why you think that Calico's use of iptables will have performance impacts for large clusters, and what type of performance impacts you expect (bandwidth / latency / cpu / something else?).
From my experience, Calico performs rather well in large clusters (e.g. 2k nodes, appx 100 pods / node, several hundred network policies).
After an initial look I think he is referring not to the kubernetes specific networking support but the GCP networking.
From the spreadsheet alone GKE has support for cross region load balancers. The author also points out cross region networking.
I don't know the exact advantage the author sees with cross region load balancers since you can only deploy a GKE cluster to a single given region. The GCP LB's are really nice though in that traffic enters Google's network at the closest location to the user and then travel within Google's network, which is wicked fast.
I have not played around with federated clusters in multiple zones so don't know how that would play out either.
They may also just be referring to just general networking from the providers. With the first two points they at least make mention in the spreadsheet of it though.
tl;dr: not sure what the author meant but here here is some food for thought on it.