Andromeda 2.1 reduces GCP’s intra-zone latency by 40%
cloudplatform.googleblog.com
cloudplatform.googleblog.com
Disclosure: I work on Google Cloud.
Regarding in-guest iptables -- there's not much we can do on GKE/Kubernetes about that. I serve about 1.2M requests per second on my GKE cluster through GLB and see this overhead very well.
FWIW I also look forward to IPVS in k8s 1.9 that should improve this slightly.
That reminds me though: you have to be communicating over the internal IP addresses, not the external ones (I hope the Services stuff does the right thing automatically). Firewall rules and such are different for internal versus external IPs (people often want a simple "deny all external traffic"), so that's sadly still a performance cliff.
I do get impressive bandwidth (1.95 Gbits/sec) using iperf.
There're 8 comments on this thread and 3 of them are from GCP people giving less marketing-y insights. Thanks @jsolson and @boulos (it's always interesting to read your comments).
I don't begrudge the folks that prefer to work silently. I'm not adding the names of any people who worked on this (and there are many!) but perhaps don't want to be publicly visible. I assume that's why there are less AWS folks for their launches here, and it's a purely personal decision. You also might be biased since Jon and I are particularly loud and slack off a lot at work :).
Fwiw, please call us out if you think we're straying into Sales/Marketing. That's not the intent, and part of why I make sure to put the disclosure on my posts. Clearly, Cloud is a business, but I'm (still) an engineer. My goal is that we should build an excellent product, and hopefully that convinces you or others to use it. If it's not excellent, keep complaining until we improve!
Disclosure: I work on Google Cloud.
With respect to AWS, in the historical "enhanced networking" case Amazon dedicated hardware by offering SR-IOV capable NICs. SR-IOV is a well understood and effective technique for approaching bare metal performance for virtualized environments, but it tends to lock you into a particular vendor, if not specific model, of hardware. I gather ENA does something a bit different, but I don't know the details.
In Google's case, we dedicate hardware to the Andromeda switch in the form of processor cores (the "SDN" block in the linked post). This allows us to be flexible in terms of NIC hardware while presenting a uniform virtual device to guests, in addition to simplifying universal rollout of new networking features to all zones/instance types.
Both approaches have tradeoffs, although I think even with ENA AWS hits ~70µs typical round-trip-times while GCE gets down to ~40µs. Amazon's largest VMs in some families do advertise higher bandwidth than GCE does currently.
(I was the tech lead for the hypervisor side of this launch — Jake, the post's author, leads the fast-path team for the Andromeda software switch)
[ec2-user@ip-10-0-1-56 ~]$ sudo ping -f 10.0.1.111
PING 10.0.1.111 (10.0.1.111) 56(84) bytes of data.
.^C
--- 10.0.1.111 ping statistics ---
115480 packets transmitted, 115479 received, 0% packet loss, time 5385ms
rtt min/avg/max/mdev = 0.037/0.039/0.226/0.008 ms, ipg/ewma 0.046/0.040 ms [ec2-user@ip-10-0-2-191 ~]$ netperf -v 2 -H 10.0.2.52 -t TCP_RR -l 30
MIGRATED TCP REQUEST/RESPONSE TEST from 0.0.0.0 (0.0.0.0) port 0 AF_INET to 10.0.2.52 () port 0 AF_INET : first burst 0
Local /Remote
Socket Size Request Resp. Elapsed Trans.
Send Recv Size Size Time Rate
bytes Bytes bytes bytes secs. per sec
20480 87380 1 1 30.00 21178.69
20480 87380
Alignment Offset RoundTrip Trans Throughput
Local Remote Local Remote Latency Rate 10^6bits/s
Send Recv Send Recv usec/Tran per sec Outbound Inbound
8 0 0 0 47.217 21178.689 0.169 0.169It's certainly a nice improvement over what we see on the c4s. Is that using a placement group to ensure proximity (I believe our tests do, but I'd have to double check)? Our benchmarking philosophy is generally to aim for "default" numbers for GCP and "best" numbers for others -- keeps us honest about our "fresh out of the box" behavior.
Also, if we should be seeing better on earlier instance types, I'd love to know what we're potentially doing wrong.
The differential between public IPs and internal IPs is tied into the path packets take after leaving the host. The path out of the guest is identical for both, but using VM public IPs (rather than internal) can result in passing through additional hops versus being routed straight to the target VM. Common firewall configurations can also impact perf here.
Original comment:
With respect to guest CPU, the approach used by Andromeda 2.1 eliminates VM exits both on transmit and for interrupt delivery (where supported by Intel). In that regard it's essentially identical to PCIe passthrough. There are customers running DPDK to further reduce variance (and eliminate the cost of interrupt handling entirely).
The choice to not pass through host hardware comes down to a few factors, but high on the list are supporting live migration and NIC vendor flexibility.
(I worked on this effort; see other comments for specifics)
For example, here's the script we install in our guest images and encourage people to run to make sure interrupts are paired to queues: https://github.com/GoogleCloudPlatform/compute-image-package...
Disclosure: I work on Google Cloud but don't know much about networking.
I was wondering what is offloaded. I guess virtio-net is a good keyword, thanks.
In terms of specific offloads, the big ones are TCP segmentation offload (TSO) and TCP large receive offload (LRO). These substantially reduce the compute burden on the guest. Less impactful (although still important) are checksum calculation and verification offload.
(I was the tech lead for the hypervisor side of this launch — Jake, the post's author, leads the fast-path team for the Andromeda software switch)
edit: Realized you might have meant Linux VMs running outside of GCE -- the improvements here are fairly GCE-specific, although as wmf points out, vhost is a similar technology in the open source world. Performance specifics down at this level (tens of microseconds and below) tend to be hardware dependent.