Google shares software network load balancer design powering GCP networking
cloudplatform.googleblog.com
cloudplatform.googleblog.com
As an aside, I couldn't help noticing these lines on adjacent pages:
> Maglev has been serving Google’s traffic since 2008. It has sustained the rapid global growth of Google services, and it also provides network load balancing for Google Cloud Platform.
> Maglev handles both IPv4 and IPv6 traffic, and all the discussion below applies equally to both.
In that case, any chance GCE will be getting IPv6 support soon? ;)
[1] https://cloudplatform.googleblog.com/2014/11/cloudsql-instan...
[2] https://code.google.com/p/google-compute-engine/issues/detai...
Internal use (2011) http://www.pcworld.com/article/245848/usenix_google_deploys_...
Internal and some datacentres (2009) http://www.networkworld.com/article/2238936/lan-wan/google-a...
Backbone dual stack (2010) https://www.nanog.org/meetings/nanog49/presentations/IPv6atG...
Facebook internal ipv6 http://www.internetsociety.org/deploy360/wp-content/uploads/...
And not just the ones that that backend was handling, but also some % of the overall traffic for re-balancing.
The degree of connection affinity from the ECMP is limited, and there's no "reliance" on it. If a connection flip-flopped between two or more load balancers, there would be no drops, thanks to the consistent fashion.
Exactly how the routes get steered to servers can vary (static routes with track objects, OSPF or BGP on the server), but the basic idea of using network gear to ECMP traffic between servers is, if not ancient, pretty common.
I suspect there's something novel in control plane or hashing used to steer ECMP but I didn't see anything about that in the article.
Well, it's "not new" in the sense that this system's been running google.com for about 8 years now ;)
This is just the first time that we've published how it works.
You still lose some connections when an LB fails, but only the ones going through the failed LB. Unrelated flows to other LBs are not impacted.
We've got multiple customers doing variants of the LB architecture Google talks about here.
Resilient hashing means that only flows going to the dead LB will get rehashed to other LBs. Those flows would break anyway, but the remaining flows are OK.
in such cases, only consistent hashing or maglev hashing might be the only option.
Would love to see the HN community's opinion on this, as someone who is not an expert in ops/infrastructure.
This is probably one of he most compelling arguments I've heard: "It will work well because we both rely on it working well. We have skin in the game."
I love the jab at AWS in the end:
> 'knowing that when your traffic ramps up, we’ve got your back.'
They've heard that AWS load balancers need "prewarming" to handle huge traffic spikes, so they're throwing one in there...
[1] http://research.google.com/pubs/pub44824.html [2] https://news.ycombinator.com/item?id=10287318
> Because the NIC is no longer the bottleneck, this figure shows the upper bound of Maglev throughput with the current hardware, which is slightly higher than 15Mpps. In fact, the bottleneck here is the Maglev steering mod- ule, which will be our focus of optimization when we switch to 40Gbps NICs in the future
Sounds like they could use RPS
On the understanding that I can't tell you any more than what it already says in the paper, the key detail that you're looking for is section 4.1.2:
""" Maglev originally used the Linux kernel network stack for packet processing. It had to interact with the NIC using kernel sockets, which brought significant overhead to packet processing including hardware and software interrupts, context switches and system calls [26]. Each packet also had to be copied from kernel to userspace and back again, which incurred additional overhead. Maglev does not require a TCP/IP stack, but only needs to find a proper backend for each packet and encapsulate it using GRE. Therefore we lost no functionality and greatly improved performance when we introduced the kernel bypass mechanism – the throughput of each Maglev machine is improved by more than a factor of five. """
My understanding is that RPS just steers packets into the Linux TCP stack. Nothing on the maglev machine wants to process the TCP layer; that workload can be distributed over the backends.
See netmap, DPDK, for other examples, or look at https://www.cl.cam.ac.uk/research/security/ctsrd/pdfs/201408...
In our case the kernel buys us very little but costs a lot, as others correctly pointed out in this thread. We spent a lot of time working on our performance via the kernel, with moderate success, but userspace talking directly to the NIC was a huge leap in performance.
Is "hardware encapsulator" just a fancy way of saying they're using tunnel interfaces on the routers, or is a "hardware encapsulator" an actual thing?
No.
Edit: expanding on this a little, it's not something that's been released so we can't talk about it. I don't think I can comment on "illumin8"s proposals other than to say that I'm pretty sure they don't work here.
This is the standard way of distributing traffic in large datacenters. That way you get extremely fast, non-blocking line rate between any two physical hosts in the datacenter, and since the physical hosts know which VMs/containers are running on them, they can pass the traffic directly to the other host if VMs exist in the same L2 network, and even do virtual routing if the VMs exist across L3 boundaries - still a single east/west hop.
In a regular IP Fabric environment we would all this device a VTEP.
They're doing basically what Google is, but with off the shelf hardware and a openly-buyable NOS.
How awesome is that?
We used this strategy at a previous company I worked at deploying large Openstack based clouds, worked really well. Though if you want to to be easy on the backends you need to use some sort of stick table or consistent hashing mechanism as described in the paper.
10gbps with 8 cores is pretty crappy performance. People using DPDK and netmap are getting 10gbps with a single core. Netmap can bridge at line rate for 64 byte packets with a single core.
I am actually surprised this is 'news'. Perhaps google should reach out to some of the DPI companies to figure out how to really scale this. Currently they are 'overpaying' on commodity hardware by at least 8x over the current cutting edge in this type of field.
Just have a look at the NFV world for the cutting edge packet forwarding rates
And I'm sure you're aware, bridging is a totally different use case.