Introducing the Fan – simpler container networking
markshuttleworth.com
markshuttleworth.com
So you ping 10.3.4.16 and your host automatically 'knows' to just send it to 17.16.4.16 where lying in wait, the receiving host simply forwards it to 10.3.4.16. I like it.
This is a vexing problem for containers and even VM networking. If they are in a NAT you need to create a mesh of tunnels across hosts, or you create a flat network so they are all on the same subnet. But you can't do this for containers on the cloud with a single IP and limited control of the networking layer.
Current solutions include L2 overlays, L3 overlays, a big mishmash of GRE and other type of tunnels, or VXLAN multicast unavailable in most cloud networks, or proprietary unicast implementations. It's a big hassle.
Ubuntu have taken a simple approach, no per node database to maintain state and uses commonly used networking tools. And more importantly it seems fast. And it's here and now. That 6gbps suggests this does not compromise performance like a lot of other solutions tend to do. It won't solve all multi-host container networking use cases but will address many.
Why isn't DHCP used? What does it not do that this service does?
This can even be done on the command line using iproute2 utilities: https://www.kernel.org/doc/Documentation/networking/vxlan.tx...
Though you should probably use netlink to do it programatically. Personally I like to combine netlink + Zookeeper or similar to trigger edge updates via watches.
I remember seeing a patch floating around that added support for multiple default destinations in VXLAN unicast but I think some objections were raised and it's not made it through. At least it's not there in 4.1-rc7. That would be quite nice to have.
http://www.spinics.net/lists/netdev/msg238046.html#.VNs3rIdZ...
Right now I am using netlink to manage FDB entries, last time I tried iproute2 utility worked too...
The only tricky thing about doing it with netlink is that the FDB uses the same API as the ARP table. Specifically RRTM_NEWNEIGH/RTM_DELNEIGH/RTM_GETNEIGH, apart from that it's pretty simple though.
We have developed re6stnet in 2012. You can use it to create an ipv6 network on top of an existing ipv4 network. It's open source and we are using it ourselves internally and in client implementations since then.
I wrote a quick blogpost on it: http://www.nexedi.com/blog/blog-re6stnet.ipv6.since.2012
The repo is here in case anyone is interested: http://git.erp5.org/gitweb/re6stnet.git/tree/refs/heads/mast...
> Well, most importantly we have stable IPv6 everywhere - including on IPv4 legacy networks.
I'm surprised that is still interesting/noteworthy.
HN Readers discount, since we're on the subject. Just email.
> rsync.net does not appear to be IPv6 capable at all :-( 0 out of 5 stars.
We're a cloud storage provider. There's not much for you to do on our website.
> More importantly, we can route to these addresses much more simply, with a single route to the “fan” network on each host, instead of the maze of twisty network tunnels you might have seen with other overlays.
Maybe I haven't seen the other overlays (they mention flannel), but how does this not become a series of twisty network tunnels? Except now you have to manually add addresses (static IPv4 addresses!) of the hosts in the route table? I see this as a huge step backwards... now you have to maintain address space routes amongst a bunch of container hosts?
Also, they mention having up to 1000s of containers on laptops, but then their solution scales only to 250 before you need to setup another route + multi-homed IP? Or wipe out entire /8s?
> If you decide you don’t need to communicate with one of these network blocks, you can use it instead of the 10.0.0.0/8 block used in this document. For instance, you might be willing to give up access to Ford Motor Company (19.0.0.0/8) or Halliburton (34.0.0.0/8). The Future Use range (240.0.0.0/8 through 255.0.0.0/8) is a particularly good set of IP addresses you might use, because most routers won't route it; however, some OSes, such as Windows, won't use it. (from https://wiki.ubuntu.com/FanNetworking)
Why are they reusing IP address space marked 'not to be used?' Surely there will be some router, firewall, or switch that will drop those packets arbitrarily, resulting in very-hard-to-debug errors.
--
This problem is already solved with IPv6. Please, if you have this problem, look into using IPv6. This article has plenty of ways to solve this problem using IPv6:
https://docs.docker.com/articles/networking/
If your provider doesn't support IPv6, please try to use a tunnel provider to get your very own IPv6 address space.
like https://tunnelbroker.net/
Spend the time to learn IPv6, you won't regret it 5-10 years down the road...
There are lots of overlay networks, are those all hacks too?
>Except now you have to manually add addresses (static IPv4 addresses!) of the hosts in the route table?
This does not appear to be true at all, based on the configuration that's posted at the bottom of the article.
Even so, there are are lots of ways to get IPv6 now, I would think anywhere where you could use this fan solution to change firewall settings and route tables on the host, you could also setup an IPv6 tunnel or address space. Even with some workarounds for not having a whole routed subnet, like using Proxy NDP.
It seems like a much more future-proof solution than working with something like this. Just my 2c...
> Also, IPv6 is nowehre to be seen on the clouds, so addresses are more scarce than they need to be in the first place.
There are a lot of people that are using AWS and not in control of the entire network. If they were in control of the entire network, they could just assign a ton of internal IP space to each host. IPV6 is great, sure, but if its not on the table its not on the table.
We will be testing the fan mechanism very soon, and it will likely be used as part of any LXC/Docker deploy, if we ever get to deploying them in production.
> Additionally, VPCs currently cannot be addressed from IPv6 IP address ranges.
http://aws.amazon.com/vpc/faqs/
And then you still have the problem of only so many IPs per host, so it doesn't help with lots of containers.
One small example: How do you implement a IPv6 firewall which keeps all of China and Russia out of your network? (My apologies to folks living in China and Russia, I've just seen a lot of viable reasons to do this in the past).
Another small example: How do you enable "tcp_tw_recycle" or "tcp_tw_reuse" for IPv6 in Ubuntu?
http://docs.aws.amazon.com/ElasticLoadBalancing/latest/Devel...
Crazy, right? Especially since new customers are forced to use VPC and don't even have the option of falling back to EC2-Classic.
It also makes routing within your VPC so much more entertaining to manage.
You do this by blocking the IP ranges that are assigned to China and Russia. Same as you would with IPv4, why would that change?
Also, tcp_tw_recycle when set for IPv4 also applies for IPv6, despite the name...
I wasn't aware that AWS still doesn't have support for IPv6, that's just amazingly bad in 2015. I'll shift my blame onto them then for spawning all these crazy workarounds.
Presumably the packets will be encapsulated so that should not be an issue.
As someone who used to manage systems in 34.0.0.0/8 and 134.132.0.0/16, oh god I hope nobody does this...
Yeah, how pragmatic of them. Instead of for pie in the sky, let's all get together to pressure people to improve tons of infrastructure we don't own, action (which can always happen in parallel anyway) they solved their real problem NOW.
>and develop a whole system that will very likely break things in random ways.
Citation needed else it's just FUD. The links explained how it works well enough for that.
Also, it's worth noting that IPV6 is nowhere near as battle hardened as IPV4; there's too many optimization and security gaps to depend on it in production. I've watched a few network gurus burn themselves out attempting to harden a corporate network against IPV6 attacks while keeping it usable.
For some values of "we". If you're stuck on EC2, yeah, you've got a problem.
We've[1] had ipv6 addressable cloud storage since 2006.
Currently our US (Denver), Hong Kong (tsuen kwan o) and Zurich locations have working ipv6 addresses.
[1] You know who we are.
It feels similar to vxlan, only without the centralized repository of IP -> host mappings.
* cross-host container networking
* no overlay IP database to sync and maintain
* only a single IP used per host
They still use encapsulation, like other network overlay technologies, it's just that by using a specific addressing scheme they can eliminate a lot of cross-host communication and all the database lookups. *Remap 50 addresses from one range to another.
*Dynamically assign those addresses to servers.
*Special Something that Fan does.
What's the benefit of using a full class A subnet when you are only using 250 addresses?So...Each physical host has its software networking pass through a NAT before hitting the physical adapter? And it's already using DHCP to assign addresses to its containers?
> But how does it route to another container on some other host elsewhere in your cloud? That's what the Fan addresses.
You use DNS to create a lookup table matching container to IP? Since this isn't being done, it must not work. What does Fan do instead?
Is there some other complicating factor to which I'm ignorant? Are we talking about having multiple Kubernetes clusters inside containers inside VMs inside a physical host?
Also, HOW does Fan address this problem? What does it use instead of one kind of database lookup or another, like a distributed database system such as DNS? Or does Fan use fancy subnet math?
I agree with you from a purity point of view that proper service-discovery is better, but in terms of practicality, IP per container is often a lot simpler to implement.
The 1.6 docker solution involved a double nat and relied on iptables - resulting in some fairly serious bottlenecks and pathalogical edge cases. It also required a third party solution for handling discovery of IPs for services.
Opening up containers to access the host network interfaces breaks the encapsulation promises of containers, and is thus not available to all people. Conceptually, it also creates holes in the idempotent service model, since they have to be aware of port conflicts.
The 1 IP per container model, such as VXLan, Flannel, Kubernetes, Docker 1.7 etc. is one of the more effective methods of countering the problem, at the cost of guzzling IP address space, and requiring a gateway to escape the virtual network tunnel.
Who made that promise? It was never a feature of Linux containers to virtualize the NIC.