Nebula, Slack's Open Source Global Overlay Network
slack.engineering
slack.engineering
I wonder if they tried ZeroTier. It sounds really like what they wanted.
To contrast, with Nebula you run your own root(s) (lighthouses) and you don't need a controller because important config (ip, group, hostname) is signed by the same CA.
https://www.zerotier.com/zerotier-2-0-status/
https://www.zerotier.com/lf-announcement/
The moon terminology will also go away since there will no longer be a difference between these and our core roots. They'll all just be roots and will be interchangeable. The use of a common underlying key/value store will allow ZeroTier to keep its unified namespace and easy ability to join anyone's network or communicate with anyone regardless of what roots they're using (as long as their roots are on the same global network as you... obviously you can't hop air gaps).
As for crypto: we plan some improvements in 2.0, but note that these days easily >90% of the traffic over ZeroTier networks (or any other VPN / overlay) tends to be already encrypted via SSH, SSL, etc. Another layer of encryption in the overlay just provides some additional defense in depth. We're rapidly moving to a world where everything layer 3 and above encrypts everything. That's why we have not prioritized sexier crypto for our L2 overlay tech. It's a bit redundant.
It feels a little different from Wireguard, in that with Wireguard your engineers would be able to connect from behind a NAT, but my reading of how it works is that machines route directly at each other. Which is good for a production network where you care deeply about routing (bandwidth, latency, costs, debugging, etc.), but it seems that here your engineers would still need to connect to a bastion host or something, i.e., it isn't a VPN in the sense of being able to join the corporate network directly.
I guess if you've also got the lighthouse node internally routable by all your machines (e.g. you have an internal datacenter network and something like AWS Direct Connect) it would work too?
It'd be nice to see a sample network design.
I think the answer is that your lighthouse(s) are the only machines that need publicly routable IPs. Your ephemeral cloud machines get any RFC-1918 address you want, with any subnet you want.
Engineers would have Nebula set up on their laptop with a configuration that knows about your lighthouse(s) static IP(s). They use the lighthouses for meeting other nodes, UDP hole punching, etc, but otherwise every connection is peer to peer.
(Which is why I suspect it's not and the readme isn't clear)
Both for overlay network[1], and/or for nodes?
Also, the end-to-end principle argues for putting complicated logic in the endpoints and making the network boring. See also, TCP is implemented at the endpoints and just requires network infrastructure to drop packets sometimes. You could imagine a congestion control protocol implemented on each router on the Internet, but it would be much more fragile and also much harder to deploy changes to.
[EDIT] MIT, that is, of course.
https://www.zerotier.com/on-the-gpl-to-bsl-transition/
I don't think the BSL is perfect. We're thinking and discussing with a number of people about potentially better licenses that would be closer to traditional FOSS while preventing "SaaSification" and similar. I think we're in the early stages of a renegotiation of the open source social contract and I don't think we've figured out the best model yet.
The AGPL is close but suffers from two problems: (1) it isn't perfect either and has numerous loopholes, and (2) there are a ton of companies out there with an irrational but nevertheless very entrenched phobia of anything associated with the GPL (as we have discovered). Maybe something a bit like the AGPL but not GPL branded would work.
BTW the closed source restriction in the BSL is effectively the same as the GPL and the only other meaningful restriction is on SaaS direct monetization. Companies can still run ZT for free and run it behind the scenes for free. It's a lot like the AGPL.
A SaaS company can get a commercial license.
https://github.com/zerotier/ZeroTierOne/blob/master/LICENSE....
> Business Source License 1.1
> License text copyright (c) 2017 MariaDB Corporation Ab, All Rights Reserved.
> "Business Source License" is a trademark of MariaDB Corporation Ab.
Which is a bit odd, as MariaDB itself is gpl2 (+commercial)?
I'd like to understand how Nebula works -- any other suggestions besides just diving into the code?
VPNs are primarily used for remote access, to get random machines access to closed IP networks. Service meshes synthesize a new network (sometimes IP, sometimes something else) to connect a bunch of related machines, almost always with policy controls for who can talk to what, usually cryptographic.
It would be weird (but not "wrong") to use a service mesh to get developer laptops access to staging Postgres.
It would be weird (but not "wrong") to use WireGuard to connect an application server to its Postgres instance.
WireGuard is a much tighter and more limited design, intended for integration directly into operating system kernels, with a strong emphasis on performance. Nebula is a much more ambitious design; it includes direct DNS support, certificates, and server infrastructure. WireGuard is a few thousand lines of very carefully written C code; Nebula is a typical Go project.
They're both very cool.
(Backdrop: I have recently moved our various prod servers into a WireGuard based VPN to encrypt the traffic between them. I found it was easier/pragmatic to do this than:
* to setup SSL for my DB
* to figure out how to encrypt traffic between my application server and Redis or my application server and Nginx )
[1] https://www.noiseprotocol.org/noise.html#introduction [2] https://github.com/slackhq/nebula#3-a-nebula-certificate-aut...
In addition to VPN, Nebula added traffic filtering and spanning different clouds and data centers. I don't think Wireguard had those as goals.
They serve very different purposes. I use WireGuard to encrypt my mobile traffic but I wouldn't have picked it to connect the various hosts in my network at work. Nebula, however, might do the trick.
I used to use tinc and have recently switched to yggdrasil, which was much easier to setup. So far it works great!
Nor per connection configuration, you can eg let any nodes that hold certs from your CA communucate.
But they want a global VPN for _everything_ including laptops. This means some level of access control.
What I like here is the use of lighhouses, to allow external nodes to punch in and discover the rest of the network. Something which is very difficult to do if you are relying on a service mesh in an unknown and unconnectable network.
The thing that immediately stands out is the routing. It looks like cjdns is a traditional-ish multi-hop network. The DHT routing table allows you to map out a route to peer A via peer B, R, & D.
What wireguard and nebula allow is for the underlying network to figure out most of the routing, and effectively create a massive point to-point network. whilst you can have concentrators/gateways, the idea is that most of the traffic goes direct from peer to peer. This can reduce load considerably.
I know there are ways to make this happen (e.g. using the techniques from Samy Kamkar's pwnat/chownat), but am not sure whether Nebula is designed to work within this constraint.
pwnat is notable because it doesn't require having a public stun-like server, but nebula already assumes there's public servers, so traversing nat is a non-issue.
The readme says "Discovery nodes allow individual peers to find each other and optionally use UDP hole punching to establish connections from behind most firewalls or NATs".
In practice, I didn't see any code that implements it, but I didn't look too hard.
To my understanding, a service mesh does not establish a common VPN-like network, but assumes it's there already. Nebula and service meshes both provide authentication, end-to-end encryption and role-based access control. A service mesh can do more than Nebula: it makes it possible to shift traffic between services for example apart from a "security group"-like filtering.
However, I might be mistaken. Any corrections are more than welcome.
That being said the code is full of TODO and other comments indicating that shortcuts were taken which should be fixed later. I would be worried about running such a thing in prod given the criticality of its function. At best you could risk performance issues under load and at worst you could have significant security issues allowing unintended traffic in/out.