Would we still create Nebula today?
defined.net
defined.net
Open-source projects not-quite-prod-ready:
- WebMesh: Golang, decentralized nodes https://github.com/webmeshproj
- InnerNet: Rust, with subnet ACLs https://github.com/tonarino/innernet
- Wesher: Golang, simple mesh with pre-shared key https://github.com/costela/wesher
- Wiresmith: Rust, auto-configs clients into a mesh https://github.com/svenstaro/wiresmith
Open source projects with company-backed SaaS offerings:
- Netbird: Golang, full-fledged solution (desktop clients, DNS, SSO, STUN/TURN, etc) https://github.com/netbirdio/netbird
- Netmaker: Golang, full-fledge solution https://github.com/gravitl/netmaker
Honorable mention:
- SuperHighway84 - more of a Usenet-inspired darknet, but I love the concept + the author's personal website: https://github.com/mrusme/superhighway84 https://xn--gckvb8fzb.com/superhighway84
It's a specification for identity based networking. There is a meshnet and a centralized implementation. You can layer IPv6, IPv4, or application traffic on top of any compatible implementation.
It's the granddaddy of mesh networking, long before Wireguard, and while it's not quite zeroconf, it's very simple to setup and maintain. It also runs on everything.
And really I've been using tinc for almost a decade and I didn't really see the benefit of changing. It's rock-solid.
With the exception of one thing: I use some central nodes on cloud VPSes and they can access everything. As far as I know a nebula lighthouse can't access any of the clients. So I've been meaning to give nebula another try.
But zerotier and tailscale aren't options for me because they rely on their cloud infrastructure. I only want stuff that's fully self-hosted.
There's a great tinc android client these days too.
Unfortunately, I cannot confirm. Sharing my experience:
I used tinc over multiple years on production servers and it would sometimes create netsplits that did not recover. I also suspect that there's a race or bug in re-keying, which also causes disconnects.
On the netsplit issue, it was me posting alone on the relevant issue [1] over multiple years without response. (I don't expect to get any from free-time maintainers, especially on hard-to-reproduce issues, but it's still important to know that such unsolvable hurdles exist.)
When I switched to Nebula, it improved this situation. But both Nebula and tinc max out at around 1 Gbit/s on my Hetzner servers, thus not using most of my 10 Gbit/s connectivity. This is because they cap out at 100% of 1 CPU. The Nebula issue about that was closed due to "inactivity" [2].
I also observed that when Nebula operates at 100% CPU usage, you get lots of package loss. This causes software that expects reasonable timings on ~0.2ms links to fail (e.g. consensus software like Consul, or Ceph). This in turn led to flakiness / intermittent outages.
I had to resolve to move the big data pushing softwares like Ceph outside of the VPN to get 10 Gbit/s speed for those, and to avoid downtimes due to the packet loss.
Such software like Ceph has its own encryption, but I don't trust it, and that mistrust was recently proven right again [3].
So I'm currently looking to move the Ceph into WireGuard.
Summary: For small-data use, tinc and Nebula are fine, but if you start to push real data, they break.
[1]: https://github.com/gsliepen/tinc/issues/218
[2]: https://github.com/slackhq/nebula/issues/637
[3]: https://github.com/google/security-research/security/advisor...
My tinc usage was far less demanding. Just a handful of nodes and light traffic, so I didn't experience any issues.
I did migrate to WireGuard about a year ago, and instead of a mesh network, I ended up with a hub-and-spoke configuration[1] which worked fine for my humble needs.
[1]: https://www.procustodibus.com/blog/2020/11/wireguard-hub-and...
I've reopened the issue so you can continue to document and look for solutions. :)
I actually looked into tinc, and really wanted to love it because of its simplicity.
Unfortunately though, it seems that the development scene around it has been stalled for several years now, and benchmark reports show that it can't keep up with cloud speed tests when compared with Wireguard:
- Wireguard: 390.3 Mbps
- Netmarker: 369.3 Mbps
- Tailscale: 62.5 Mbps
- ZeroTier: 56.8 Mbps
- Nebula: 38.4 Mbps
- Tinc: 34.7 Mbps
- OpenVPN: 22.3 Mbps
Source: [disclaimer, article written by Netmaker CEO, so implicit bias]
(https://medium.com/netmaker/battle-of-the-vpns-which-one-is-...)
I found this other benchmark[1] that places them much closer to native WG performance, with Netmaker still in the lead, and somehow faster than a direct connection. This is probably due more to Netmaker using the in-kernel WG, instead of the userspace implementation. I'm kind of curious to give it a try, TBH. :)
Hell, even OpenVPN is _much_ faster than that[2,3].
Also, yeah, tinc development has been slow for years now, but I didn't experience any issues with the prerelease versions. It's certainly behind the times now, but I also enjoyed how easy it was to configure and use.
[1]: https://techoverflow.net/2022/08/19/iperf-benchmark-of-zerot...
[2]: https://www.zerotier.com/blog/benchmarking-zerotier-vs-openv...
The folks over at the Nebula team had an interesting discussion regarding the the "battle of the VPNs" article published by Netmaker I sourced in my parent comment:
Tailscale gets most of the attention on HN, and I'm sure that it's a wonderful product too, but Nebula is a nice, simple, "do one thing well" product.
Our decision to leave lighthouse hosting in the hands of users has one primary rationale: We want users to have complete control their network availability. Any downtime of our service should not impact their network availability. You can even host some of your lighthouses inside of network boundaries to ensure that an internal network functions properly if its connection to the internet is interrupted. Other overlay options may continue to work for some time, but new connections are often not possible, and the network can degrade rapidly.
Relays are are a similar story, but with an additional reason: We don't have to limit our customers' relay bandwidth due to cost. When hosting relays on behalf of others, we would be transiting a lot of traffic, which has an associated (sometimes unpredictable) cost. By letting our customers host relays, they can ensure relay traffic is just as fast as direclt connections.
Having a mesh means only the lighthouses need fixed IPs, and then everything can talk directly with no extra work. That is significantly more scalable, adding a new client requires nothing from any of the other clients. And it cuts out unnecessary legs in the trip and the traffic and latency that involves.
Wireguard is awesome and it sounds like you probably don't need to bother in your use case, but if you start providing distributed services not funneling through hubs starts to be worth considering. Particularly if you start caring about one service site going down affecting others.
I certainly have my gripes about the closed nature of Slack itself, in particular using a closed protocol when the model is clearly "federated" between multiple servers internally. That said, the contribution of something on the scale and quality of Nebula back to the open source community is hard to argue with.
[0]: https://github.com/anderspitman/awesome-tunneling#overlay-ne...
They added in tag support [1] a few months ago which I have yet to try out but it looks very promising. The defined.net API [2] is very easy to use for host management and I am able to auto enroll new hosts and remove them after I deprovision them.
I also made a GitHub Action [3] which I use to allow for my Actions to communicate with resources on my overlay network.
[1] https://docs.defined.net/guides/creating-firewalls-using-rol...
Thanks for sharing this on HN! I'll keep an eye on the comments and try to answer questions that come up.
Do you thing the tech landscape today would have allowed for nebula to be born? Lots of companies now have strict IP agreements they have team members sign.
The most important concern (IMO), was considering whether we could commit to properly maintaining a project. Before open sourcing anything, you need to discuss how you'll go about managing an issue and pull request backlog, so that people don't come across "dead" projects under your stewardship.
In a high growth startup, I do think something like this could happen again, but as a company grows, there are certainly more layers that can make it difficult to share things openly.
I don’t really think it’s the size or layers of a company that prevent; it’s the culture. This culture of creation permeates everything I’ve seen Stewart Butterfield do. At least from the outside. Admirable and extremely profitable.
I'll be honest: If I could do it again, I'd use Nebula. The primary issues I have are that Tailscale has a lot of magic which I can see some cases it being nice, but it does make some of the routing and firewalling I'm doing on machines, and in particular the thing where it sets up Tailscale routes to network routes as higher priority than local interfaces leads to problems in my environment.
The other thing is just Headscale itself, it works quite well but does have some rough edges. It's entirely too easy to kill your whole mesh by flubbing an ACL, and currently restarting headscale to pick up ACL changes is taking 3-5 minutes.
I do, however, really prefer the Tailscale ACLs over Nebula's.
One thing that led me to Tailscale was the ability for it to relay around network routing problems, and it looks like Nebula has added that since I started. Around the time I was evaluating Nebula vs. Tailscale we had a ~1 day network routing issue where some of my users were blackhole routed in Comcast, and Tailscale just worked around it.
After having searched (and implemented) this myself for work, the only practical solutions I found were 1) smallstep [1] or 2) Terraform (with the nebula provider [2]) and a CM tool of your choice. The latter can be nicely combined with the ansible provider if that's your CM of choice.
0: nebula-cert-py 1: https://smallstep.com/docs/step-ca/integrations/#nebula 2: https://registry.terraform.io/providers/TelkomIndonesia/nebu...