Most people do not use TailScale. I'd encourage doing the work of understanding why, there is likely a big opportunity somewhere here.
When you use Tailscale extensively, it becomes your new network. Now all your systems depend on a piece of software that you do not fully control. The control plane is not open source, and it is a key component of Tailscale. Headscale is a great effort, but it doesn't have feature parity with Tailscale (1). Moreover, the dedicated team at Tailscale keeps releasing amazing new functionality regularly.
That being said, if I had to buy software from a company, Tailscale would be my first pick. I respect and trust the founders and the early engineers working there.
As a side note, I'm planning to contribute to Headscale. This technology is crucial, and I want to help ensure its success.
(1) The functionality offered by Headscale is sufficient to build a robust mesh network and enjoy its benefits. Kudos to the team and to Tailscale for supporting it.
An attacker with access to coordination or relay servers would be able to change whatever is in the admin console, which are basically ACLs.
Am I missing anything?
1: https://github.com/DefGuard/defguard discussed at https://news.ycombinator.com/item?id=36056080
This is why I asked, the phrase "I decided to reinvent the wheel which has honestly been quite fun with learning about eBPF, and recently clustering and HA with etcd" makes it sound like it's doing a bunch of cool stuff (which I want to hear about!), but the readme says nothing about those.
Primarily because I want people to have a reasonably good time setting it up, rather than having to go through my explanation on things!
I'd be super interested to know how they track "session state" as their do seem to rely very heavily on adding proxies and other additional software layers in front of the wireguard connection itself (https://defguard.gitbook.io/defguard/admin-and-features/wire...)
With wag specifically it's all just wireguard and a tiny bit of ebpf to do the management, along with tracking the external IP to determine if its time to re-challenge a user.
I didn’t take it as malicious, but trying to understand more about this method. I’d love for the author to tell us a bit more about how it works. I’m curious about what obstacles the author hit and how they got around them.
Note: re: the flagged sibling comment. Yeah, that one doesn’t get the benefit of the doubt and was out of bounds.
How it works:
In short, Wag adds an eBPF program to a WireGuard device that it instantiates. The eBPF program uses a number of hash maps and LPM (longest prefix matching trie) maps to determine the policies that are applied traffic coming in on the wireguard device. These policies based on the ACLs defined per user/group, and contain MFA/Allow/Deny rules which require mfa, allow without auth and deny always respectively.
Wag also watches all the wireguard peers ingress IP addresses, and when an address changes it deauthenticates the user and requires the user to complete a login challenge. This is done by basically setting a bit in the maps exposed to eBPF that says "unauthorised"
Challenges:
First and foremost with WireGuard there is no good way of determining if an "external ip" i.e where the user is connecting from has changed. There was a patch set submitted for review in 2022~ that was never actually added to the kernel that would have added netlink compatibility and thus event based notification that things had changed, but alas that was never reviewed by Jason Donenfeld and has quietly died the death.
Secondly was defining multiple policies per route was quite difficult as eBPF doesnt do dynamic memory even in userland exposed maps and I wanted multiple rules per route, i.e you might allow port 80/tcp when MFA has passed but otherwise always allow 22/tcp. So to do that I had to define a maximum number of rules that could be inserted as one memory blob into the LPM map that the ebpf program would then linearly search to make its decision.
Thirdly has been making everything highly available which has been a bit of an on-going battle with ETCd mainly around how it manages cluster certificates as they dont (as of 2024 but it may be coming soon) expose the right structures to allow for dyanmic certificate creation, so you have to kind of make a wrapper around that in order to get everything going.
Im sure there are other things that I've had struggles with, but these are what come to mind immediately!
Best of luck with the project!
Like the time I had to optimise map insertion because the linux kernel does some truly insane locking when you use specific types of eBPF maps:
https://github.com/NHAS/wag/issues/84
This is slated to be improved (or has already been improved in kernel 6.8?). But for now wag sort of just side steps it in a horribly stateful way.
On the patch, maybe try reposting it on the list, with the pointer to your project to see if that provokes a new review?
I just didn’t get that vibe from the now flagged comment by @aragilar. I saw it as a genuine curiosity about the design choices. Maybe I was wrong.
“¯\_(ツ)_/¯“