And if any tailscale employees are reading this - https://github.com/tailscale/tailscale/issues/15724 please fix this too. Regular users not using some sort of enterprise saas DNS (whatever their thing is?) deserve DNS privacy too.
And if any tailscale employees are reading this - https://github.com/tailscale/tailscale/issues/15724 please fix this too. Regular users not using some sort of enterprise saas DNS (whatever their thing is?) deserve DNS privacy too.
That said, note that if you run your own DNS server on your tailnet, the regular UDP DNS is automatically private because it’s carried over Tailscale. That’s the most common setup for non-SaaS DNS servers. DoH doesn’t really add anything in that arrangement. (And it’s more fiddly because you need to get and refresh a TLS cert.)
So it's not that simple: it's impossible for Tailscale to use any existing kernel or accelerated WireGuard implementation. They could derive inspiration, but a kernel module for Linux won't fix Windows & Mac. With that said, I feel they have enough funding to maintain a few platforms (:
1) the WireGuard kernel implementation, despite not even being zero-copy, exceeds the performance of the userspace implementation
2) implementations utilizing the userspace network stack have a maximum potential performance (context switch + memcpy is very slow, and that affects UDP disproportionately). It's the wrong approach for meaningful improvement.
Memory copying is on the order of 100 gigabytes per second. You can do 80 full payload copys and still out-pace a dog-slow 10 gigabit per second connection. If your bottleneck is memory copying you are either doing something very wrong and doing way too many copys or congratulations you have implemented one of the fastest network stacks.
Supervisor calls are also very fast, on the order of 100 ns up to maybe 1 us with all the Spectre mitigations. Even if you did something as stupid as one supervisor call per packet, you would still be getting on the order of 10 Gbps at the long end there and 100 Gbps at the short end. Which, again, means congratulations are in order because you have implemented one of the fastest network stacks. If you do batching and add just 10 us (us, not ms) of latency then that entire cost is so small as to be irrelevant. You are going to bottleneck on your memory copying first.
Network stacks are so slow almost entirely due to poor protocol design and poor protocol implementation. Usually both.
If they would have taken that advice, tailscale instances would have been pwned by copy.fail