- What other benefits does it give you?
- This isn't a new problem and presumably has prior best practices for mitigation. What is this replacing, what was the landscape like before this? What was the most similar already?
- Have people been looking for a better solution like this for some time?
- What's the over/under on the maintenance cost? You added in another TCP/IP stack to look after. You maybe save on static configuration and can make your system more dynamic. Pros, cons ... let's list them.
The talks a bit about problems they were trying to address, but not in a way that clearly answered the above for me. It's of course valid to write a piece with a more informed audience in mind, but in something that aims to spread the virtues of an idea I think it could do more.
What I think user-mode TCP/IP gives us is the ability to build arbitrary services --- Postgres, Redis, SSH, network management, whatever --- without having to make infrastructure changes. We don't have to have some weird API or application proxy that knows what's running and who's allowed to run what. Instead, that's simply baked into the network, and flyctl, by dint of netstack, can just use it. If somebody comes up with a cool network service to plug flyctl (or any other tool someone wants to write) into, it will just work.
But things like the maintenance cost, well, yeah, that's most of what the post is about. The maintenance cost was not especially low.†
There's a natural inclination to read any post like this as a kind of brag, but I'm really just experimenting with trying to show the good with the bad here. User-mode TCP/IP is a weird choice! Nobody else I know of does it! It might have been the wrong choice! Even though I love it!
† It's actually not low even right now; I'm spending the first half of the day deploying code that relays stats from Netlink on our gateways through our GraphQL API, so that flyctl can check WireGuard gateway health. That is not a thing we would be spending time on if we had just written an explicit proxy for Postgres or whatever, rather than providing a generic network transport.
I'm in automotive/embedded at the moment, and our daily battle is making decisions on how static vs. dynamic we want our system to be - static (e.g. resource allocations or baked-in scheduling decisions) makes it easier to reason about the system and provide guarantees, but generally lowers efficiency at runtime. Dynamic can make the system much better at serving a wide range of usage scenarios, but makes it harder to eliminate the risk of pathological cases. It's hard not to see things through that lens. The way to construct and run services you've described here to me is an interesting option on that type of axis, in the sense of where the costs/friction goes.
Sure, just how many distinct WireGuard configurations would the gateways be comfortable with per-Fly app? 1K+? 100K+? 1M+?
I keep thinking of the nightmare of keeping up with the world of Internet middleboxens, broken net layer implementations and icmp hacks that the Linux kernel supports and makes 'just work'. The jump to usermode tcp seems interesting if you're not worried about that (and I've been watching the formal-proven ip/tcp stack space like a hawk for years), but I've been burned so many times with non standard stacks and 'oh you need to connect to that non-updated lynxos system and huh' or 'hah could you enable ecn or this obscure tcp option because... Legacy?'... And sometimes I need tc/netem and netlink and I don't know...