Our User-Mode WireGuard Year
fly.io
fly.io
Modified SLIRP code is also found in VirtualBox, Qemu, UML and other virtualization software, for sharing the host connection in NAT mode.
Maybe it's just rosy nostalgia.
I thought that's just how you got into the interwebs without running that obnoxious, buggy faux winsock client.
My solution was to write a program (I called checkcpu) that would spawn a process (slirp) and periodically check its total cpu usage. When it hit the threshold (110 seconds), it would spawn a child and suspend the parent (seamlessly passing the current run state to the child). It worked great and they either never noticed what I was doing, or they did not care. Over time, the number of suspended parent processes would rise, but it never became a problem.
I'd have to dig into the details more, but something like this might allow you to implement a simple tunneling system based on WireGuard that runs in the kernel if you have the privileges, otherwise falls back to usermode and is no worse than QUIC in terms of performance. That would be awesome.
2. Tailscale is user-mode WireGuard.
3. "User-mode WireGuard" in the sense this post uses the term is a misnomer and refers to the fact that we run TCP/IP itself in userland (Tailscale normally runs through a tunnel device and uses your native TCP/IP stack).
4. But Tailscale also has code to do user-mode TCP/IP (they've got it running in a browser with wasm).
Running on wasm sounds awesome. This[1] looks like it. Do you know how they're doing the actual networking? WebRTC tunnel?
[0]: https://news.ycombinator.com/item?id=24483173
[1]: https://twitter.com/bradfitz/status/1451423386777751561?lang...
Tailscale's gvisor/netstack-based userspace networking mode has been supported and in wide use for quite some time. It's the default on Synology DSM7, for instance.
You don't need root when you run tailscaled with `--tun=userspace-networking`.
Peers can still connect inbound to the non-root tailscaled, but to connect _out_ to other peers, you need to use tailscaled's HTTP or SOCKS5 proxy, which are also flags to tailscaled, to specify what port they listen on.
Do you have any links that talk more about how the wasm stuff works? I'd love to read more about that.
(iOS 15 bumped the Network Extension memory limit to 50 MB, but we still need to be super trim for iOS 14's 15 MB limit)
Is there actually a preference for user-mode networking? I assume that’s primarily about control and flexibility?
Either way, I hope that the PacketBuffer changes can help reduce footprint after issues are shaken out.
You can use Tailscale here but you'll need to separately run a reverse-proxy on a public machine. There are more moving pieces but if you're using Tailscale already then it's a good option.
It looks like Fly.io has all the bits, they just need to be packaged as a stand-alone tool rather than built into flyctl and only talk SSH.
tailscaled --tun=userspace-networking --socks5-server=localhost:1081I should try to recreate the logs and issue for the tailscale folks.
An example of using Tailscale to access VSCode in the cloud: https://render.com/blog/host-a-dev-environment-on-render-wit...
Found a gap, Linux Foundation's FD.io's VPP (a high performance network virtual switch) has native wireguard support as well, all in userspace. Support here means you can do full kernel bypass from the app all the way down to the NIC card (e.g. via DPDK).
https://docs.fd.io/vpp/20.09/d5/d54/wireguard_plugin_doc.htm...
I'll open a PR on this later.
Not sure it fits your "mainstream" qualification, but many projects ago I used Airhook to help create a userspace, application-layer('ish), multi-channel virtual network: http://airhook.ofb.net/ (https://github.com/egnor/airhook)
Airhook is a relatively low-level library that handles framing and flow control; it's not a functional solution on its own. But that seems to be what you're getting at--something you can deeply integrate into your application, not a separate service. Though, I guess containers have sort of muddied that distinction.
There isn't kernel driver like in real WireGuard.
I am both impressed and horrified, and I mean that in the best possible way.
This sentence style sounded familiar in my head. Looking at the username. Oh, of course, that's him!
(for those who don't know, http://trout.me.uk/mstcat.jpg )
In her post she doesn’t mention fly.io’s motivation for doing userland TCP-IP: a nicer end user client experience.
If I read this correctly then fly.io did all this work to make their CLI user experience markedly better. That’s pretty cool. Twist yourself in interesting knots to make your user’s lives better, maybe even in ways they won’t notice!
Not sure what happened there. The processor seems to score less than an RPi4 on Geekbench.
The apu2 is an embedded AMD quad core 1GHz SoC consuming 5W. It is not a powerful system by any means and not surprised rivaled by a >1GHz quad core Arm.
I wish that more companies could be like this and skip the corporate BS, it shows that they really have something outstanding to offer.
(And happens to be a top HN contributor)
The nature of the blog typically cater towards the intended audience.
The CIO of Disney doesn't give a sh*t if the protocol is called WireGuard or OpenVPN or that if it uses AES-256 encryption - he/she wants someone to tell them that their developers are securely accessing their infrastructure. Full stop. If/when Fly gets to that level (let's say $500M in revenue) their blog tone will likely change - their audience is almost primarily developers and startup CTOs...for now.
Whitepapers. They want whitepapers and magic quadrants.
> Whitepapers.
I probably should have said "content" instead of blog because I agree 100% with this. Point still stands.
What the OP was referring to will likely become part of an engineering blog. a la : https://codeascraft.com/
I kind of miss the IBM ITSO Redbooks. No idea if they still maintain the same quality today, but in the 90s and pre-internet/google/wiki etc, they were fonts of deep knowledge.
Pro:
+ Can just run "native" SSH directly over it (or, in our case, use x/crypto/ssh, without modification).
+ Lets `flyctl` offers a `flyctl proxy` command to users, so they can plug their own programs into whatever application they need to use, without asking us to change some proxy we run in our infrastructure.
+ Offers a single security and access control model (IPv6 private networking), rather than something we have to think about on a per-app basis.
+ In theory, we get all this right and never have to think about another network protocol in our infrastructure.
+ Allows existing network management tools and libraries to function directly with Fly.io infrastructure.
+ With the WebSockets gateway, we can do all of this stuff directly in browsers as well; that is, we can present TCP/IP as an API to browser Javascript to do UI stuff (and we're doing more and more UI stuff these days, in Elixir.)
+ Puts more IPv6 in the world.
+ Get to talk to Jason Donenfeld more.
+ Get to write blog posts like this.
Con:
- Spends one (maybe multiple) innovation tokens or whatever you want to call them.
- Way more things can go wrong; relies on state synchronization and on a clear network path between our users and our gateways. Right now, we have to care whether you can speak 51820/udp.
- User-mode TCP/IP via Netstack is probably significantly slower than a simple TCP proxy would be.
- Required `flyctl` to run a background agent process to manage multiple connections through WireGuard.
- The agent process adds to the list of things that can go wrong (hopefully we're ironed most of them out now).
I can probably come up with more cons.
...
It's buried in the middle of the post but I want to say it again because I think it's important: this sort of started out as a stunt; it's what I put together to allow people to SSH into their instances without having to install WireGuard locally, and that's all it was. I don't have to write a soul-searching pro/con list on stunts I use to give people SSH access, because lots of providers have super janky "pop a private terminal" setups. But all this stuff took on greater importance when we used it to run Docker for our remote builders.
I like the approach we're taking a lot! I don't... regret it? I don't think? I think I'm happy with it. But it's complicated.
The experience of...
1) build a thing because it's immediately useful for a specific use-case, 2) someone reuses it for another use-case because it's already there and saves some work, 3) times passes 4) oops really important stuff now relies on this thing in ways that weren't originally intended
... seems like a common pattern (see: JWT succeeding as the format for interoperable tokens by dint of just being around).
In this case, it seems like the pros are basically user-centered pros (`flyctl proxy`, existing tool interop, etc.) and the cons are basically fly-centered cons (state synchronization, maintaining the agent and making it work right).
The cons that do affect users (slowness, maybe they can't speak 51820/udp) seem _annoying_ but not deal-breaking for a lot of use cases. If the slowness persists over a long time it will be interesting to see how users opt to route around it (architect applications / processes to not rely on this channel).
While on the topic, we're also eagerly awaiting improved autoscaling (e.g., more responsive, using additional metrics, and scaling down properly). I'd be really curious if you could leverage the more detailed access to instance-level metrics to implement some cool new queue-theoretic modeling: You know roughly how long it takes for an app to launch, you know the current request rate, and you know the time to service requests. You could apply a lightweight Markov model to predict the probability of a given queueing delay in each region within the average launch time and, if so, preemptively launch a new instance before queueing delays even occur. This could be configured to balance a client's tolerance for queueing delays with over-provisioning budget.
If you like, you can roll a Dockerfile that runs OpenSSH directly on your internal network address (bind it to `fly-local-6pn`), and then use native WireGuard to talk to it.
I've got a branch on hallpass that does port forwarding, but I never merge it, because you're right: using port forwarding on Fly.io is weird, because we already provide you direct access to any port you're exposing, and you can't talk to any of this stuff without WireGuard already. I think it would just confuse people more if I made port forwarding work.
We're working on a replacement for Nomad (called flyd) that gives us full control over VMs. Once apps are running on that we can do a lot of cool things. Better autoscaling is one, but I'm really excited about suspending idle VMs that our proxy wakes up on demand. That'll cover most use cases without forcing customers to worry about counts or blowing through a budget.
We haven’t had too good a time with nomad, but not sure if it’s just our limited understanding. It doesn’t help that there are very few people out there that know it.
At that point it doesn’t matter how good the tech is.
You probably shouldn't do port forwarding on Fly.io; if you're running into an actual need for that, we should talk about extending our network access control model.
If I were running PoE (Postgres on Edge) I'd probably want to connect a local client for poking around, but without the bother of meshing my laptop into the cloud.
$ flyctl proxy 15432:5432 -s -a fizz-db
? Select instance: [Use arrows to move, type to filter]
> gru.fizz-db.internal
iad.fizz-db.internal
lax.fizz-db.internal
lhr.fizz-db.internal
ord (fdaa:0:446b:a7b:20db:0:77a5:2)
ord (fdaa:0:446b:a7b:20dc:0:784c:2)
yyz.fizz-db.internal
That forwards whichever you select to local port 15432.Tangent: that's debatable IMO. In my company's current AWS infrastructure, there's no shell access to either the production containers or the host machines. I did write a script to create an ephemeral container that lets me (and future staff) run a shell inside the production network. And the thing I usually do in that shell is run psql; I suppose that's not ideal for auditability. But still, I can't poke around in the live containers or the host machines; in theory they could be distroless, with no shell at all. I'm trying to take immutable infrastructure to the max here; this seems like a good thing for security. It does mean that for debugging production problems, I can only go by what I find in logs and the database. But that has been acceptable so far.
Edit: Given tptacek's security background, I was surprised that he considered production shell access essential for a new app platform.
If you are running Docker containers and you can shell into local containers, that is usually "close enough" that you can do useful troubleshooting. But fly.io (and CloudFlare workers, etc) are different enough from off-the-shelf containers that it is very important to be able to poke at containers when they break, even if they are not the actual production containers.
It's turning out to have been a misfeature that confuses people more than it helps anyone, and we may get to a place soon where we just automatically provision a root cert for new organizations.
I was more on your side when I wrote the feature, and I'm less on your side now. Also, I SSH into instances to debug things all the time. :)
I'm a fan of this approach, it is hard improve upon the security of a server that doesn't exist.
I recently setup an AWS serverless (mainly ECS Fargate) stack for a project and took this approach of spinning up an ephemeral EC2 server as part of a "breakglass" runbook for the rare cases where such access is needed.
This, combined with Tailscale userspace networking [0] (so the "breakglass" servers can run on a private subnet and are never exposed to the internet), Pulumi [1] (for managing the lifecycle of the "breakglass" instances) and Yubikey based MFA for short lived credentials [2] (required to spin up the server via Pulumi), I found to work well.
This approach is also useful for ensuring that whenever a "breakglass" server is started it is using the latest AMI version (Pulumi's `aws.ec2.get_ami_output()` is useful for this) and runs the usual security updates on startup. The ssh keypair for the server can also be created (and later destroyed) on the fly so there is no need to manage any long term ssh credentials.
[0] https://tailscale.com/kb/1113/aws-lambda/
[2] https://aws.amazon.com/blogs/security/enhance-programmatic-a...
Hah! That one got me. They know their audience :)
Is it something like: "Consul bears a heavy weight (has a lot of responsibility and so responds slowly)"?
1. Could you include more info on which userspace tcp/ip stack you use and why? I presume doing userspace UDP is relatively trivial/fast compared to what slirp had to do with tcp.
2. How does flyctl hijack syscalls across Linux and Windows? Is there some abstraction to do that? Wasn't even aware this was a pattern on Windows.
I realize I could read the code, but would appreciate some direction.
Kudos for the websocket proxy too! Would be really cool if this and the unprivileged wireguard became standard parts of wireguard toolset.
I could try to give you reasons we use it but the truth is that I mentioned something about wanting a user-mode TCP and Jason Donenfeld said netstack was there, and then got it to work.
We don't do any system call hijacking! Don't have to. Go abstracts Dialers and Listeners, and the WireGuard part itself is just a vanilla UDP protocol spoken over a PacketConn.
I expected this to be a headache but it took less than 5 mins to download WG, generate the conf with the fly CLI and paste it into WG. Done.
I followed the instructions in their docs and that's what was recommended.
https://fly.io/docs/reference/postgres/#connecting-to-postgr...
Also, wouldn't I need to install Wireguard anyway to use fly proxy?
I've just been doing research on setting up my own wireguard mesh (currently using a spoke/hub setup with pi-hole/pivpn).
I found https://github.com/HarvsG/WireGuardMeshes today which is awesome, but I'm curious what fly.io / other readers here may be using.
On the other hand fly.io looks really interesting, I want to try it out. The infrastructure described (other than the SSH hack) feels like how a modern cloud platform should be built.
Also, this post makes me update my prior on top HN contributors being unproductive (i.e. that they spend their time on this board all the time instead of working).
So we need assured guarantees that data is not traversing outside of the region. most cloud vendors have specific data residency compliance for India - https://aws.amazon.com/compliance/india-data-protection/
Many EU companies are increasingly concerned about data location, with some going into full panic mode.
> Metrics, logs, and volume snapshots end up on servers in the USA
This will unfortunately prevent me from using/recommending Fly for EU customers at the moment. It should also be pretty prominently stated in the docs.
- What other benefits does it give you?
- This isn't a new problem and presumably has prior best practices for mitigation. What is this replacing, what was the landscape like before this? What was the most similar already?
- Have people been looking for a better solution like this for some time?
- What's the over/under on the maintenance cost? You added in another TCP/IP stack to look after. You maybe save on static configuration and can make your system more dynamic. Pros, cons ... let's list them.
The talks a bit about problems they were trying to address, but not in a way that clearly answered the above for me. It's of course valid to write a piece with a more informed audience in mind, but in something that aims to spread the virtues of an idea I think it could do more.
What I think user-mode TCP/IP gives us is the ability to build arbitrary services --- Postgres, Redis, SSH, network management, whatever --- without having to make infrastructure changes. We don't have to have some weird API or application proxy that knows what's running and who's allowed to run what. Instead, that's simply baked into the network, and flyctl, by dint of netstack, can just use it. If somebody comes up with a cool network service to plug flyctl (or any other tool someone wants to write) into, it will just work.
But things like the maintenance cost, well, yeah, that's most of what the post is about. The maintenance cost was not especially low.†
There's a natural inclination to read any post like this as a kind of brag, but I'm really just experimenting with trying to show the good with the bad here. User-mode TCP/IP is a weird choice! Nobody else I know of does it! It might have been the wrong choice! Even though I love it!
† It's actually not low even right now; I'm spending the first half of the day deploying code that relays stats from Netlink on our gateways through our GraphQL API, so that flyctl can check WireGuard gateway health. That is not a thing we would be spending time on if we had just written an explicit proxy for Postgres or whatever, rather than providing a generic network transport.
I'm in automotive/embedded at the moment, and our daily battle is making decisions on how static vs. dynamic we want our system to be - static (e.g. resource allocations or baked-in scheduling decisions) makes it easier to reason about the system and provide guarantees, but generally lowers efficiency at runtime. Dynamic can make the system much better at serving a wide range of usage scenarios, but makes it harder to eliminate the risk of pathological cases. It's hard not to see things through that lens. The way to construct and run services you've described here to me is an interesting option on that type of axis, in the sense of where the costs/friction goes.
Sure, just how many distinct WireGuard configurations would the gateways be comfortable with per-Fly app? 1K+? 100K+? 1M+?
I keep thinking of the nightmare of keeping up with the world of Internet middleboxens, broken net layer implementations and icmp hacks that the Linux kernel supports and makes 'just work'. The jump to usermode tcp seems interesting if you're not worried about that (and I've been watching the formal-proven ip/tcp stack space like a hawk for years), but I've been burned so many times with non standard stacks and 'oh you need to connect to that non-updated lynxos system and huh' or 'hah could you enable ecn or this obscure tcp option because... Legacy?'... And sometimes I need tc/netem and netlink and I don't know...
What's wrong with that? ;)
It's also one of the three major projects that use it besides User-mode Linux and rr.
Author is @strigeus of uTorrent/Spotify fame.
flake: curl -v https://fly.io/blog/ -sS -o /dev/null -D-
* Trying 2a09:8280:1::a:791...
* TCP_NODELAY set
* Connected to fly.io (2a09:8280:1::a:791) port 443 (#0)
* ALPN, offering h2
* ALPN, offering http/1.1
* successfully set certificate verify locations:
* CAfile: /etc/ssl/certs/ca-certificates.crt
CApath: /etc/ssl/certs
} [5 bytes data]
* TLSv1.3 (OUT), TLS handshake, Client hello (1):
} [512 bytes data]
* OpenSSL SSL_connect: SSL_ERROR_SYSCALL in connection to fly.io:443
* stopped the pause stream!
* Closing connection 0
curl: (35) OpenSSL SSL_connect: SSL_ERROR_SYSCALL in connection to fly.io:443
Seems to work when forced to IPv4 though.Thank you for helping with this!
It helps that it was designed and implemented by a kernel exploit author.