Remote access to production infrastructure (death to the VPN)
mattslifebytes.com
mattslifebytes.com
Obviously, this isn't practical for everything. But, if the thing you were using VPN for is already a web application, you are basically halfway there. Ideally, you just directly expose a secure web application to clients, but in some cases (i.e. very old legacy systems) you probably want to put an nginx box in front and then put the authentication at that level.
Web access has a huge range of benefits. Users are scoped directly to the system of concern rather than an entire network of hosts. You can take security to the next level with server-side rendering of web content in order to avoid additional required channels of communication or revealing of implementation secrets to the client (e.g. SPA client source).
We are at a point of placing our actual application servers directly on the public internet (with TLS1.2/MFA/ACLs/etc). Hiding behind VPNs or layers of reverse proxies seems to cause more harm than good.
If you have the engineering resources to back it up, it definitely can be. Internal services at Google usually trust the office network the same as any other -- well documented in the BeyondCorp paper if you're interested.
Someone at Google please feel free to correct me on this.
(FWIW, I don't think there's anything secret here. This stuff is very explicitly described in the whitepapers.)
You don't get direct SSH access to production machines or any other lower level network access like packet sniffing on the production network.
1) I want the websites to do certificate verification on the certs I'm using on my desktop.
2) Then on top of that my website should use usb security key verification as well.
Easy enough to do #2, but I want #1 to be ubiquitous as well.
... So basically my HTTPS server will use my public key as my identity, not my username/email and password.
Happy to be proven wrong, I'm just unaware if any popular open source HTTPS servers offer this as an integrated solution.
Or better yet, I'd like raw access to the certificate info FROM the application layer on server-side so I can manage that as needed.
In the end you'd easily get e.g. daily U2F with monthly cert rotation. Or whatever you end up wanting.
Open source have had this convered for years, but you'll have to look for it.
Disclosure: I'm the main author of all those three projects.
If you don't care about the user (TLS client certs + standard username/password) you can get away with proxying the application through nginx and calling it a day.
Basically you turn on client verification and you're done. If you want to show an error to unauthenticated users, you can make verification optional and add something along the lines of: if ($ssl_client_verify != SUCCESS) { return 403; }
Disclaimer: I haven’t don’t this myself (yet), but have read about it a bit.
But in this case users could provide public keys they will use when accessing the website from internet (as opposed to intranet).
I personally use Envoy as the proxy and cert-manager to manage certificates internally. You can peruse my production environment config for my personal projects at https://github.com/jrockway/jrock.us (dunno if that's the real link, my ISP broke routes to github tonight, but it's something like that).
The flow is basically:
1) At application installation time, a cert is provisioned via cert-manager. Each application gets a one-word subject alternate name that is its network identity. The application is configured to use this cert; requiring incoming connections to present a client certificate that validates against the CA, and making outgoing connections with its own certificate. (This integrates nicely with things like Postgres, that expect exactly this sort of setup.) This lets pure service-to-service communication securely validate the other side of the connection. This is nice because, in theory, I don't have to configure each application with a Postgres password, Postgres can just validate the client cert and grant privileges based on that. (I have not set this up yet, however.) I also like the ability to reliably detect misconfiguration; if you misconfigure a DNS record, instead of making requests to the wrong server, the connection just breaks. Saves you from a lot of debugging. And, of course, if the NSA is wiretapping your internal network, they don't get to observe the actual traffic. (But probably compromised your control plane too, so it's all pointless.)
2) The other half is letting things outside of the cluster make requests to things inside the cluster. I use an Envoy proxy in the middle; this terminates the end user's TLS connection, and routes requests to the desired backend, like every HTTPS reverse proxy ever. I wrote a "control plane" that automates most of the mTLS stuff (it's production/ekglue in the repository; ekglue is an open-source project that is agnostic to mTLS, my configuration adds it for my setup). At this point, users outside of the cluster will see a valid jrock.us cert, so they know they've gone to the right site, and applications inside the cluster will see that traffic is coming from the proxy, and can decide how they want to trust that. Right now, everything I run in my cluster just passes through to its native authentication, so it's pretty pointless, but the hook exists for future applications that care.
3) For applications that want a known human user (or human-authorized outside service, think dashboards or webhooks), I wrote an Envoy ext_authz plugin that exchanges cookies or bearer tokens for an internal request-scoped time-limited access token. Applications can then validate this token without calling out to a third-party service, so no latency is introduced. (They do have to be configured to do this, and the state here in the open source world is pretty abysmal. OIDC is helping, and it's trivial to write it into your own application framework. A few applications will just accept an x-remote-user HTTP header, which I found to be adequate, especially if they can trust the proxy with mTLS. Compromising the proxy lets you compromise all upstream apps, though, so I'm looking for a new design.)
I actually wrote this at my last job and don't have the code (it's theirs)... but am slowly rebuilding it in my spare time. Second system syndrome is a bitch. You can follow along at my jsso repository on Github, but it is not ready to be used and I think that most of the stuff I wrote in the design document there is going to change ;)
Anyway, where I'm going with all this is... all the pieces exist to make yourself a secure and reliable production environment. mTLS is pretty straightforward these days, and in addition to the easy route of just doing it yourself, a bunch of frameworks exist to let you get even more security (SPIFFE/Spire, Istio, etc.) For authenticating human users, most of the work has been done in the closed source world; Okta, Duo, Google's Identity Aware Proxy, etc.
Defense in depth is a concept that should be applied with some thought, it would be good if your additional layer did something different. For example good reactive security, endpoint attestation, etc.
Multiple layers, security in depth...
1. With NAT and metropolitan area networks, hundreds of thousands of devices could share the same public IP.
2. Large networks with many devices often connect to the public network through trunking (load balance the connections through multiple routers), so the HTTP connection between OKTA and my browser can VERY well originate from a different IP address than my SSH session, and I would never be able to connect.
3. Many devices are mobile, and they can change their IP address when they pass from WiFi to LTE for example. This would force an unnecessary re-auth.
The correct solution is somewhere in the middle: block everything by default to get you to an inner courtyard, where the zero trust model is deployed... (which ironically he suggests by deploying port knocking (port knocking is a bad idea (TCP/UDP ports are sent in the clear and the "key" is never rotated)))
The best model is probably "block everything" by default, then allow access to the inner courtyard via a VPN, where then the Zero Trust model is deployed. You remove the ability for an attacker to have unlimited retries, but access to resources still requires individual authentication.
A VPN means an individual connection is authorized into the interior courtyard.
A Lambda with 2FA to whitelist an IP, then a cron job to cleanup means everyone at your local cafe wireless access point is also authorized into the inner courtyard.
The answer is of course defense in depth.
I wonder why the author seems to reply to every comments other than this.
Obviously defense in depth can go as deep or shallow as you see fit, given an organization's resources. We believe that the short-lived SSH certificates, IP whitelisting (via "enterprise port knocking"), endpoint authentication (device trust), password authentication, and multifactor authentication are enough to protect a single production deployment. Encompassing all of that with a VPN seemed unnecessary when other protection mechanisms like the above, and additional mechanisms that we won't speak to publicly, are taken into account.
Like with anything, it's a game of risk, and it is up to each organization to decide what risk level they will tolerate. I believe most organizations have deployed VPNs in a way that gives them a higher exposure, and simply wanted to share some of the things we have learned through the process :)
The only issue I've run into with port knocking is places that heavily restrict outbound ports/protocols. Though technically that is solvable too I just haven't bothered.
The main open source option I'm aware of now with support is Pritunl Zero. Was going to actually stand that up today before I read the article.
ZT is when you move [strong] authn and authz to the endpoint itself.
IAP supports HTTP and TCP connections, so you can put it in front of your website (say an internal admin webapp), or use it to tunnel SSH onto a machine that doesn't have a public IP, using your IAM roles.
If you're running Kubernetes in GKE, you can also wire IAP up to an Ingress, to protect any TCP/HTTP services you have in your k8s cluster. This one is a bit tricky to configure, but is very nice once you have it up and running.
What we used VPN for is allowing us to establish SSH connections to the equipment. I would really, really like a low-resource mechanism to replace this but everyone wants to deploy their solution in a 90+ megabyte Docker container, or a Snap, which is about 1.5 times as large as the entire Linux system image for our oldest equipment. So these are great solutions for when you control the entire network path from the server(or when you are using a server!) to the Internet including the firewall, but they suck terribly for eliminating VPN in cases where you can't just open an inbound port on the firewall.
As it is I'm trying to figure out how to configure an OpenSSH client to punch out through the firewall to an OpenSSH server, then immediately turn around and provide a shell to the server. This seems to be entirely contradictory to how OpenSSH is designed, but I'm hopeful I can hack something together.
If so, then that's not as much a hack as a pretty standard reverse ssh.
> As it is I'm trying to figure out how to configure an OpenSSH client to punch out through the firewall to an OpenSSH server, then immediately turn around and provide a shell to the server. This seems to be entirely contradictory to how OpenSSH is designed, but I'm hopeful I can hack something together.
This is trivial. But if you don't control the firewall, how will you get the outbound SSH access? PCI requires that both inbound and outbound traffic from the secure zone (CDE) be controlled. If you can impose upon the customer that they punch an outbound hole, you can impose inbound requirements as well. Your inbound connection does not come from "the public internet", it comes from your managed in-scope network.
https://github.com/pomerium/awesome-zero-trust
PRs welcome.
Also, are there any concerns about IP timeout vs explicit VPN disconnect? Obviously the latter works better in shared environments (e.g. shared terminals, wifi's that reuse IPs frequently, large NATs that have many devices behind a single IP).
For web apps, we simply front using an OAuth2-aware proxy. Back in the day, we used this: https://mattslifebytes.com/2018/08/07/protecting-internal-ap... Now, we utilize Kubernetes for hosting most production internal apps, so we run the oauth2-proxy Helm chart (https://github.com/helm/charts/tree/master/stable/oauth2-pro...) to handle verification of identity before sending traffic back to its destination service. Conceptually similar, as auth has to be completed before the request is sent to the back-end.
VPN is dead because some customers want you to route the internet interfaces of all machines through the VPN server.
How does this even make any sense?
Is it actually possible to have a technical control like this?
Why can't I create a container or virtual machine that just runs a VPN client, and then use the virtual machine network controls to decide what host traffic gets routed to the VM and through the tunnel? How would the VPN client running inside the VM know about anything I'm doing one level up?
Or is this just another bullshit "you don't actually control the software running on your machine" technical control?
Realistically, the usual plan is to create controls that are impossible for most non-technical users to bypass, inconvenient for anyone else to bypass, and back them up with the threat of disciplinary action.
There may be real cryptography over the wire, but there's nothing "strong" about the assumption you mentioned, or the disceplenary threats. If the threat model assumes that I can't extract a key from a laptop, or clone the behavior of some garbage Cisco client, that seems pretty broken to me.
Commercial VPNs are mostly just shitty software for enforcing shitty corporate policy, disguised as a remote access tool.
The PC you use with the VPN is never to be directly connected to the internet. It connects to a piece of dedicated VPN hardware. (could be a Raspberry PI with special software or something far more expensive) That PC can use the VPN, and thus get to various computers within the company, but it can't go elsewhere. No other business is reachable.
The company can allocate IP addresses without NAT and without regard for the rest of the world. There just isn't any connection to the rest of the world, so conflicts can't happen.
Split tunnel vs not split tunnel means nothing if the client doesn't want it to mean something.
More specifically it's a pain in the ass if you use AWS ALB load balancers and whitelisting. Those IPs aren't consistent and you typically can ONLY route on IP.
We do it because it's better than the alternatives, but our setup wouldn't scale past more than a few applications.
I blog about things that I encounter at work and find interesting. That happens to often be a cross-section of infrastructure and identity!
We have used multiple OpenVPN servers with password protected cerificates and TOTP. Even if someone were to obtain access to my credentials and certs, they wouldn't be able to access the production services without also obtaining access to the authenticator device. Once your machine is enrolled in ScaleFT and while you're authenticated with your identity provider, malware or just a malicious coworker could access the production services with a single command line.
There are upsides to ScaleFT as well, though. As long as you're all in on Okta or can federate with it, user management is a no brainer. And having the IdP integration is much more user (and malware) friendly and is likely more reliable for server to server use compared to OpenVPN. Limiting access to particular services is likely easier, too.
Downsides with this product include having all sorts of reoccuring configuration problems where a server just disappears from the list of available services, which requires ops involement to restore access. If you're using macOS and RDP (I just outed myself to Matt...) you have to use the sub-par FreeRDP client. And ultimately you're tunnelling TCP over TCP, which works ok in the office but which might not always work as well in mobile or higher latency network situations.
To the best of my ability, my goal was to make the post more about the network architecture (esp around the concept of SSH bastions) and less about the actual OASA product itself. I think there are a number of fungible solutions which would be just as effective (though I think the integration with Okta is a key product feature). What I find interesting and novel is more what we can do to only open ports to authenticated IP addresses, and to address connections between a single source and a single destination. To me, that's where the real power lies.
WHAT?
AWS already has a vastly superior solution for this, called AWS Systems Manager Session Manager (it's quite a mouthful). You create a session with AWS using a federated login service (SAML-based SSO) and then craft IAM policies to allow a single user to ssh into a single server, over the AWS API. Not only will this be more secure, you don't have to maintain a wacky custom solution.
Logging into servers is an anti-pattern, and wherever possible you should be running away from it. Get metrics out of the server and analyze them, run commands remotely using some kind of persistent system agent, stop storing state on your servers. I know this is not the point of the article, but I want to remind people of it so they can can avoid the ssh trap early.
https://www.openbsd.org/faq/pf/authpf.html
The FAQ entry is about building an authenticated gateway, but the same technique can be applied to open individual ports.
Unless you want permanent connection with routing and everything, ssh socks proxy work awesomely.
Point a firefox profile to use it and you can really act as if you were in a different subnet: it can proxy dns resolution too.
[0] https://github.com/elpy1/ssh-over-ssm <-- not made by me, but a good example
This has some serious downsides for non-SSH applications. For example, to connect to a production database cluster, one would need to ssh through the proxy to a bastion host, and then set up port forwarding from the bastion host to the database. Setting up a simple database connection now requires shell access to a production server. This is less secure and more complex than using a traditional VPN.
I'm curious: why is utilizing port forwarding over these mutually authenticated SSH tunnels less secure than employing a VPN? From my perspective, port forwarding still adds a level of intentionality which reduces the likelihood of an incident/accident.
If intentionality is desired, one can use per-server VPNs.
> OASA also protects these hops by issuing client certificates with 10-minute expirations after first verifying your identity through our single sign-on provider, and then also verifying you are on a pre-enrolled (and approved) trusted company device.
Let elect ZeroTier to be the president of remote, secure access :)
hopefully ZeroTier makes some strides in 2.0.
So now I gotta go decide if I want to rip out all my existing ZT infra or not.
I'll keep using fwknop-protected OpenSSH on OpenBSD and WireGuard, others can do whatever they want without thinking about the security vs. convenience.
From a post awhile back about using Bastions with ScaleFT:
> One of our values at ScaleFT is to do our best to support our users where they are, with the decisions and tools they’ve already selected. This means treating SSH bastions as an SSH feature, parameterizing and centralizing the associated configurations, and seamlessly integrating it into our users’ daily workflows.
https://www.scaleft.com/blog/bastion-hopping-with-ssh-and-sc...
So, if you want to layer on top VPNs, or SSH Jump Boxes, we try to let you. We also try to make parts of the chain better whenever we can.
(disclaimer, I'm ScaleFT co-founder)
Yes, I completely agree. This post is literally an endorsement of that idea, with enterprise port knocking mixed in for additional security. At no point in this post do I advocate simply opening all servers to the Internet. Quite the opposite.
If you have suggestions for how I could be clearer in the post, please let me know.
Really?
Bridge a network device – such as a laptop or even another server … into a larger network of servers – such as in the cloud or on-prem … across the Internet – protected with an additional layer of encryption