BeyondCorp is dead, long live BeyondCorp
mayakaczorowski.com
mayakaczorowski.com
As far as I could tell it was a huge group of bored people chasing Impact and not solving real problems or even bothering to read tickets. Most of those problems were the product of previous generations of Impact-chasers.
I've long been curious how early boot works on Google servers (and TIL workstations too, although it makes perfect sense) - primarily because I want to copy the techniques myself! :D
How is key storage and device attestation actually done?
In reality, what we can expose, is that during the early boot process, our systems reach out to another system that register's its interest, and as a user (from another trusted device, phone, laptop, etc.) you can visit the web service and click a button.
How and where keys are stored is a great big "?" that's up to the implementer to solve..
Of course I want to know how and where the keys are stored :D so I can do the same thing myself! Arguably security systems in this class demonstrate their integrity *because* their architecture is fully open, documented and straightforwardly reproducible. It isn't science if it isn't reproducible, right? Something something computer science...
That's the idealistic view, of course. In the <insert cartoon punching fight cloud here> real world, we have The Legacy PC Problem™, where secure boot isn't, TPMs can be bus sniffed, SGX doesn't really support the hacker/tinkerer exploration necessary to power defense in depth, ME is a ginormous black box that eats authentication headers like they might as well be glue... and it doesn't matter that I am a dog residing on Mars because the MDM my 2FA device is signed into has decided I'm legit.
Hence my interest in real security. It's a giant debacle, surely there are some genuinely cool wins to be had out there that truly make a dent ._.
You can then monitor whether your service is up, and if you get a notification that it's down use the hosting service's management interface to reboot the machine/VM and then do the SSH thing.
- Box with Ethernet port that connects straight to CorpIT (something something service account) and client/gadget USB port that shows up as mass storage; you talk to the box via a remote server using a userspace tool that speaks FUSE. Everything staged/uploaded gets archived. You could go one step further and tie this in with goma and have the box receive securely built code, so you'd only need to archive git metadata instead of giant blobs.
- Engineering-specific Drive frontend could give you token/TOTP/straightforward based access to files, with download audit logging/data archiving (...this would be so simple for CorpIT to provide...)
- PCI card with two Ethernet ports, one special one that plugs straight into CorpIT and a data port for single devices or even a local n-port switch; Linux on the card signs the special port into CorpIT via a service account and then uses custom RPC/whatever to mirror all traffic flowing over the data port. Chances are this setup would likely only be able to viably do 10Mbps given that all traffic would be mirrored. The card would require a wakeup (and/or firmware upload) event before the PHY appears, preventing EFI boot confusion.
- "Secure bringup services" branch of CorpIT implements processes and makes services available that focus on the fundamentally open-ended requirements of hardware design, which would ideally provide more cohesive solutions that the workarounds described above.
IMHO I reckon all those bored engineers would have an absolute field day doing stuff like the above (chasing Impact in a hardware context); it's a great pity they aren't adequately enabled to demonstrate the security implications and priority of doing so.
And so when he tried to work with a machine without exactly one(1) IP and one(1) active NIC going up the NAT router to the Internet it immediately failed. This is surprising to me that of MANGA, Google seems to be on the weak side with respect to IP networking.
For almost 2 years of mostly remote working during the pandemic, beyondcorp has caused me trouble maybe 4 times, and 2 of those times were when my office re-opened and I had to revive my security keys that hadn't been used in over a year.
There is a VPN, but I've never had to use it. They kept my corp desktop running in the office which I could use as a jump host when I needed it. Tunneling X11 over SSH worked well, and the browser-based remote desktop worked ~ok.
That doesn't reflect my experience as a Googler. I have one corporate issued laptop that works great for everything from software development, flashing testing devices, managing production services, and browsing the internal meme boards. I don't have to switch devices. I'm not sure what relevance personal devices have. You are pretty restricted in using personal devices for work (phones being an exception).
I think BeyondCorp ended up working fantastically well during the pandemic as basically the entire company has been working via untrusted networks. It's hard to imagine using a clunky old corporate VPN from the past anymore.
Maybe the 2 devices are a PC and Mac because the employee needs tools that only work on one platform or the other. But that's orthogonal to BeyondCorp.
I was discussing about ZT with a friend recently and we were agreeing that one of the problems with the USGov memo (and most of ZT advocates) is referring to ZT as an "Architecture". The memo paints a picture of ZT as a destination whereas it really should be understood as a framework, culture and design philosophy. And that makes it, by definition, a journey. Its principles are supposed to guide your architecture design but they are not the architecture i.e there can never really be a point where you can call a friend and be like "Look at this, I've finally 'built' a Zero Trust Architecture". And you can't have a consultant come in and go back a few months later telling you "Alright, here's your Zero Trust Architecture". ZT has to be continuously entangled into your dev flow, ops, policies and day to day technical decision making.
I also suspect that another important missing piece (whether you look at it as a journey or a destination) is how to quantitatively MEASURE progress on Zero Trust. Having precise reference metrics would help in actually enforcing the goal of the memo or at least being able to tell that company A has a better measured ZT progress than company B.
I guess, like they say, "Zero Trust is like teenage sex: everyone talks about it, nobody really knows how to do it, everyone thinks everyone else is doing it, so everyone claims they are doing it."
Disclaimer: Googler but I don't work on the BeyondCorp team.
I think the whole DevSecWhateverOps thing fails to account for the severe antipathy large organizations have for outside-the-box solutions. A solution that requires people leave their silos, learn new concepts, or adopt new practices is just too much for them.
You can also see this publicly in GCP's Workload Identity and ALTS primitives, which enable very sophisticated policies.
Most BeyondCorp concepts seem simple, and they are, but they depend on a lot of existing machinery, almost all of which is non-existent in pre-existing corp networks. The average tech company is currently struggling to catch up.
My point was it is simple for an app developer to integrate with.
To measure ZT progress, count the services and security controls which rely on VPN/LAN. The closer it is to zero, the closer you are to ZT.
Putting your internal apps behind an OIDC proxy instead of the VPN is a straight upgrade at that point. Especially if your provider already does some checks for you (e.g. Chrome Enterprise, requiring Cloudflare WARP app)
For example, all VPNs give you a single lockout/change password point, so this has to be at least common authentication. And VPNs have good logging for connections, so you got to have this too. And if your VPN client had device attestation, your solution should have have this too. And also services on VPN are not nearly vulnerable to auth bypass bugs, so your webserver should have some protection on this too (authenticating proxy?).
I think !VPN is exactly hallmark of ZT... as long as security is preserved. If you disagree, what do you think ZT's hallmark is?
(There is a philosophical question: if a company had a crappy VPN with no logging, and multiple VPN servers with no central auth.. and went to equally crappy user/pass webapp on internet.. does this count as "ZT transition"?)
The point of all of this is perimeter security becomes an optional, small contributor towards a secure context. It turns the entire legacy corporate network on its head.
Especially at large megacorps, it's borderline impossible to perform incremental migrations towards a zero trust network without first killing perimeter security as a concept. You can always introduce it later, but it's a different beast entirely.
Proxy (BeyondCorp): services remain naive but employees and offices are evicted from the trusted network. Employee requests go to the proxy, which checks authentication and authorization before passing along the trusted network to services. Services and potentially engineers/SREs who can get on the production network can still make whatever requests they like.
Zero Trust: all services on the production network authenticate all requests. Even if you are root on a production box you would need tokens/certs to get useful responses from other services on the same production network segment.
I entered the industry at a “proxy” company that is pushing towards zero trust. It’s hard to believe that the “perimeter” model is real. But you look at something like Target getting owned by thermostats in its stores that merely needed internet access, it is clear that some enterprises do work that way.
Having an nginx loadbalancer doing all the beyondcorp stuff and forwarding on any authorized request is pretty straightforward and covers a large chunk of what your employees will be wanting to do. It easily lets you allow employees to access low risk services from their personal unmanaged and untrusted phones. It also means any home grown internal service doesn't need to do Auth - it can just look at the trusted header from the loadbalancer and know who the user is logged in.
You have then at least substantially lowered the remaining attack surface, which now consists of just the non-http services (shared drives, SSH, remote desktop, etc.).
For those, a VPN server allows you to authenticate the user and device, and a big dynamic set of iptables rules lets you decide which sessions can access which service.
It's perfectly capable that you control things like SSH and other services through a Zero-trust system as well.
Personally, there's things like Pomerium.io that can handle HTTP(s) services, as well as TCP services.
Even internally at Google, the machine, and user are used to validate SSH sessions, and the remote system has a unique policy on who's allowed to access it.
Disc: Googler, not on BeyondCorp
All you need is a centralized authentication server (preferably including access logging) and a simple way to use it for all applications (the simplest is to have the app listen on localhost and expose it via nginx configured to use the authentication service).
I guess for real security you also need to forbid logins from non-secure devices, which is trivial if you can trust employees and only hire security-competent engineers, and otherwise needs dedicated devices, secure boot and remote attestation (which is also easy but will take some work to do).
The old-school variant has the same issues, IMO? A previous employer of mine was "everything was on the company wide VPN" i.e., no SSO, half the stuff is insecure.
Externally, you could connect to the VPN with a fairly bog standard VPN client. But how was the trust established? Well… it trusted DigiCert's CA cert. Meaning anyone who purchased a DigiCert cert & could obtain a privileges position on the network could MitM the VPN connection.
Inside the office, the WiFi was connected to the VPN (i.e., VPN client was only required off-site) … and the office WiFi was WPA2-PSK. And of course the PSK a. was not a good PSK and b. did not rotate when employees left the company.
Worked for a larger corp. Same idea, internal net was trusted net. STP packets on the ethernet ports, which IIRC my network training probably meant that I could convince a switch on the network that I was a switch, please start routing me traffic.
> They’ve tried to use the tools already available in the market themselves
I've tried to use more old school solutions, like LDAP. LDAP tooling is horridly difficult to set up. I gave up on my first attempt; I think today I know where I went wrong. Things like pfSense & VPN are also incredibly complex, and my understanding on the security community consensus on IPSec is that it is that its complexity guarantees that it is insecure.
> The tools that are on the market today aren’t even doing the hard part of zero trust yet.
But I don't disagree with that, too. The tools are definitely not up to snuff. Getting working SSO, even with just a decent MFA experience & then getting a service to authz with it has been considerably complex. & like the author surmises … even if I ignore device security.
Now we're all remote … yet security would still like to allowlist network access by IP address, meaning one of these days I am creating a common, centralized white pages of employee IP addresses because those are getting numerous. (I have mixed feelings on this. The list is a PITA to maintain. But it does keep random Internet riff-raff from ever hitting an open SSH port… and while yes, people should use keys, all it takes is one mistake. And people make mistakes.)
ZT is IMO the right approach; the network is already compromised. Whether that's at the Internet or the port on your laptop doesn't matter. Yeah, people definitely forget the device bit. (And ought not to.) But the other approaches aren't implemented with any more diligence.
Yeah, I've never found network topology to be a good way to manage trust. Even ignoring employees sticking Raspberry Pis under desks or reconnecting using non-rotated credentials, it's ridiculously easy to convince apps to make internal requests expecting them to be external requests. Slack had a big security incident when its url unfurler, which runs on its internal network, started making requests to other internal services, thinking that they were external websites. The problem is that internal addresses were implicitly trusted, and people have endpoints like /quitquitquit hanging around, and that combination ends up being game over.
Ultimately, every request depends on at least two pieces of information: what user is making this request, and what downstream application is making this request. Most people only take into account the first, and thus these problems recur. (Because operators get made when they "kubectl port-forward" in and the application rejects their debug requests because "random unauthenticated HTTP request" does not meet the security requirements. This, of course, is a good thing. For security, anyway.)
> Getting working SSO, even with just a decent MFA experience & then getting a service to authz with it has been considerably complex.
Yeah, the industry seems to have decided upon OIDC, which is significantly more complicated for both the operator and the application developer. I really like the way Google's managed auth proxy works, and I'm surprised it's not more popular. If the request goes through the proxy, it injects a signed header with the user information in it. The application simply uses a few lines of code (ok, JWKS is involved, so a lot of lines of code to keep the list of trusted public keys up to date) to verify the signature and extract the username, and then can make an authorization decision. No cookies, no redirects.
I ended up writing my own proxy that uses username + WebAuthn to authenticate and pass this information on to applications behind the proxy, and it's nicer than any auth solution I've paid 100000x more for. I can FaceID into internal status pages when I'm out drinking, impressing everyone! OK, not very many people are impressed, but they can at least see the thing I want to show them. I'm surprised there's no maintained OSS thing that works like this.
2) Netflix runs a single application, whereas organizations maintain and provide hundreds
3) Netflix doesn’t need to do any device assurance or inventory
4) Netflix has essentially one level of access to enforce, vs hundreds of roles across hundreds of applications.
2) Okay but a web client will get you 80% the way there.
3) Yes/No. Netflix doesn’t but that’s because Google does it for them for Widevine deployments.
4) Widevine has three, pretty complex, policy levels. No reason there couldn’t be more. Plus device authentication is just a metric used in authorization decisions. Nothing is stopping you from having arbitrarily complex authz.