First, agreed. Certificates can't and shouldn't attest to the user's identity. You need another mechanism for that.
If you're securing requests, then securing the channel might be redundant. However, I do want authenticated service identity (not just user identity). The shape of policy I want to enforce is: "<service x> should be allowed to make <rpc> to <service y> in the context of responding to <request> from <user z>". I need the authenticated identity of both the service and the user to do that. Either that, or the authorization could take place higher up in the stack and I need a capability token passed through. Even with a capability token, it's nice to know the authenticated identity of the calling service for other purposes, like audit. Not to beat a dead horse, but I realize that macaroons satisfy all of these requirements :).
I also want confidentiality. Macaroons don't do this. This is an important point that I want to make more strongly this time, because I feel like it's getting lost in the noise. Authentication isn't the only thing people are using TLS for.
I appreciate your perespective. I think the best way to respond is to actually answer @halvarflake's question: what's the threat model that justifies service-to-service mutual TLS?
One answer is: that's the wrong question. Many people use mTLS for flexibility, not security.
Most microservice systems are born as a bunch of web services thrown in a VPC (or, more recently, a kubernetes). They can all talk to each other. Some of these systems are successful, and eventually outgrow their perimeter. Sub-components need to be able to communicate across untrusted channels (e.g., the web). TLS can give you a "cryptographic VPN" to do this.
There are a lot of VPN-y solutions with a lot of pros and cons. TLS isn't always the right way. The big pro for TLS, however, is that it works everywhere. Even if you don't control the host stack (e.g., functions as a service) or if you're on a constrained device, etc. TLS is also baked into most infrastructure: proxies, queues, databases, etc. You can easily terminate TLS easily at your perimeter if that's what you want. Alternatively, you can connect directly from a service in one perimeter to another service or to a piece of infrastructure like a database in another, without a proxy. Again, TLS is ubiquitous and it's flexible.
At this point I think we're at your "you must be this high to get on the ride" model. Why go further? This is where it makes sense to start talking about threat model.
First... I already mentioned this, but, once again: confidentiality. I think this is less controversial, and it doesn't require client certs. So I'll not elaborate further except to say that people need it.
The interesting topic is: service authentication. Why use client certificates for service authentication? I contend that service-to-service mTLS plus a bit of simple, coarse authorization actually does address a number of common security threats. Specifically:
* Malicious insiders: if you're using microservices and practicing DevOps and agile and have engineers on-call you probably have a lot of people and a lot of code that can get inside your perimeter. Coarse segmentation of service-to-service interactions can help here. An engineer who has the ability to `ssh` or `kubectl exec` or tunnel into your perimeter won't immediately have access to every sensitive service you run.
* Lateral movement: very similar to the above, but this time a service is popped or an engineer's laptop is compromised. Service authentication helps control the blast radius if someone gets inside your perimeter.
* Credential leaks: secrets are hard to manage. They get committed to repos, accidentally logged, dumped in error messages and served to end users, pasted in to slack channels, walk out the door with dismissed employees, etc. Using certificates with a relatively short lifetime gives you assurance that compromised credentials will be eventually rotate once an attack vector is closed. I contend that it's actually a lot easier to set up client certificates that automatically rotate than it is to rotate secrets for all of your proxies, databases, services, etc. An asymmetric crypto key is also inherently less likely to leak to logging infrastructure or end users than a password or bearer token since it's not actually in any requests.
Request authentication could mitigate these threats, too. And, again, carrying an authenticated end-user identity through the stack to downstream services is valuable. But, in a microservice system, I'm less worried about code lying about the end-user it's acting on behalf of than I am about malicious humans. If I already have TLS for "my VPN", authenticating clients and coarsely authorizing service-to-service interactions starts looking pretty attractive.
Regarding compromised keys, certificate lifetimes, and revocation... that's all important, but less critical to the threat model than it may appear. If a key is compromised I'm still much better off than I was without my coarse segmentation: the only stuff you can access is the stuff that the service is authorized to access. If an engineer exfiltrates a key to "test something", that key will eventually get rotated, even if I never know it happened. Same with an attacker exfiling a key: they have to maintain a presence on compromised systems to exfil new keys as they rotate. (FYI: by default, the smallstep toolchain issues certs for 24 hours. Personally, I'd push that down to a few minutes.)
It's also worth remembering that we still have our perimeter (or perimeters). If we find an attacker, we can lock them out at the perimeter instead of revoking a certificate. Even if you can TLS through the perimeter, we only need to enforce revocation at stuff directly exposed to the internet (e.g., in a proxy, not necessarily in every service). Active revocation is nice. But it's not necessarily critical. As another example, since you mentioned SSH: smallstep's SSH product uses certificates without active revocation. We revoke access in policy, by pushing ACLs to hosts and locking users out via NSS and PAM. The certificate may still prove I'm "mike", but mike doesn't have access anymore.
Why not do all of this host-to-host instead of service-to-service? Well, depending on your requirements, that may make sense. Certainly in kubernetes-land the line between "a host" and "a service" are blurring. TLS still has a few advantages:
* It works in places where you can't control the host (e.g., functions): it's a consistent mechanism that works everywhere
* It makes applications more self-contained and, therefore, more portable
* It's logically closer to what you want: you're trying to authenticate services not hosts
* There's more deployment flexibility: you can terminate in a sidecar or in the application
* It's easier to pass authenticated identity information into the application
* It works for authenticating to infrastructure: proxies, databases, queues, etc (again, it's a consistent mechanism that works everywhere)
Fin.
(Are we using macaroons yet?)