If you find someone malicious in your network, you lock them out of your network. You don’t need to revoke a certificate to do this.
OR, you can do CRL. This isn’t jazz hands, it’s a choice. As you push mTLS harder, and need things like active revocation to satisfy your threat model, it’s worth reconsidering whether mTLS is the right choice. Sometimes it is, sometimes it’s not. Sometimes you don’t have a feasible alternative.
You can't "lock someone out of your network". That's handwaving. If you can do that reliably, you don't need any other security controls. Just lock all the bad people out!
The problem with mTLS in configurations where you rely on it for authentication is, if an attacker manages to obtain a client certificate key, it no longer matters if you nuke the machine from orbit, because they have the keypair! The keypair still works! They can use that keypair from any other place on the network they have access to (if there's no such place, you just defined away the need for mTLS).
Which is why you need active revocation if you're going to rely on mTLS: when you lose confidence in your sole custody of a client keypair, you need to actively revoke that keypair, immediately. You can't just wait 7 hours for the step-ca default key lifetime to expire!
"Or, you can do CRL" is literally active revocation. Your project has a prominent call-out about how most people don't need to do this. But that's literally the opposite of the truth. It's true in the WebPKI, but I think you've become confused about the difference between WebPKI server certs and internal client certificates. Telling people they should deploy lots of mTLS but not worry about revocation is malpractice.
What's going to happen to people in practice is that they're going to have to re-key their entire fleet any time there's any question about the integrity of a service. Or, more likely: they won't, because the mTLS deployment mode you're encouraging creates so much goddamn friction to protecting a key that people will roll the dice with the safety of their users rather than confront the fact that the correct engineering solution is an outage-inducing all-hands-on-deck rekeying.
What blows my mind about this is, at 24-hour expiry, in a medium-sized application, you're going to have machines needing to refresh keys practically every hour of the day; your CA will need to be available 24/7. At that point, you've basically reinvented Kerberos. Ops teams fucking hate Kerberos! And for all that effort, you still face fleetwide updates any time something sketchy happens anywhere.
I like mTLS for things like "we set up a Consul cluster; let's make sure just the machines that use Consul can reach it". It works fine for that. I don't think it's a good idea to take it much further.
If you have a perimeter with well-defined entry points (e.g., a bastion for HTTPS and SSH) then you can lock people out at those entry points. Mutual TLS is defense-in-depth, and how you implement least authority for services inside the perimeter.
I know that CRL is literally active revocation. If you want active certificate revocation you can literally do active certificate revocation. Whether you need it depends on your threat model. And it’s not actually that hard in many scenarios. It is harder than not doing active revocation, and if you find yourself needing it to satisfy your threat model you probably should consider other alternatives for client authentication.
I’m not telling people they should use lots of mTLS but not worry about revocation. Any content I have written that discusses short-lived certificates is very careful to point out the fact that a compromised certificate will be considered valid until it is revoked. Most people who are using our open source toolchain are actually using it for internal server auth TLS, for narrow use cases like issuing certificates for database authentication or VPNs, or for mTLS to satisfy the sort of threat model I’ve described here.
We haven’t implemented anything around active revocation yet. But it is planned, because it is necessary in many scenarios.
If you’re worried about keeping a simple web service running 24 hours a day then you shouldn’t use our stuff. If you’re far enough along to need our stuff, you’ve figured out how to keep a web service running... you have to keep your own web services running 24 hours a day otherwise your microservice system is already going to have availability issues.
You don’t need revocation because the credential is useless outside of the perimeter. If you find a malicious insider or a compromised service, you fix the service or lock the insider out of the perimeter. You’re glad that you had client certificates in place as it reduced the scope of the compromise. It would be nice if you could actively revoke any compromised credentials, but it’s not critical because you’ve locked out the offending user and/or fixed the compromised service.
I think we're probably at an impasse.
Credential rotation is good security hygiene. To suggest otherwise is malpractice. Our toolchain makes certificate rotation trivially easy. Why not rotate frequently?
Hopefully the threat model stuff made sense. It still feels like you actively want to disagree with me, and I’m still not sure why. But I agree that this is starting to feel unproductive.
I do appreciate the discussion. I understand your position on client certs better now. Your concerns are valid.
Maybe one day we can discuss over beers or something. It feels like that would be the right atmosphere.
Furthermore, it’s not arbitrary. Credentials leak and services come and go. Having active keys around that aren’t in use is worse than not having them around. If someone accidentally commits a key to a GitHub repo or something, it’s nice to know that key will only be useful for a little while.
If you still want to rotate less frequently, change the default.
This really has very little to do with the topic at hand, so I’m not sure why we’re debating it. Do you want me to change the default certificate lifetime in step-ca? What do you think it should be?