Reading it over, I think we mostly agree on the facts. It's easy to do mTLS and x509 wrong. The question, then, is what's easier / more secure: doing mTLS/x509 right or doing something else? I think that's somewhat subjective: it depends on your requirements, your environment, and your skillset.
One point that I'd like to reiterate is this: if you want a consistent cryptographic solution that works everywhere, TLS is pretty much your only choice. You could use something else for client authentication, but you probably still need TLS.
As a strawman, here's a sketch of how I'd recommend doing TLS in a microservice system. I consider this "right" for most garden-variety microservices-in-cloud scenarios and don't think it's particularly hard to do. Most of this is already implemented in https://github.com/smallstep/certificates:
* Deploy the root cert via automation (so it's quickly rotatable) and/or keep it in a managed HSM/KMS. You might harden root rotation a bit by signing your new root with your old root. But, generally, trust config management or container orchestration to push root(s) (you already trust it to push code and secrets). Root rotation (and, thus, bulk revocation) is now as fast as secret rotation (secrets are generally pushed the same way).
* Issue short-lived certificates per logical entity. If it gets a box and a name in your architecture diagram, it should get an identity and each instance should get a certificate. Use domain names and email addresses that you control for names. Keep certs simple: one SAN. Certificates bind a name to a public key. That's it.
* Automate certificate issuance. ACME can work for this, but there are other options (single-use tokens issued by config management, cloud-managed instance identity documents or service accounts, an existing device certificate issued by a manufacturer, etc.)
* Automate certificate renewal. A simple mTLS HTTPS request works for this. This is easy to implement and easy to scale out with multiple intermediates. "Revoking" a certificate just marks it as "not renewable". To reduce risk of outage, in this architecture, it's safe to renew an expired certificate as long as it's not revoked (ACME-STAR basically does this, but it's push instead of pull).
* If you really need active revocation, fine. One good solution is to push CRL to a cloud storage bucket. Short-lived certs will keep your CRLs small. If you need to do a mass rotation, rotate roots (push new root, wait for rotation, pull old root).
* Use secure NTP for time.
* Index issued certificates. CT (trillian) is cool if you want to be fancy. Your existing database or SIEM also works. zcertificate can parse x509 and output a JSON representation of a certificate that you can map to something like an Elastic Search schema: https://github.com/zmap/zcertificate
I want to respond specifically to your first and final points.On your first point: I understand that in theory an attacker could slip a request across a secure channel, and binding authentication to a request could in theory prevent that. I don't understand how that's likely to happen in the context I'm thinking of here. Which may be different than the context you're thinking of. So let me clarify.
Suppose I have `<end-user> -> <service-a> -> <service-b> -> <database>`. Let's focus on `<service-a> -> <service-b>`. I don't see how using end-to-end mTLS, terminating in `<service-a>` and `<service-b>` application code, would be any more vulnerable to this variety of attack than an HTTP Basic header like `Authorization: Basic base64(service-a:password)`. Surely, the logic in `<service-a>` is simply "insert HTTP Basic header into requests on their way out to `<service-b>`". It doesn't matter if we're authenticating the request or the channel. If you're able to smuggle something malicious into that request, it's gonna get sent over to `<service-b>` with proper authentication attached.
Are we talking past one another? Are you trying to make `<end-user>`'s authenticated identity carry through `<service-a>` to `<service-b>`? If that's the case, then yes: I see what you're saying and you shouldn't use mTLS for that. I'm not sure if there's a term-of-art here, but I call this "end user identity propagation". You need something like a top-of-stack ticket service (a bearer token) for that. Or, better yet, macaroons. I consider those two separate things, though. mTLS is for authenticating your immediate peer. For end-user identity propagation mTLS is a poor choice.
On your final point: you could, in theory, express claims in x509. I'm sure you're aware, but it's been tried before (e.g., SPKI/SDSI). However, I agree that, unless you really know what you're doing, x509 is too complicated for that. Don't do it. You'll likely screw it up. If you're parsing x509 and ASN.1, you're doing it wrong. If you're processing strings that you've extracted from a certificate, and you're not in the habit of writing your own formal languages, you're definitely doing it wrong. Just put a flat name in a SAN. The only thing you should ever need to do with that string is an exact string comparison. If you need to know roles or groups or some other metadata look them up in a database.
(Or use macaroons)