What if your Pods need to trust self-signed certificates?
blog.alexellis.io
blog.alexellis.io
This means that trust store updates are not dependent on the OS vendor, they’re consistent between native/Java/etc, and updates don’t require rebuilding container images.
When certificates rotate the system builds a new image which gets pulled down by watchtower, which then in turn handles dependency management and restarts things as needed.
I wanted my homeprod setup to be as hands off as possible while still allowing easy management. Each physical host is running Alpine. During provisioning I install docker, Tailscale, and manually start a "root" container that runs[2] docker compose and then starts a cron daemon. The compose commands include one or more "stack" files and are generated based on a yaml file listing the stacks for each host. Watchtower runs with a 30 second cycle time to keep everything updated, including the root container. Adding or updating services means committing and pushing a change to the root container repo, then CI builds and pushes a new image. Watchtower picks up the new image and restarts the root container, which re-runs Compose which in turn starts, stops, modifies, etc anything that's changed.
For certificates, I tried a number of different things but ultimately settled on the method I described earlier. The purpose of the container image is to 1) transport the certificates and install them in the right spot and 2) be updatable automatically with Watchtower.
Certificate changes are very similar to the root container, except the git repo self-modifies upon renewals (yes I keep private keys committed to git, it's a homelab, it's really not a big deal).
[1]: https://www.petekeen.net/homeprod-management-with-docker
[2]: https://github.com/peterkeen/docker-compose-stack/blob/main/...
I build a bundle (though I may just move to trust-manager [1]) and replicate it into all namespaces with kubernetes-replicator [2], and then I can annotate any pod with
[0] https://github.com/microcumulus/ca-injector
trust-manager also supports pulling in the Mozilla trust bundle which most Linux distros (and therefore most containers) use!
Handling trust of private [2] certificates is done poorly generally across many orgs and platforms, not just Kubernetes. There are lots of ways of shooting yourself in the foot - particularly when it comes to rotating CA certificates. I think there's a lot of space here for new solutions here!
[1] https://cert-manager.io/docs/projects/trust-manager/
[2] I try to avoid "self-signed" in this use case because its literal meaning is that the certificate signs itself using its own key, which is what root certificates do. The Let's Encrypt ISRG X1 root certificate is self-signed but it's definitely not what I'd call a 'private CA'; see https://letsencrypt.org/certificates/
An internal CA is also, while achievable with cert-manager these days, not exactly trivial to set up.
…it's not exactly hard, though? Install cert-manager, and it's all of 3 resources? 1 self-signing Issuer, one Certificate (the CA cert), and then an Issuer that issues certs under that CA cert.
Then any service that wants an internal cert adds a Certificate issued by that Issuer.
The annoying part is getting it mounted in every pod that needs it and needs to consume it, which is the point the original article made. Node has its own way of doing it, Java has its own way of doing it, etc etc and they all have their own little ideosyncracies that you have to worry about.
Yes, I suppose, in that it does need to be available to the pod.
This is one of those problems you're solving regardless of Kubernetes. I think a few lines of YAML to say "pull this secret into this Pod" is pretty acceptable to trying to figure out how it ends up on a VM with Ansible, or do I pull it from SSM, etc.
> Node has its own way of doing it, Java has its own way of doing it, etc etc and they all have their own little ideosyncracies that you have to worry about.
Yes… all the world doesn't use the same HTTP/TLS libs. That's just dealing with encryption, period, though; you've got this regardless of whether the CA is internal, or you're on k8s, etc.
Java, though is particularly annoying, most of its ecosystem stubbornly refuses to use the same formats for key material as the rest of the world does, necessitating a translation into its format. That is annoying, but nonetheless that's a Java-ism.
I think there’s a use case to be made here for running internal services on public certs. Obviously it won’t work for every use case but if your services are low traffic I don’t really see an issue of doing that and could save end users a lot of headache as they don’t have to worry about anything and you as the operator only have to worry about automating the DNS challenge
Your customer sees one cert at the LB level. LB talks to ingress which also has own cert by CAM. And any other pod on the cluster uses mtls via ebpf.
I think the better approach is to just use certificates from public CAs on private networks. There's some big issues to address: how does the public CA verify internal servers which don't permit external connections, is leaking internal subdomain names a security issue? But I think these are solvable.
getlocalcert.net[1] is my project to simplify the process. Currently you can register a subdomain and use the ACME DNS-01 protocol via my site to issue certificates from from providers like Let's Encrypt. All you need to issue a certificate is a getlocalcert API key and the ability to connect out-bound to the Let's Encrypt API and the getlocalcert API. It supports a couple ACME clients and it should support cert-manager, mentioned in the article, but I haven't had a chance to test and properly document it[2]. I'm not doing anything that other subdomain registrars or DNS providers can't do, but I'm trying to address the public-CA-on-private-server niche as best as possible. Longer term I'm looking to add bring-your-own-domain support and other improvements.
[1] https://www.getlocalcert.net/
[2] https://docs.getlocalcert.net/acme-clients/cert-manager/
You can also use MutatingWebhooks to inject initContainers or side cars to achieve that pattern.
MutatingWebhooks increase the runtime complexity since you're jamming more config in but they can handle vendor/3rd party software.
Depending on the CNI, insecure TLS flags might not actually be that bad. There are some CNIs floating around that can handle encryption. In addition, software defined networks like AWS might not offer encryption but they do guarantee messages are authenticated so you can trust the source IP
It could also save on container size theoretically. Implementing TLS is a daunting task and requires many kilobytes of code. Implementing HTTP is easier task. So you can strip TLS implementation and achieve lighter containers.
It adds complexity with the side car but usually you get standardized RED metrics regardless of if the service implemented them.
You can also have the sidecar do connection pooling and maintain persistent or long lived TCP connections which ends up speeding things up without doing that at the app level.
That's a horrid idea, why would you do that ?
But the solution is internal CA, not self signed certs that defeat near-entire point of encrypting communication.
First, cost. Any CA that issues unlimited certificates will charge tons of money. Free CAs like letsencrypt do have rate limits that we would frequently hit with autoscaling environments, CI jobs, and such.
Also, CAs require the use of certificate transparency logs. Which will expose your internal infrastructure data to the public. It will, by exposing autoscaling data, also expose financial data (at least in hints), e.g. by showing that last christmas, your scaling peak was far higher.
And external CAs are a security risk because you need to provide firewall exceptions and/or transfer mechanisms for certificates into your internal infrastructure that you would usually want to isolate.
Lastly, an external CA is an availability risk. Should your external CA be unreachable for some reason, you might not be able to run any CI jobs or auto-scale-up your infra.
Before this (d)evolves into a zero trust, security-by-obscurity discussion - some auditors won't certify you in some edge cases related to this, and you may be operating in a regulated sector where such a certification is necessary. Just because it doesn't impact your use case, doesn't mean this is the case for someone else.
- Letsencrypt and friends don't give you whatever purpose certs you would want. I.e File Encryption, Code Signing and friends, individual client certs
First, cost. You're not just going to slap the root CA onto the network drive and let the developers have at it - you're going to need infrastructure to keep the key safe and handle automated cert provisioning and suchlike - that's going to need people to maintain it. And it'll be an important part of your infrastructure, so you'll need enough experts you can maintain round-the-clock support.
Second, it reduces your security because your users will inevitably learn to ignore certificate errors.
Thirdly, you'll never stop the certificate errors. Oh, you're going to set them up for both Chrome and Firefox and Edge and Safari on Windows, Mac, Linux? Oops, you forgot Android and iPhone. And your CI system. And Java and Docker and Git. And the network printer and the electronics team's network-enabled oscilloscope. Think you've covered everything? Surprise, Slack is distributed in a Snap now, it's generating certificate errors.
Just have your cloud provider take care of it. Not in the cloud? A wildcard cert on your load balancer will get you 99% of the value with 1% of the work.
And talking about security risks, wildcard certs are especially dangerous and should be forbidden from ever existing. They just lead to "copy it everywhere"-keys that, sooner or later, will leak. And that won't be revoked or replaced, because of course everything will break at once.
Oh, and the certificate errors will also come with external CAs. Chain too long? Error in some browsers. ECC signature? Error in some browsers. Chain with different paths? Error in some browsers. 4096bit certificate somewhere? Error in some browsers. Two different valid roots? Error in some browsers.
Recall the post that started this subthread:
> Having to communicate with outside is kinda overkill if you just want to have container A talking to container B.
The article here is about the same.
#2 and #3 in your post don't apply; we're not talking about browsers or end users at all here. #1 may apply but I think you're overstating it; Active Directory Certificate Services takes care of all that. Remember that you don't have to follow the CA Baseline Requirements as a private CA. It's harder to get rid of an ADCS PKI than to set it up.
Y'all don't have internal tools implemented as webapps? Self-hosted version control servers? Nexus? SonarQube?
Oh, I'll agree that you can outsource all that stuff if you want to - but any business with that philosophy would surely also outsource their certificate provisioning. Especially considering how easy and cheap AWS make it.
> although I'm confused about the mention of network printers and Slack
Do you not want graceful handling of internal URLs when mentioned in slack? Such as previews, image unrolling etc? Do you not need a certificate for the internal file server your scans upload to, and so on?
Stated another way, I believe you are saying "don't use internal CAs for things you'd otherwise use public certificates for" but what we're saying is "use internal CAs for things you'd otherwise use self-signed certificates for". I believe both statements are correct but we weren't talking about the first thing at all until you brought it up.
The goal is not to make cert errors go away.
The goal is to establish secure connections between parties, with the smallest possible trust boundary and attack surface.
You're expected to design around that. Deploying should never create a new certificate, those should live in secure storage and get deployed when needed.
The main rate limit also doesn't apply to renewals, so you could potentially issue N*50 domain names on your Nth week using Let's Encrypt.
"Renewals are treated specially: they don’t count against your Certificates per Registered Domain limit." - https://letsencrypt.org/docs/rate-limits/
If you need to issue certificates for new domains at a higher rate, you're very likely a large company that can afford to pay some money for any excess certificates you need. Failing over to ZeroSSL (zerossl.com) on rate limiting should be an easy engineering task since both use the ACME API.
It's less secure than a self signed cert. OK you can't just sniff traffic and decode it thanks to Elliptic Curve, despite having the cert+key, but you can MITM just as much, and not throw an error.
With a self-signed if someone has accepted and cached the self-signed cert you can't MITM them without throwing an error.
So, self-signed certificate is the easiest and fastest way to solve this problem.
PS. Also, custom CA certificates don't work well with browsers -- browsers like to cache them and won't let go of them easily. This has adverse effects on automated UI testing with Selenium. I.e. say, you need to modify something in your cert (usually something like a domain name) and then the browser will refuse to accept it because it remembers the old one somehow...
Another problem is that errors related to certificates are very hard to debug. The popular libraries (eg. essentially, OpenSSL) are trash when it comes to error reporting. You never get clear error messages, nor a suggestion on how to fix errors. The whole thing lives in obscurity, and CAs are even worse documented than the process of creating self-signed certificates. So, people looking to solve a problem choose a solution they can find, instead of the best solution.
There are many, many good reasons to have your own PKI that are not tech-debt related. PKI use-cases are not limited to internet-exposed server certificates.