TLS certificates for internal services done right
tuxnet.dev
tuxnet.dev
Use DNS validation to allow these internal services to pull ACME certs. There's so much less headache, long-term.
Split-horizon DNS (and the tedious make-work it can create when you start needing to mirror public-accessibly records in the private DNS) has always been something to aspire to move away from in my experience.
Services themselves are constrained to the server, bound to listen on 127.0.0.1 only.
Key is only available to the reverse proxy.
https://community.letsencrypt.org/t/dns-persist-01-deploymen...
I wonder if the interim version has been rolled out to some CAs.
Once it's supported I think my next iteration will be DNS persist + internal ip addresses on the public zone.
Thank you all for comments and feedback! It's cool to see real interest in this blog post
... and it is even already documented at https://github.com/acmesh-official/acme.sh/wiki/DNS-alias-mo...
With this setup, I don't have to grant 3rd parties DNS access.
I actually made a webhook that allows per hostname API keys to wrap dnsimple because they only had per zone keys and I didn't want each VM to have access to the entire zone. These challenges would have solved that by allowing the DNS record automation to pull record values from the VM instead of having the VMs push values.
I think someone told me dnsimple might have more granular keys now but I haven't checked. Iirc we have the same concern with external-dns at work (some things need subdomains on the TLD but we don't want to give external-dns access to the whole zone so we usually cname the TLD subdomain to a per environment zone external-dns is allowed to update). With this, we could have the same pull based setup that applies arbitrary rules to decide if a requested record should be created.
I think the main takeaway is allowing pull instead of push model
Could be achieved with cnames but it's an extra layer of indirection to deal with and doesn't fully solve the "semi trusted 3rd party" case
Lets encrypt supports IPv6 for validation.
You need a DNS provider which supports API calls (I use DNSimple) but the core is all very straightforward.
To prevent having to include DNSimple authentication on the client's internal server I have a small API server on the web which does the Acme work.
https://github.com/acme-dns/acme-dns
CAUTION, though, the last time I downloaded a binary release, ClamAV triggered on it, so I kept my old version which worked. I was using the 1.0 series (without any problems!), and now it seems the project has picked up development again with a 2.0 series.
Library/CLI that speaks the API of several dozen DNS providers so you don't have to re-invent the wheel:
For this edge case, the issue is not with DNS; it's that the router is not configured to allow internal clients to access the network through the public IP, vis-à-vis hairpin routing.
With one of self hosted services being Adguard-Home can do both ADs blocking and internal DNS... The public DNS records for your "internal use only" domain remain empty.
Saying something isn't bad without pointing to right direction makes my insides hurt a little bit.
> Use DNS validation to allow these internal services to pull ACME certs...
The major ACME clients support DNS validation. If your DNS host doesn't have an API that your ACME client supports either get a new ACME client or a new DNS host.
I thought you have a solution to overcome not having to do split DNS where I define public IP in public DNS and internal IP within some local hosted DNS.
Like how to make so that when connecting from inside local network the router recognize that by connecting to public IP, he has to route it back onto some local IP address?
For a homelab situation split-horizon DNS is just fine. You're going to have minimal duplication of records from the public DNS into the private DNS.
The canonical frustrating "bad split-horizon DNS" world I've seen, time and again, is a corporate network with a MSFT Active Directory named the same as the company's public-facing web property (e.g. "example.com" rather than "ad.example.com"). This creates the need to duplicate all the public-hosted resources into the internal "example.com" zone (and keep them updated when records inevitably change on the public Internet). It's make-work for no practical upside in that example.
Totally agree on the AD part - luckily we had an option to migrate to different domain and name our AD correctly :)
This means that I can always use public DNS servers like 1.1.1.1, 8.8.8.8, nextDNS etc
This is not "done right" by any stretch but it's extremely low effort to set up and has never once failed me, unlike countless complex meshy things.
Removing attack surface is better than trying to hide it.
The juice isn't really worth the squeeze for the token spend any more than it was worth the human energy.
I'd prefer this over split DNS, any day.
I use the form of hostname.int.example.com for everything inside my home network. None of which is accessible to the outside world. I use LetsEncrypt with DNS validation to get the certificates.
If it does then you don't have to mess with your public DNS whenever you want to add or renew certificates for home machines.
I'm using the free DNS my registrar provides, which doesn't provide API access unless you upgrade to their paid DNS service and so if I could use a local DNS server for the ACME challenges for the home network I could pick one that is friendly to automation.
I use Cloudflare for DNS and it is free to use the API.
Note that int is a valid TLD:
Lots of folks were using "dev" as a sub-domain which was fine until ICANN decide to give Google a TLD:
* https://en.wikipedia.org/wiki/.dev
So if you generally had "search example.com" in you resolv.conf, and were in the habit of having "web01.dev" in places, behaviour may have changed if you were suddenly on a machine that had the "search" line missing (or something else).
That won't prevent me from getting a ticket saying "the network is down".
I always use FQDNs for everything.
I did set up tailscale, way back. After using it a few times to test, it failed me when I really needed it (I was out of the country and it failed - can't remember exactly what went wrong but it wwas 100% 'in my tailscale account'). I immediately dropped it and went back to OpenVPN (shit but reliable) before building my current setup.
Several Rust libraries also tend to default to a predefined set of certificates (which makes sense for libraries supposed to run on bare metal as there are no system certificates there, but that's not really a problem on most Linux installs). I make it a point to always use the native OS roots in the code I write, but unfortunately that's not universal.
Having to bind-mount certificates inside of docker containers is also always an annoyance I forget about until I see the first TLS errors in the logs, but that's by design and probably a good thing.
A lot of non-language tools bring their own certificate bundle as well, like uv, git, curl and Firefox. (I think they might all be the same Mozilla bundle even).
However, it seems like the situation here has improved slightly: git can read Windows certificates now once a flag has been set, Firefox has a flag for Windows and macOS, and it supports p11-kit. curl can be built with Windows/macOS support and respects OpenSSL environment variables. uv can be configured to use system-certs.
And obviously any VM or container has to be set up separately. That often includes stuff like pipelines in forges. As mentioned in the sibling comment, at least here it's definitely by design.
Python? The widely used 'requests' relies on 'certifi' and skips the OS store - but 'pip' on the other hand does use the OS store.
A tool that might use java, like a database or IDE, means it might have its own store.
Node? Make sure you set NODE_USE_SYSTEM_CA=1
Firefox and Chrome AFAIK both have their own stores.
Building a Docker container? That's intentionally isolated from the host, of course your container won't inherit the OS trusted CAs. Running a VM locally? Same.
Installed something using snap? The container-like isolation means it won't pick up the OS trusted CAs.
And of course you need the certs set up right on your cloud servers, your CI servers, the dozen different smartphones the mobile team uses for testing.
* with some notable root certs that I have… questionable… trust and confidence are not simply controlled by certain state actors.
I have to secure the CA's key, but I also have to secure all the keys for the certificate it signs, both being a similar level of challenge.
For personal use, or for very small organisations, using a passphrase-protected Yubikey as a "cheap HSM" should suffice.
* https://www.yubico.com/products/hardware-security-module/
See also perhaps less expensive option:
It looks possible, on paper, if you don't poke at it too often.
Or even a more extreme example: https://crt.sh/?id=27555237869 (sorry for any possible crt.sh downtime) - the domain name in question never existed in public or private DNS by itself. It is used only for a WPA3-Enterprise network, as the CN that WiFi clients expect to be present in the RADIUS server certificate, but never resolve. In the public DNS, only the "_acme-challenge" TXT record exists.
The point was that you can obtain a certificate for a domain name without creating any records other than the _acme-challenge TXT record. I.e., that the domain might be completely empty all the time except for this record.
I also use Tailscale so I configure my DNS to use my Tailscale IP addresses. If you don’t want to expose them on a public DNS server you can add them only to an internal DNS server.
https://www.eff.org/deeplinks/2018/02/technical-deep-dive-se...
See also perhaps DNS aliasing in case you are not able to dynamically update your 'primary' domain, but can update a secondary or sub-domain:
* https://github.com/acmesh-official/acme.sh/wiki/DNS-alias-mo...
So if "example.com" is control by Corporate IT, and they don't want 'random' folks fiddling with it, then you can create a "dnsauth.example.com" and point the dns-1 challenge record from "…foo.example.com" to "foo.dnsauth.example.com" (or a completely different domain, like "…example.net").
There are DNS servers written strictly focused on this use case:
* https://github.com/acme-dns/acme-dns
Also code that handles a bunch of DNS provider APIs so you don't have to roll your own for ACME client hooks:
I get it, I could just do *.mydomain.com and slap that wildcard cert everywhere, but it's still in the public logs..
- Don't use split DNS. Don't use any special internal or dev domain. Leave it to your infrastructure to route/NAT those public IPs to your internal network.
- Don't use the HTTP-01 challenge. Use DNS-01.
- Don't run your own internal CA. Use Let's Encrypt. If you care about name leakage (CT Logs), use wildcard certs. Use a central reverse proxy/load balancer for termination.
If you do that anyways, you could also use something like oidc authN/authZ on the reverse proxy level and just expose it to the internet.
You dont even need to self host the oidc idp, you can use Google/Github or even something like ATProto
Split-horizon DNS for a publicly usable domain is almost always a bad idea, but running your own ACME server is pretty easy (maybe 10 lines of Caddy config) and using an internal domain (an actual one, not a randomly picked TLD you don't think exists yet) solves the problem pretty easily. You'll want a safe backup for your root certificate private key, of course, but that's pretty much all you need to really worry about.
So what's your solution when you have a wholly private service that will never have a public v4 address, nor a publicly routable v6? How do clients get the address for a nice domain name without the addresses in public DNS?
I use acme.sh with DNS validation, and use common domains that have both public and private services on subdomains. I use split horizon so private.domain.example resolves only on LAN and VPN, and public.domain.example resolves everywhere, but the address changes depending on the network one is connected to.
They don’t. You put the addresses in public DNS.
If you have a heavy enough tech stack to run fully internal services, than you can also run an internal DNS service (even pihole is enough) and load internal only entries there.
Or add everything to your hosts file if you have a central config service.
>Wait…what? You have a fully internal service and you dont have an internal DNS server?
You can't do split horizon without an internal DNS service. By definition it's a non-public service that gives separate results to public.
So what's your solution when you have a wholly private service that will never have a public v4 address, nor a publicly routable v6? How do clients get the address for a nice domain name without the addresses in public DNS?
Which, if you claim you use private DNS, you know you can indeed solve by only having the private records sit on your private DNS service that can only be hit by clients on the internal network. Without the architectural bad practice of implementing split horizon.
https://en.wikipedia.org/wiki/Split-horizon_DNS
I could not have been clearer that the private DNS servers I use provide different results from public, for the same names, which is the split. Not all private DNS servers are split horizon, and nothing I wrote implies that that's my understanding. If that's what you've concluded, then you've either failed to comprehend what I've written, or you're a troll.
But at least we both agree private dns =/= split brain. :)
1. Register a domain ("server.com") and put it on some public DNS that can do DNS validation with acme.sh.
2. Use DNS validation to get a certificate on your domain from Let's Encrypt. You can just grab a wildcard one ("*.server.com").
3. CNAME all of your services on a public DNS to an internal address ("email.server.com" → "server.internal", "plex.server.com" → "server.internal").
4. Resolve your internal address on a local DNS server with an A record ("server.internal" → 192.168.0.123). This can often just be done on your router.
Since you use DNS validation, you just API keys for your public DNS service that acme.sh can use. No need to have any VPN network interfaces for getting your certificate. Your wildcard certificate also doesn't leak any details about your services.Would it work if a user's device that is already connected to the VPN, but has custom DNS override to say 8.8.8.8 ? How can I allow my users to be able to use 8.8.8.8 DNS override and still work seamlessly?
Specifically grafana is nice to be able to see on the phone, and split horizon DNS and corp VPN is a hassle, to say the least, on phones.
I bet you can do it with HA-Proxy, but I use https://github.com/ThomasHabets/sni-router
Most browsers support trust on first use for leaf certs
And later if something changes, then they can do the whole DOING SOMETHING NASTY! thing, which is effectively the experience today
Using a browser in an air gapped environment is so much more pain than it should be.
You can also invert and have k8s cronjobs provision the generated certs into other infra
With this setup, you don't have to worry about the RHEL certbot snap updating to a broken version which gets blocked by SELinux...
But on hosts you control, you should absolutely provision them with an identity and join the local CA. You're going to need it for a multitude of other reasons.
In that case there's no need to validate anything as names, dns records, certificates and anything else should already be in place.
Unless you enjoy that sort of thing.
And then the issue is protecting the private key of the issuer and monitoring certificates (it's a good idea to do that anyway).
acme.sh --install-cert -d grafana.tuxnet.dev --key-file /etc/ssl/private/grafana.tuxnet.dev.key --fullchain-file /etc/ssl/certs/grafana.tuxnet.dev.crt --reloadcmd "systemctl reload nginx"
Then you can ditch your custom cron and let acme handle everything on its own, as intended.
This introduces all sorts of aberrant behavior. Most notably that clients will get different responses depending on their query and route that you now have to test for. It also is a “silent” deviation for audiences that may not be aware of.
I’m not sure what problem that is trying to be solved here? That the OP wanted internal users seeing a different site than external users? Or using a different route in? In both those cases the correct user behavior is to use a different DNS record.
You now also have to build infrastructure to distribute the wildcard from (presumably) central place where you generate it to all the different places where it is desired.
And hope the wildcard's private key does not leak from one of myriad of places it now lives.
Leaking is an issue but we're talking about internal services too.
It also forces you to use DNS challenges, which means you don't require the publicly route able address.
Personally what I do is: - Predominantly use wildcards
- Run a tiny coredns instance with my A records for a internal subdomain of my normal domain
- Configure tailscale to use the coredns resolver
- Run two haproxy instances, one for internal services, one for public facing. The public facing one can't route directly to the internal services.
When even the obscure DNS provider I'm using is supported for DNS challenges, I really don't see much upside to using HTTP challenges anymore
This made me consider my haproxy architecture. I use haproxy acl rules and the use_backend directive to enforce security policy and routing. It denies http requests based on source ip address per service (host header)
I think a good middle-ground is to use a single haproxy but two frontends; one binds to the public interface and one for private. haproxy is now no longer responsible for routing, the router is.
* "internal services" = on a single server that is publicly routable
Not quite done right.
I do have a luxury of all the homelab VMs being rebuildable via IaC, so I've just injected CA trust at that step.
The biggest PITA so far were 3rd party docker images, each with its own way to inject custom CA.
iOS devices were surprisingly easy to handle.
I did it multiple times, most recently using YubiHSM as root key store for offline enterprise root CA.
There's no single trust store: the OS has one, Firefox, Java (cacerts), Python (certifi), Node, Go containers, all your Docker images, ...
Failure? People just toggle TLS verification off.
Easy to do in a one-man shop, but almost impossible in any big company. Try grepping for verify=False / -k / InsecureSkipVerify in any fortune 500 and you'll find plenty
The only reason is laziness.