> So, this is not a bug and all is working as intended.
Caddy folks had better never restart the caddy service (or server) while LE happens to be down, even if you already have a valid cert!
Update: Mholt pushed a change where caddy only refuses to start if the cert is expiring in 7 days or less. https://github.com/mholt/caddy/commit/410ece831f26c61d392e0e...
Having the server be unable to start through circumstances outside of the system's control is just such a huge no.
edit: looks like the dev added a fix to only refuse the start if the cached certs are dangerously close to expire. that satisfies me and I'll be continue to be using caddy.
Am I missing something?
The way Let's Encrypt works, it makes a lot of sense to have the functionality be part of the web server.
That's an odd thing to say in this particular conversation thread...
Why would you want to tightly couple your webserver to the availability of another service provider?
No, with Caddy you have several vaguely related pieces glued together with superglue.
> The way Let's Encrypt works, it makes a lot of sense to have the functionality be part of the web server.
I think this thread is pretty much proof that that approach will bite you in the ass.
What bit me here is the fact that I'm running alpha software instead of a battle-tested web server; I'm doing so willingly, with full awareness of the risks that that entails.
Drawing the conclusion you did from the variables at play is shortsighted. If anything bites people in the ass, it's prejudice and shortsightedness. I wouldn't want you handling my ops/infrastructure.
This wasn't caused by a bug. This was a deliberate decision to fail to start if the certificates on-disk were <= 30 days away from expiring and the CA can't be contacted.
> Drawing the conclusion you did from the variables at play is shortsighted
Using caddy is the web-server-stack equivalent of "putting all your eggs in one basket". If one thing about it isn't working the way you want, you have to either a) replace it completely or b) work out how to disable the bit that's not working how you want, and replace that part of it.
> Drawing the conclusion you did from the variables at play is shortsighted
- People use a piece of software that serves as both ACME TLS certificate client and web server
- Said software by design won't start if the CA can't be contacted 30-days out from expiry
The conclusion I drew is that such integration leave the operator with less control than if they followed a separation-of-concerns approach, and left web serving to a web server, and TLS certificate renewal to an ACME client. The former doesn't need to care about how old the certificates are, just use what it's given.
There should be at most a warning but it should start. Otherwise you end up with an external dependency that can cause your web server to not start through no fault of your own.
Ahem. https://github.com/mholt/caddy/issues/1680#issuecomment-3026...
Emphasis mine:
> So, this is not a bug and all is working as intended.
Failure to start if a CA is down is a "sane default" ?
1. The cost does not scale with the amount of instances since it is the one-time cost to create the configuration package.
2. If you decide to go for Caddy instead, you'll have to spend the same time, if not more, learning Caddy.
If I am hosting 5 sites on caddy, and add a 6th one, I restart the server. If the 6th site doesn't work (for example, if DNS didn't resolve for lets encrypt), the other 5 sites which were working before the restart, all fail to start as caddy completely crashes.
This is basic resiliency you'd expect from your web server. Why should the other 5 sites fail to start if their configs are completely valid?
No, you don't. You reload the server, not restart it. Restarting a web server should only be required if you get an upgrade for it (or for OpenSSL etc.)
notifies :reload, "service[caddy]", :delayedCaddy will just refuse to even handle them.
Every single other server on this planet handles them properly, but caddy doesn’t – and mholt considers that working as intended.
Try out: https://www.google.co.uk./ https://www.microsoft.com./en-us/ https://www.amazon.com./ serve the page directly; https://www.facebook.com./ redirects to the relative domain
and then https://caddyserver.com./ (That said, traefik is equally dumb, as seen with https://traefik.io./ )
Separation of concerns means you are in control, and using separate layers means you can swap one out when a vulnerability/show-stopper bug is discovered.
What exactly do you do when your look-ma-no-hands server won't even start?
Edit: maybe "all-things-to-all-people" was the wrong term to use here.
From https://caddyserver.com:
> The Most Beloved Server
They started the hyperbolic claims, not me.
You're claiming "caddy does everything". As opposed to what? If you're running apache or nginx, your server does far more than caddy, so you're quite simply mistaken.
Serving content over http(s), and obtaining TLS certificates are two very different tasks.
> If you're running apache or nginx, your server does far more than caddy
Far more, that is directly related to serving content over http/https.
Except that with let's encrypt one actually needs the other.
This fix gets almost everybody where they should be, the next time the same thing happens (and it will) Caddy isn't a problem for three weeks, which is definitely enough time. Meanwhile we're going to see the same Apache crappiness for OCSP again each time until someone over there finally snaps out of it and asks someone who actually knows how OCSP stapling was supposed to work.