My whole reason for switching to LE was that it was announced as an automatic way to upgrade certificates. So it required less of my attention than before. Or at least, that was the spiel.
My whole reason for switching to LE was that it was announced as an automatic way to upgrade certificates. So it required less of my attention than before. Or at least, that was the spiel.
And that is why we run it. But we still monitor.
First we pipe stdout things to syslog via logger(1) to record each run in cron; however stderr is not redirected so we get that sent to us. Further, in our hook scripts, we have an "apache2ctl configtest" when a cert is renewed to verify that Apache (or equivalent) can read the files correctly, then we do the restart/reload. The configtest is sent to stderr as well so we can track when renewals happen and when there's an issue.
On top of all that, we don't actually trust our ACME client(s) to work properly all the time, so we check the expiration time on our certificates. We're monitoring the HTTP(S) service anyway, so adding another check doesn't add much cost.
And all this runs in the background silently and it only becomes top-of-mind really when it shows up in our service alert dashboard.
The difference for me is that Let's Encrypt has required this kind of manual intervention more often than the previous status quo. Previously, I would put an item on my calendar to renew the free StartSSL cert, would do it on schedule, and things would generally go off without a hitch.
For me, the automatic renewer has fallen apart multiple times (for several reasons, including failing to reload nginx after renewing the cert), more than once a year on average, and worse still, they don't fail on any kind of regular schedule. I'm not a pro, just someone who wants a cert for my personal site, so I don't have any complicated monitoring system set up. This means I have to regularly remember to check up on certbot and see if it renewed the cert this time.
If spending $x per year solves the problem for you, then go for it, but it does seem odd to me that you're having so many problems.
As for monitoring, a simple thing to do is to install some well-tested, pre-canned check scripts (e.g., from Nagios):
* https://packages.debian.org/search?keywords=monitoring-plugi...
and have them kick off via cron on a regular basis. A non-zero error code will cause an e-mail to be sent to you.
While we use Nagios (for now) to monitor from a central host, we also leverage them in things like keepalived for running HA clusters to trigger fail-overs. No sense re-inventing the wheel.
Criticism of the "official" client was also very loud from the start.
I do agree that LetsEncrypt has been a constant source of maintenance problems over the years. For example now, I have to switch to another alternative acme client, since the old one doesn't seem to be in active development anymore...
However, certificates used to be expensive, even unaffordable for some projects.
Of course that was a bit trenchant, however you're complaining about a promise that was never made. No one ever promised that you could setup a old certbot instance and it was out of sight and mind for perpetuity. There are any number of issues that can occur, and honestly if one expected certbot to run without issue, having it automatically updating as well seems to be a base minimum.
Also worth noting that LE was early with ACMEv1, but a lot of alternatives started with ACMEv2. ACMEv2 became the common standard.