We've had one minor hiccup when LE changed the company that they used for front-end load balancing, but other than that it's set and forget. Generally we don't think about LE/ACME at all at $WORK. We still have some annually-renewed certs for 'appliances', including a wildcard, but practically we don't think about either 'system' more than the other.
If spending $x per year solves the problem for you, then go for it, but it does seem odd to me that you're having so many problems.
As for monitoring, a simple thing to do is to install some well-tested, pre-canned check scripts (e.g., from Nagios):
* https://packages.debian.org/search?keywords=monitoring-plugi...
and have them kick off via cron on a regular basis. A non-zero error code will cause an e-mail to be sent to you.
While we use Nagios (for now) to monitor from a central host, we also leverage them in things like keepalived for running HA clusters to trigger fail-overs. No sense re-inventing the wheel.