So many human & tech error factors lead to this occurring and they're all the same old things. Staffing changes, spam filters, ignored warnings, skipped emails...
So many human & tech error factors lead to this occurring and they're all the same old things. Staffing changes, spam filters, ignored warnings, skipped emails...
Thanks for the reminder.
There's a remarkable number of ways for this simple thing to go wrong. To prevent a future repeat, we got rid of our calendar reminders (which we started ignoring once we both thought the change had been made) and wrote a script that emailed us based on the time to expiration of the live cert. This is a much better method.
Of course, give us enough years and I'm sure we'll manage to find a way to get this new setup wrong.
Add a test that just unconditionally fails on a certain date, like a week before your cert expires. Don't let any code review sign off on a merge of a fix the test until the new cert is in prod. Don't let any code promote between environments while tests are broken.
The problem with emails and warnings is they're all ignorable and therefore completely unsuitable for managing something as critical as a cert.
The secret is to create a straight-up error that absolutely interrupts every engineer in the organization's day until the cert is renewed. A development-halting error a week before cert expiration is a hell of a lot better than a business-halting error when it expires in prod.
Or just write a test that checks the date on the cert and conditionally fails.
If it's within 4 weeks, send email to x,y,z.
If it's within 2 weeks, send it to VP of x, y, z.
combination of failing test and email should lessen chance of it going unnoticed(email could possibly fail for whatever reason).