Some background ...
We have a fairly robust drive-proofing procedure prior to deployment - I think we still use 'badblocks' and really beat them up for 5-7 days. It's not enough to merely not fail - we have to see perfect returns on the SMART diagnostics and zero complaints from FreeBSD.
So we're starting with a known-good population of drives.
Then we monitor them very closely - again with both SMART and FreeBSD/ZFS system data and we are fairly aggressive about failing them out early if we see them start to misbehave.
So we don't see things like straight up drive failures - the bad ones have already been rejected and drives that misbehave are culled out.
The last time I got pulled in to make a judgement call (and got a little nervous) was (IIRC) we had a 15 drive vdev that had a drive failure and a candidate for early removal that had started to misbehave and the choice was made to yank them both because why do two big long resilvers ... and then during the subsequent resilver a third drive, while not failing, started to spit out errors. So the resilver completed and all was well but we had a potential third failure that could have died in the resilver and then we would be running with no protection.
But that's why I love raidz3 - even if that had happened, we would have needed to lose yet another drive to lose the vdev (and the entire pool).
I think the actionable recommendations here are to burn in your drives when you get them because we do, indeed, find rejects in most batches. Also, pick an error threshold that is low and be disciplined about sticking to it - don't let drives spin out SMART errors sporadically for months ...
EDIT: Here is another thought and this goes back to pre-ZFS days and old-style RAID, etc. In normal operations we're aggressively failing out drives that misbehave BUT if you're in a marginal situation you need to flip that logic around - especially if the pool is in-use during resilver/rebuild.
If you've lost 2 drives in a raidz3 (or, say, 1 drive in a RAID6) that remaining protection drive even if it is failing is like gold. If you fail it out, it's gone forever and has no relevance to the pool/array - but if you keep it, even as it's dying, you can either limp along OR you can even offline the pool/array and send that last drive to recovery or clone it or whatever ... the point is, when an array goes sideways every single bit of parity, no matter how poorly behaved, should be treated like gold.