... most computers aren't particularly reliable.
I'm curious - are you not monitoring or at least keeping a vague eye on ECC correction events with your existing hardware? If so, are you just not seeing any?
I've never really operated at any sort of "large" scale - a handful of racks at most - but I've always found correction events to be about as routine as IO errors, and certainly way more common than outright disk failures.