* Fault tolerance: near-infinite MTBF
* High availability: near-zero MTTR
With HA there is a blip. It might not be visible to an application because of retries, but it is visible outside of the HA system/component to some degree.
Special bonus thought: as a guide to designing or implementing an HA/FT system, I always found it helpful to think in terms of what happens to system reliability as size increases. In a traditional system, system reliability goes down because of dependencies between nodes/components. In some systems this degradation is even worse than you'd think because it's tied to the number of connections - O(n^2) rather than O(n). In an HA system, system reliability should go up because of nodes being able to cover for each other.
The key question was always: if X fails, what other part of the system can make up for (not just survive) it? If it's a whole node, what other node(s) can take its workload? If it's a disk, where is another copy of the data? If it's a network, how else can nodes communicate or at least synchronize? That last was interesting but because it led to things like serial lines or pinging through shared disks as a last-resort way to convey cluster state. Fun times.