Self-healing system: increase cluster size and replace servers randomly. It works because it was a problem of threads occasionally entering an infinite loop but not corrupting data. And the whole system can tolerate these kind of whole server crashes. IMHO an unusual combination of preconditions.
It's not explained why they couldn't write a monitor script instead to find servers having the issue and only killing those.