This is really interesting. Impressive to see a startup with so much bare metal hardware (800+ servers in 47+ data centers), due to their need to power autocomplete-style search and hence offer the lowest latency possible.
One detail I particularly liked (buried deep in the article) was this one: "Once a machine is detected to be down, we push a DNS change to take it out of the cluster. The upper bound of propagation for that change is 2 minutes (DNS TTL). During this time, API clients implement their internal retry strategy to connect to healthy machines in the cluster, so there is no customer impact."
Offering client API libraries that have a retry strategy baked-in and relying on that for part of your high availability strategy is very neat.