How Algolia Reduces Latency
stackshare.io
stackshare.io
One detail I particularly liked (buried deep in the article) was this one: "Once a machine is detected to be down, we push a DNS change to take it out of the cluster. The upper bound of propagation for that change is 2 minutes (DNS TTL). During this time, API clients implement their internal retry strategy to connect to healthy machines in the cluster, so there is no customer impact."
Offering client API libraries that have a retry strategy baked-in and relying on that for part of your high availability strategy is very neat.
One day, during a problem, every single query will take over a second, and this will be an exciting Slack channel to be in.
Really love the docsearch project btw, great UX.
EDIT: Just found this blog post right after posting this comment https://stories.algolia.com/algolia-s-fury-road-to-a-worldwi...
Implanted search for a client using Algolia last month and was completely blown away. The speed queries returned at were amazing. (Although it wasn't my idea to use Algolia I'll definitely be looking for opportunities to use it again).
We cannot use any local network IP or load-balancer as we distribute a cluster on several providers with different autonomous systems. This is how we are able to offer SLA of up to 99.999% with a big refund strategy: https://blog.algolia.com/for-slas-theres-no-such-thing-as-10...
Why not CloudFront? Moving binaries through CF CDN might not do for you. It's also seen that they don't really like you moving large data through their service [0].
That's a big difference.
Is this to ensure data integrity, or some other purpose?