Spooky action at a distance, how an AWS outage ate our load balancer
blog.hostedgraphite.com
blog.hostedgraphite.com
Also kind of terrifying from a security perspective, since it's implied that Slack has been effectively delegated permission to modify AWS resources in their account.
I guess I've dealt enough with load balancers that as soon as I saw 'AWS network issues' and 'Many of our clients are hosted in AWS' I immediately jumped to the LB connection pool being full of mostly-dead connections, but this is an important lesson for anyone utilizing LBs. This is a very easy avenue to attack for bad actors, so setting strict timeouts and high connection limits (within reason vs. the capability of your load balancers) is important.
I know it’s easier said than done, but there should be active mitigation in place, rather than only monitoring.
We had a similar haproxy session saturation problem a couple of month ago. By now, our alerting would pick this up within a minute and trigger alerts. Those include a runbook to resolve it. Our standard resolution would fail in a case like this, but I'm pretty sure we'd solve it in a second iteration.