When you're dealing with millions of users and response times go up above timeouts, there isn't much difference between "at capacity" and "down" if most users can't reliably use the service.
exactly, I think it will be an interesting story in retrospective, what happened here