Unless I'm misunderstanding something the system did not perform as documented. It should have scaled, it didn't.
When a critical piece of infrastructure fails under massive load I'm not sure it it'll help much when you politely tell your engineers they fucked up for not anticipating it.
You learn lessons. Both Slack and AWS seem to have learnt lessons here.