Incident review: Intermittent downtime from repeated crashes
incident.io
incident.io
Sharing a technical write-up of our recent incident where our service repeatedly crashed due to a poison pill event in the async workers.
Interesting from a general incident coordination angle, but also in terms of mitigations some of which apply generally to web apps, such as splitting work by category/type for improved reliability.
Hope you enjoy, feedback always welcome!