https://news.ycombinator.com/item?id=6589508
I remember the week after this. Everyone I knew who worked at a fund was going over their code and also updating their Compliance documents covering testing and deployment of automated code.
As a side note one of hte biggest ways funds tend to get in trouble from their regulators is to not follow the steps outlined in their compliance manual. Its been my experience that regulators care more that you follow the steps in your manual than those steps necessary being the best way to do something.
I came away from this thinking the worst part of this was that their system did send them errors, its just that when you deal with billions of events emailing errors just tend to get ignored as at that scale you generate so many false positives with logging.
I still don't know the best way to monitor and alert users for large distributed systems.
The other take away was that this wasn't just a software issue but a deployment issue as well. It wasn't just one root cause but a number of issues that built up to cause the issue.
1) New exchange feature going live so this is the first day you are actually running live with this feature
2) old code left in the system long after it was done being used
3) re-purposed command flag that used to call the old code, but now is used in the new code
4) only a partial deployment leaving both old and new code working together.
5) inability to quickly diagnose where the problem was
6) you are also managing client orders and have the equivalent of an SLA with them so you don't want to go nuclear and shut down everything