Tokyo Stock Exchange Blackout: One Piece of Hardware Took Down a Market
bloomberg.com
bloomberg.com
The one that really sticks out to me as an engineer is the fact that the whole system in most cases seems to be tied together by a fragile arrangement of 100+ different vendors' systems & middleware that were each tailored to fit specific audit items that cropped up over the years.
Individually, all of these components have highly-available assurances up and down their contracts, but combine all these durable components together haphazardly and you get emergent properties that no one person or vendor can account for comprehensively.
When the article says a full reset entails killing the power and restarting, this is my actual experience. These complex leviathans have to be brought up in a special snowflake sequence or your infra state machine gets fucked up and you have to start all over. When dependency chains are 10+ systems long and part of a complex web of other dependency chains it starts to get hopeless pretty quickly.
Or worse, "compliance" line items, that some tool or some company identified in their cookie-cutter processes. As long as that line item goes away, noone really cares what the long term implications are.
Basically, I think you can remove your first sentence and leave the last to make a truer statement. A great system should have design considerations and tools to make it understood at various scopes by various people of various skills. I think people have falsely conflated event-based architecture with systems health due to having to rewrite a system almost from scratch, which results in better tooling since it's something you're thinking about actively.
+1 not sure how event driven arch would have mitigated here. A well designed (with redundancy where necessary, better tooling, monitoring, scalable, etc) would.
I've had to explain this to a dozen teams so far in my career and most of them go ahead with the design anyway, regretting it not even 6 months later.
This is a characteristic with all event-based systems, but persistence-enabled event systems (such as Kafka) make it even harder because now there are events already "in flight" that have to be taken into account. Event-based systems that do not have persistence (and thus are simply message queues used as a transport mechanism) have a strong guarantee that _no_ events will be in-flight on a cold-start, and thus you have an easier time figuring out the current overall state of the system in order to make such decisions.
The only other way around this is to make every possible consumer of the event-based system strongly idempotent, which (in most of the problem spaces I've worked in) is a pipe dream; a large portion certainly can be idempotent, but it's very hard to have a completely idempotent system. Keep in mind, anything with the element of time tends not to be idempotent, and since event systems inherently have the element of time made available to them (queueing), idempotency becomes even harder with event based systems.
A rule of thumb when I am designing systems is that a data point should only have one point of persistence ("persistence" here means having a lifetime that extends beyond the uptime of the system itself). Perhaps you have multiple databases, but those databases should not have redundant (overlapping) points of information. This is the same spirit of "source of truth", but that term tends to imply a single source of truth, which isn't inherently necessary (though in many cases, very much desirable).
Kafka, and message queues or caches like it (e.g. Redis with persistence turned on), breaks this guarantee - if the persistence isn't perfectly synchronized, then you have, essentially, two points of persistence for the same piece of information, which can (and does) cause synchronization issues, leading you into the famously treaterous territory of cache invalidation problems.
As with most technologies, you can reduce your usage of them to a point that they will work for you with reasonable guarantees - at which point, however, you're probably better off using a simpler technology altogether.
Then the events will wait in a queue and be consumed when sub side is ready.
You reboot the sub side. The event never gets processed because pub already sent it and recorded this fact.
Or, the sub side doesn't reboot, the pub side does. The pub side accepts an event for publishing, send it to the sub side, and promptly loses power.
The pub side reboots and either it resends the event and the sub side receives the event twice (because the pub side didn't record that it had already sent it before power was lost), or it doesn't resend the event and the sub side never receives it (because the power loss killed the network link while the packet was on its way out).
If you think you can make these and other corner cases go away with a simple bit of acknowledging here and there, good luck!
The corner cases can be solved, but it's not half as simple as "wait in a queue and be consumed when ready".
I'm literally saying it can't be fixed with a simple solution, in the parent comment and other comments, so I'm not sure where you get the idea that I'm saying it can.
Dealing with inconsistent states from failures in a distributed system is solvable but it's not simple unfortunately. It's not even simple to describe why.
Full 100 services should have end-to-end integration testing and any change made to that chain of tooling should have to run through a massive integration test. If anything fails, the change is no longer acceptable.
The answer is that tests are never perfect. If you want to create an integration environment that mimics prod, you have to fork an entire parallel universe into your integ environment to run the test. Anything else will diverge from the reality of the future.
Even if every vendor's service or hardware had integration tests, that doesn't mean that the integration tests covered every case. It doesn't mean there's not an emergent property of two systems behaving in a slightly unexpected way that turns into a catastrophic result.
It's not necessarily even possible to have two copies of some of the systems; who knows how expensive a given vendor's hardware box is.
It's definitely not possible to exactly mimic future traffic. Perhaps in the test environment it works, then the prod environment, requests are different so it fails.
Hardware errors happen, and integ testing those is difficult to say the least.
Given you are running a stock exchange, chances are you have enough money for at least one copy.
"Everybody has a testing environment. Some people are lucky enough enough to have a totally separate environment to run production in."
The good design approach is obviously superior in many ways. But the downside of it is that you have to trust the competency of a lot of different business units to maintain it. A cumbersome bureaucracy on the other hand ensures that incompetent/lazy/low-initiative internal actors can’t impact your compliance. If you fail, at least you fail in a compliant and auditor-approved way.
That said, a lot of the failures I’ve seen in organisation like this stem from silo’d expertise. People don’t know that much about the systems that are outside their remit, so they will make changes that impact connected systems in ways they failed to imagine. As an example I have seen 3 seperate banks have non-trivial service disruptions stem from the same independently made mistake. A person enabling debug logging on the SIP phones. The traffic DOSes their networks, and all of a sudden, core network infrastructure starts to die. Afterwards they send the right reports off to the right people, make the correct adjustments to the bureaucracy, and proceed with their compliance intact.
There is also a 3rd way to approach compliance, via negligence. But the more you are in the regulatory spotlight, the less of an option that is.
You know, just from a human perspective, talk about a bad day.
It's like seeing the cruise liner heading for the port too fast, knowing it is going to crash and cause immense damage, and realizing there is absolutely nothing that can now be done to prevent the damage.
Except, in this case, it's just way worse.
I was talking about monetary damages - excluding any kind of loss of life or personal injury.
I could have picked a better metaphor!
How much was lost by being unable to sell something today that you weren’t willing to sell yesterday?
So you have the opportunity costs for those people being paid to do nothing for a day and generating nothing in the form of revenue. Rightly or wrongly, a significant amount of profit generated in the stock market is from short term moves / games, not from buying and holding for long periods.
Second, and I don't know if this is the case, but what about companies that were issuing equities to raise cash on that day? They had geared up entire investment banking teams, with customers orders lined up, and had a whole game plan for executing the offerings.
That all went up in smoke and they will have to re-prepare everything for another day. That's true to IPOs, general follow-on offerings, to some extent ATM programs, etc.
Those are a couple ideas, maybe there are more?
Could be a lot. Markets are not disconnected. If only Japan existed, sure. Some orders would no longer exist, but nothing major, because noone else would be trading anyway.
But suppose this was during an economic downturn. You are trying to SELL SELL SELL because all countries worldwide are feeling some economic pressure and you can't because the exchange is down.
The next day, whatever assets you had are now worth a fraction of what they were one day before.
Oh, and there are some financial instruments that expire.
Some exchanges deliberately close in those circumstances. Or when big news is about to be released.
But as much as people were unable to sell, there are nearly as many situations where being unable to sell is a good thing.
I guess we’re eventually going to argue for 24/7 stock exchanges and for some reason very few are.
My own experience in Canada is that a lot of big Canadian stocks also trade in NY, and when one market is closed and the other isn’t because of different holidays, not a lot happens.
I guess an unexpected crash is different, but my guess is that everyone takes the day off, avoids releasing any big news out of respect for the situation and gets back to it tomorrow.
In fact the trend is in the opposite direction, towards shorter trading hours as well as more turnover in the opening and closing auctions. Shorter hours especially would be a huge win for work-life balance and diversity in the industry, and just as importly for reducing costs.
Having the markets open longer only spreads liquidity thinner across the day, and moreover having announcements made, corporate actions processed, etc etc is better done when the market is closed so everyone can be on the same page when the happen.
They just lost at least a whole day of revenue.
After compiling some number for their annual report[1], I would say the loss from TSE revenue alone is roughly 2 million USD. But considering the effect trickling down the revenue stream where other security partners whose revenue depends on earning transaction fees, I would say the real damage would be many folds of that 2 million loss from TSE.
[1]: https://www.jpx.co.jp/english/corporate/investor-relations/i...
However, you see, TSE YoY growth from 2017 to 2018 is roughly 4 million USD. Growth is hard as it is for TSE, leaving money on the table like this time definitely hurts.
The line of reasoning “well, no one died” is a classic example of humans’ unfortunate tendency to round small numbers to zero in utilitarian calculations, which leads to all sorts of bad decisions when you’re dealing with things that effect lots of people in small ways.
If the market closing is so bad for its participants, why aren’t they lobbying to keep it open past 4PM? Or on Saturdays?
Or at least adopt the retail model where they get some more retail traders by operating some weeknights/weekends.
Most equity exchanges are only open or liquid during local business hours. US markets are open longer, but generally illiquid and there are fewer protections outside of "market" hours.
I think for equities, nobody wants to be a market maker 24/7 without a sufficiently wide spread to protect them from news events such as the death of an executive. So even with 24/7 markets you'd probably see very poor conditions outside of core hours. With derivatives, there's more interest in never sleeping. The world is a very interconnected place and less hinges upon a single person's death.
That doesn't mean I think anything can't be measured or fit into some sort of framework, I just think there is some fundamental truth in the rounding to zero you refer to, that needs to be accounted for. It's too glib to dismiss it.
Why not?
[0] https://en.wikipedia.org/wiki/Japan_Airlines_Flight_2#The_%2...
This failure highlights how frequent, scheduled testing of your failover system is needed in order to be able to say, honestly, that you even have a failover system, and not just another box burning power and doing nothing. When you have a choice, it is often better to have both systems running all the time, at less than half capacity and sharing the load; or each doing all the work, and throwing away half.
If you choose the former, traffic can sometimes peak over half capacity, without loss. If the latter, you can check that they are producing the same answers, too.
The exchange suffers, paying salaries and rent without income, but they deserve it for providing crappy service.
While Japanese so-called journalists are completely blind at technology, even worse, they don't even have a basic literacy or listening skill, questioning things they already told, or fart out a question like "But computers won't break, isn't it?" while the cause is likely be soft memory error.
"When the error happened, the system should have carried out what’s called a failover -- an automatic switching to the No. 2 device. But for reasons the exchange’s executives couldn’t explain, that process also failed. That had a knock-on effect on servers called information distribution gateways that are meant to send market information to traders."
But anyway, it also shows that the failover was not properly tested.
It seems they had a press conference with more details in Japanese.
On a failure, it might work. But there might not be a failure. Same difference.
I've seen backups and failovers not work so many times that it's an amazing surprise when one actually works--usually after laborious manual off-script intervention and invention. I was being somewhat kind to failovers, I have seen smallish backups work on several occasions. An untested disaster recovery is the worst.
Edit: think of an unexercised procedure as not "It works for me" but rather "it worked for me once."
Ideally, you periodically test your ability to failover. But if it doesn't work, well, there's a chance that you just caused a user-facing outage with your test.
I've been involved with a number of failover systems where even when it worked there was the possibility that you might hit a _known_ condition that causes the fail-over to fail.
Pretty scary knowing the product your working on has a couple critical holes in the fail-over that management papered over, which while rare could happen. A lot of these solutions are the equivalent of pull the power on one machine move the disk to the other and power it on. The assumption being that the storage mirror/replication/etc being used to maintain transnational consistency for the "move the disk" part is actually going to be consistent when that happens.
What happens if the failover fails as well?
There is a large disconnect with large system integrators there and the actual developers / architects; usually 2 or 3 levels of subcontracting. This means there are 2 to 3 levels of intermediaries taking a margin too, and as such, the actual developer doesn't have the fiscal wherewithal to push back on requirements or deadlines too strongly. (How do you save money on a 300k yen salary in Tokyo?)
If I was to guess, the requirement was there, but by the time someone technical got into the real nitty gritty, they discovered that the timeline was too tight to effectively do the testing. Instead of pushing back, they just rushed it through with a lot of overtime work.