Ideally, you periodically test your ability to failover. But if it doesn't work, well, there's a chance that you just caused a user-facing outage with your test.
I've been involved with a number of failover systems where even when it worked there was the possibility that you might hit a _known_ condition that causes the fail-over to fail.
Pretty scary knowing the product your working on has a couple critical holes in the fail-over that management papered over, which while rare could happen. A lot of these solutions are the equivalent of pull the power on one machine move the disk to the other and power it on. The assumption being that the storage mirror/replication/etc being used to maintain transnational consistency for the "move the disk" part is actually going to be consistent when that happens.
What happens if the failover fails as well?
There is a large disconnect with large system integrators there and the actual developers / architects; usually 2 or 3 levels of subcontracting. This means there are 2 to 3 levels of intermediaries taking a margin too, and as such, the actual developer doesn't have the fiscal wherewithal to push back on requirements or deadlines too strongly. (How do you save money on a 300k yen salary in Tokyo?)
If I was to guess, the requirement was there, but by the time someone technical got into the real nitty gritty, they discovered that the timeline was too tight to effectively do the testing. Instead of pushing back, they just rushed it through with a lot of overtime work.
"When the error happened, the system should have carried out what’s called a failover -- an automatic switching to the No. 2 device. But for reasons the exchange’s executives couldn’t explain, that process also failed. That had a knock-on effect on servers called information distribution gateways that are meant to send market information to traders."
I've seen backups and failovers not work so many times that it's an amazing surprise when one actually works--usually after laborious manual off-script intervention and invention. I was being somewhat kind to failovers, I have seen smallish backups work on several occasions. An untested disaster recovery is the worst.
Edit: think of an unexercised procedure as not "It works for me" but rather "it worked for me once."
On a failure, it might work. But there might not be a failure. Same difference.
But anyway, it also shows that the failover was not properly tested.
It seems they had a press conference with more details in Japanese.