NYSE Tuesday opening mayhem traced to a staffer who left a backup system running
bloomberg.com
bloomberg.com
In a week or so there will be a comprehensive internal post mortem, and every engineer in the company will read it because that's why they work there. "The staffer" will not be named, nor will they be fired. The process will be changed. The systems will be changed. You probably haven't heard of Pillar, but the NYSE in your head was replaced by some pretty amazing, distributed, low latency systems. The culture is to over-engineer, over-provision, plan for black swans. And test. That it works. Test that it scales. Test that backups work. Test, test, test. firmitatis, utilitatis, venustatis. This failure was due to daily testing.
Sometimes things still fail. That's true anywhere. In most places your failures don't make the papers, and accidents are swept under the rug. That doesn't happen at NYSE for obvious reasons. They're not building large language models (that I know of), or self driving cars (pretty sure on this one), but they're a modern, cutting edge, "soft" real-time engineering shop. If you haven't looked already, you might find something interesting there: https://www.ice.com/careers
They're a big company, some groups are better than others, some customers get more attention than others.
But they are. The consequence of a one-day or one-hour shutdown on their system is exponentially worse than most any other. I would expect them to have more rigorous systems, including more rigorous attention to development. Comparing the NYSE to any other business is like calling Fort Knox just like any other bank vault.
Interesting point that teeters on false equivalence. I think AWS or Azure might make for a better analogy. Your point identifies the inherent risk of actually operating a platform business. A bank vault is (mostly) synonymous with Cloud, in this context. If a vault is robbed or a cloud goes offline, losses extend beyond the business which inherently compounds the severity of downtime.
Linear loss vs. parabolic loss.
If a stock exchange executes trades at incorrect prices, even for a short amount of time, all of a sudden you're in a kind of non-linear sigmoid regime, where investor confidence can suddenly tip into panic selling and recessions can be triggered. Thankfully, that didn't happen here, but it could have. If you're going to give a company that power, you should better hope that they're held to higher standards than most dysfunctional tech organizations!
This is false equivalence and slippery slope.
There's far more critical snowflakes out there... FAA Airspace management, a medical radiation device, avionics in an aircraft, and facebook.
Them being a securities exchange does not somehow provide immunity from developing rigorous systems which have oversights, or make bureaucracy magically go away.
Likewise, the impact of an outage being more extreme does not mean the people there are infallible. Things slip through. Especially random customer requests being bounced around from team to team, the thing in question.
Main difference being that most bank vaults aren't actually empty. ;)
But as the company grew, the postmortems became more about blame since now you're not blaming an engineer, but an entire team so singling them out isn't personal. The postmortems were no longer a single engineer describing what happened in his code, but were team leads talking on behalf of teams. They were all about shifting blame from your own team and talking about why a service from another team led to the problem, even if your team could have (and should have) been able to work around it without melting down.
I'm no longer at the company, but Postmortems are much more useful when they really are no-blame because you can get to the real root of the problem, but I don't know if that's possible in a large company.
A few years back, a task to modify an index was given to a scrum team. The lead was away and the senior people could not be bothered. The junior developer stack overflowed an answer, asked for review, tested the script and let it rip. She missed that the change deleted everything if you noticed. Every environment, every data center wiped out. 10B records in each prod instance. Lessons were learned and processes fixed. She was not fired, but rather became one of the people safeguarding the keys to our prod kingdom as we fixed out broken process. I stole her away as my first report when I switched groups.
illinois.edu has one of the top rated CS programs, but once you graduate there are not a lot of options but to move to one of the coasts, or move back/to Chicago and try your luck there. Second City has a good deal of fintech.
If I can take that same skill set and apply it to something with a much better culture surrounding it that affects people in a positive way, then I would definitely choose that over finance any day of the week and twice on Sunday.
At the end of the day, if the NYSE did not exist, the world would continue to turn. It's just not that big of a deal to a heck of a lot of people.
This is startlingly ignorant of the complex machine that is the modern economic system. If something like the NYSE was to shut down today it would be pandemonium.
There is a difference between 'I don't understand how something works' and 'I don't understand how something works, so it is worthless'. The former is healthy and the first step to understanding, the latter is ignorant, and the first step to getting more ignorant.
The current state of business within the current iteration of how people interact with one another isn’t some necessity.
Yes, the world may fall apart for a relatively brief moment in the grand scheme of things — but then life will go on.
The first step to understanding this is to drop the superiority complex.
Very little is actually needed to keep the world turnin.
The real answer here, in my opinion, is that yes there would be pandemonium, and then yes, the world would go on without it, but then something else just like it will pop up. And that is because a liquid market for financial assets (whether that is securities, options on securities, futures, etc) will always be a massive benefit to the ability of businesses to conduct business, and the ability of individuals to preserve and increase wealth.
The movie portrayals don’t match my experiences at all and I saw a lot more bad behavior in the tech companies I worked for.
Heck I saw more people working for the intellectual challenge of it in trading than I did in SV style tech firms where money drove nearly every decision.
It’s really hard for me to buy that SV style tech companies are a better place to work when for the last 2 decades the business models that have been front and center are panopticon style tracking to sell ads and legal arbitrage.
It's not an either or, I can hate both ;-) I'm a big boy and get to make up my own mind on the matter.
Maybe they aren't as smart as they think they are? Or they find that there are interesting problems to solve in fintech? Problems they can tackle and see resolved in a realistic time frame vs 'tangible' (?) self-driving cars or chat bots.
I know AI encompasses a far larger range of things but right now, what problems is it solving? Artists, writers, and others can do that work. What do self-driving cars resolve beyond continuing the dominance of car culture in a world that could have better public transit and safer infrastructure?
There are problems in Fintech that are absolutely worth solving for altruistic reasons. One that I think is very important and might even need to incorporate AI is this :
Larger financial institutions have access WAY more and WAY higher quality data surrounding stocks and options. For example, publicly available SEC filings contain extremely useful information about companies. Professional traders have access to services which provide this data accurately in programatic form (like an API). Us normal people have only the SEC filings themselves, which are enormous documents. It would be impossible to read them fast enough to ever catch up on all of them in the last say year. There are free APIs, but they are absolute dogshit and provide incomplete and inaccurate information.
If someone could democratize this and provide this info for free or cheap to the public, it would be an enormous benefit to the general public.
Here's a quote from their Pillar product page:
> Up to a 95% Reduction in Latency: The roundtrip latency on NYSE Pillar order entry sessions via Pillar matching engines has been reduced from ~592μs to ~32μs for FIX and from ~96μs to ~26μs for Binary, getting client orders into the market much faster. With a 92% improvement in the 99th percentile latency results, clients can also have more confidence in improved performance consistency regardless of market conditions.
Reading stuff like this makes my current work feel stupid by comparison.
It makes our economic system seem stupid. Jesus, we're not calculating astrophysics or quantum mechanics. A made-up system should not require or depend upon this kind of speed or precision. Maybe we should chill.
Reminds me of those pro StarCraft players who keep unnecessarily clicking the mouse to keep their APM (actions per minute) stat high.
This is just another step in the endless journey of widening the gap between your average Joe and someone with access to high level financial services.
Systems with such significant potential impact, and in industries where lack of financial investment in their continuity is a deliberate choice have very little excuse to be passing the buck to grunts for basic process flaws that can be triggered by individual error.
The list of managers stating that "they were taking responsibility" and then immediately stepping down was always fairly short.
Wait, what? You think a CEO should step down because their management over-hired a relatively small proportion of employees and had to do some layoffs?
The people laid off and the people not needed were a different set of people, at the time of the layoff.
"Kill one man, and you are a murderer. Kill millions of men, and you are a conqueror"
If you make some idiotic financial decision near the bottom of the management tree, such as... over hiring, you'll likely lose your job or get demoted.
Do it as a CEO, and get a huge bonus.
Some things are cyclical and you need more people for some amount of time, and then you find you need less. It’s not always predictable/seasonal like farming or holiday rush.
Is it wrong for a company to respond to market effects? That there was a layoff isn’t necessarily a sign a company did anything wrong… I think how they actually do the layoff certainly can be done well or poorly.
I’ve forgotten which FAANG it is. But one of them still has more employees than last year even after layoffs. It’s offensive.
It was likely the right call to hire then, just like it might be the right call to reduce headcount now.
Maybe they should step down any time they fail to accurately predict the future?
Frankly - People seem to be forgetting that until 2013, MS was still doing stack ranking and routinely letting go of the bottom 10% of their workforce (and they were hardly the only ones doing it...)
I don't see it as unusual AT ALL that these companies are doing a wave of cuts to headcounts after the large hiring sprees during covid. Especially as interest rates rise, so they're looking to lower debt burdens in the short term and pay off loans made at low interest rates instead of rolling into a higher interest loan in the new environment.
If anything... I'd expect the exact opposite - a CEO that fails to address cost centers as debt becomes more expensive is a liability, and someone the board might be looking to replace (ask to step down).
---
Does that mean I'm not sympathetic to those who've lost jobs? Of course not.
But tech had to rev the engine pretty hard to handle the extra load during covid when everyone was indoors and doing things online, and now that demand has dropped. So they're letting off the gas pedal.
If folks don't like it - blame the game. Work to unionize. Work to incentivize co-ops and shared ownership. Work to increase taxation on these companies and their highest earners (which... if you're in the tech industry almost certainly includes YOU). Don't go work for giant tech conglomerates and then act surprised when they act like giant tech conglomerates...
Of course, once you take advantage of such a representation at scale by deploying tremendously more complex infrastructures, you then have to deal with the dependency network meta challenge lest you inadvertently fall into dependency hell. While towards there lies NP-hard problems, they're still computable to a reasonable degree and I dare say a more robust situation than doing it all by hand like we do today.
The real challenge is the vast majority of devops staff today would really dislike reasoning about such a representation when it blows up in their faces, and I can't blame them for that kind of reaction.
What if multiple sensors fail or it's an ambiguous situation like say you are deciding whether or not to fail over a power circuit and it's a brownout but not a complete power failure? What if there is a systemic problem and it's likely the backup power source is going to brown out too? At some point you need highly skilled individuals, like say trained airline pilots flying a plane who have the authority to override systems immediately without having to jump through hoops.
This is especially true for mission critical systems. Many of the mission critical systems we rely on are NOT built on the cloud, i.e. other people's computers because you want to be really careful about what hardware you are using, precisely how your data center is setup and want to make sure things like a noisy neighbor do not impact you.
Like it or not, these highly trained individuals are going to make mistakes every now and then. A failure like this once every decade or so really isn't so bad. The individual who made this error is likely not a "grunt". I suspect the individual in question will not necessarily suffer any major consequences as a result of this unless it wasn't a mistake but a flagrant disregard for the rules like say bringing a bottle of water into a data center that then spilled or something.
Have you built a mission critical, distributed system that hasn't failed for 10 years? It's a lot harder than it looks. That's how often the NYSE has a problem like this, about once a decade. A lot of things that work in theory, don't work for the edge cases and things that lead to problems once a decade or so are extreme edge cases.
In the grand scheme of things a mucked up opening auction is a minor problem and anyone who did not take the precaution of sending a limit order and sent a market on open order despite it being standard practice to essentially always use limits and go hurt badly will be made whole.
It pretty much boils down to: it depends upon what the business wants to prioritize; operating margin or resiliency. There is an entire subfield investigating the statistical foundations of resiliency, and the general case of N-modular redundancy is in practice implemented as triple modular redundancy in most commercial systems that want to spend in this vector.
> Like it or not, these highly trained individuals are going to make mistakes every now and then.
Absolutely, and here is where the organization's no-blame learning culture swings into action for the well-led teams.
> It's a lot harder than it looks.
We all know this, and we can all help each other get better to deliver ever increasing value to our customers by sharing what works for the context we deployed within!
In real-life outside of a journal article, it's a lot harder than just deciding whether you want to prioritize operating margin or resiliency at 5000 feet.
In real life when these sorts of edge cases happen, you have to understand in minutes or sometimes seconds the tradeoffs in terms of costs to your own company and your customers of one of n specific possible failure modes and risk-manage so you minimize the probability of the catastrophic outcomes. This sometimes may involve increasing the probability of low cost bad outcomes. You can't reason about this stuff before hand. If you could, you would have designed your system to not fail in that manner.
Exactly.
I wonder if any of the people claiming "it's management's process fault!" would be the first to complain about their workplace where they have no autonomy.
There are many kinds of "outcomes". A simple backup would make outages far more rare.
It didn’t.
> [NYSE execs] plan to examine the platform’s procedures and management, potentially reworking rules to be more flexible and provide further protections.
Sounds like they know it's a management issue. The headline probably focuses on the staffer leaving the backup system running simply because it's a better headline.
I will not be surprised if nothing gets fixed with the issue at NYSE.
You're reminding me of the difference between engineers and non-technical managers; to many of the latter something's only a problem if/when the customer or senior mgmt are on the phone complaining about it. Until then it's all naysaying engineers being too pessimistic about process and risk.
At some point you need to strike a balance between freedom/flexibility and stupid proofing.
HN goes real hard on the "people are idiots and we should design things that no matter what buttons get mashed it all works out fine" side of things but in the financial world the balance is struck a little further on the "train our employees to not be idiots" side of things.
Furthermore, it's usually better optics to blame things on people because people can easily and cheaply alter their behavior cheaply (per incremental change). If you blame the outage on systems it raises questions of when it will be fixed and how much $$.
As an aside, it was almost certainly not individual error. At places like NYSE you pretty much always have 2-3 people who should be in a position to catch a mistake like this.
That's exactly the point that is being made here. Either the message being put out by the NYSE claiming this was an error by one individual is true -- in which case, NYSE leadership is to blame for setting up a process that allows catastrophic consequences for a single individual's error, OR the message being put out by the NYSE is a fabrication designed to redirect blame at some scapegoat, in which case NYSE leadership is to blame for putting out a false or misleading statement.
[Edit: It seems I misunderstood -- attributing this to an individual was done by reporters and rumors, not by a formal statement from NYSE.]
It was simply the explanation of what happened. I didn't get any hint that the said "staffer" will be fired or otherwise punished.
Is there a problem with the system that did not have enough safeguards to let this happen. For sure, but then no system is perfect. This glitch does not happen every day. From memory, I remember a NASDAQ glitch at Facebook's IPO. Let's say there are 2 or 3 glitches like that for major exchanges in one decade. How can you design a system that prevents bugs that show up once a decade?
In fact, blaming anyone is unhelpful unless baltent misconduct is the problem, and I don't think it is the case here. As always, shared responsibilities. I just wished a different wording, something like "NYSE Tuesday opening mayhem traced to a backup system not properly shut down". Leave the "staffer" part to the technical report. It is useful information for investing the problem and fixing what needs to be fixed, but it is inconsiderate for a press release.
If you're going to move the blame up the food chain, might as well blame the shareholders for giving the company money and choosing to keep the upper management in place.
At the end of the day, the engineers are responsible for the engineering. Managers are responsible for managing. Shifting all responsibility for execution issues on to management can give warm fuzzies, but in reality managers aren’t all powerful in shaping execution by engineers.
Companies that put all blame on managers when things fail are inevitably encumbered with excessive micromanagement, as the managers are effectively saddled with responsibility for execution as well.
The article was purely anonymous. I don’t think it’s fair to assume they’re jumping to blame or fire individual engineers.
Some updates are highly regimented, but a couple of the more operational teams have discretion to deploy things outside of that process, and most teams can flip feature toggles whenever they want.
Point is that sometimes people will comment, or even veto changes. We have a major customer visiting today, or the sales team is at a conference. Don’t touch anything or you might break something.
Stuff like this will happen more and more. We treat software driven systems rather recklessly.
- William E. Vaughan https://quoteinvestigator.com/2010/12/07/foul-computer/
It's clearly not optimized for efficient use of personnel, but the personnel complement will have been designed to provide sufficient people at all times and the cost of getting it wrong can be very large indeed.
Though in that state it can't do much more than fly: combat capabilities are strongly diminished, maintenance doesn't happen, post-combat repairs are out of the question, science missions would be much harder. On occasion the Enterprise has transported 150 passengers, so I imagine there's a lot of kitchen staff, security, etc. You only need 5 people to fly the ship, maybe 40 to fly sustainably with maintenance, but to actually accomplish their reglar mission you need the other 300 people.
On the other hand, the Enterprise won't even warn anyone when command staff are injured, cloned, mind-controlled or vanish from the ship altogether unless a human asks the computer where a specific person is first.
Of course there are Doylist reasons for all of this but I do like the premise of a general fear of AI and possible weird space BS being a factor.
The canon explanation is that automation-in-charge was experimented with and went really badly, though periodically they try something approaching it again.
https://memory-alpha.fandom.com/wiki/The_Ultimate_Computer_(...
(AI, human genetic engineering, and a number of other areas of technology are affected by variants of this issue in the Trek canon.)
I disagree. When the sub sinks, all these people die. That's far more inefficient use of personnel.
It feelf like you're putting more emphasis on the material cost ("the cost ... can be very large indeed") than on what actually matters.
In wartime you would care about efficiency, build the largest number of subs staffed with the minimum crew. And since training crew quickly becomes the bottleneck you would probably go for the highest degree of automation that doesn't impact production times too much. In peacetime, efficiency isn't as important. What is important is the bad PR of losing one of your submarines in a training excercise or on patrol, so crew safety becomes a much bigger concern.
Losing those sailors in war would have been a noble sacrifice for the cause, losing them to the exact same accident in peacetime is a national tragedy.
Losing one during wartime can cost you the war.
An inefficient one that probably won’t sink due to combat damage can stay in the fight long enough to matter.
As a former naval officer I can only say that technically losing any vessel could be the one that loses a war, just as any soldier lost could be the straw that breaks the camel's back. But if your navy is so rickety that the loss of a single vessel is enough to lose then the main deficiency was in planning rather than any specific warship loss. Losing one in peacetime should never happen but is not unheard of even in modern times. See eg the Kursk or the Fitzgerald.
It could also be argued that the Romans capturing a single Carthaginian warship turned the tide for their entire empire.
It was “too important to automate” so a trading assistant keyed it in every morning. One morning he typed the wrong number and the mistake was in the billions digit.
At 2:45 “the cage” called the repo desk and said “You know you guys are still short a billion, right?”
There was then a flurry of activity as traders got on the phone to try to borrow a billion dollars in in fifteen minutes, while also trying to not let on we were kind of over a barrel. The head of fixed income prepared his explanation to the Fed about why we needed to borrow a few hundred million overnight.
The number got automated in our next release, and the open procedure was changed to the trading assistant verifying the number against the “cage” report.
Process changes that people have to remember or more systems to prevent the issue. So I don't get your statement related to this article
Shutdown and wake-up time in bios of server and switch ;)
We will simply create a second system (B) to monitor the first system (A). Now we have two systems to maintain. System B will not be capable of steering A by itself. So we still need to know how to diagnose and repair A, and we also need to know about B too. Maybe system B can talk to a Prometheus/Grafana stack (if it's up). And that can put alerts into Slack (which we ignore because there's always alerts in Slack). And after standup we can take turns looking at graphs with consternation.
> Stuff like this will happen more and more. We treat software driven systems rather recklessly.
That sentence is where I go when I hear the word 'automation'.
More people -> more entropy -> much more things to go wrong
I did last year, and my company is in the process of de-automating certain processes that can endanger that company if they go wrong.
There are many things in tech that are too important to automate.
I'd even posit that the more experience you have in tech, the more you've seen how things go wrong, and the more you realize that automation is a tool for humans to use, not a replacement for humans doing a task.
Much much much more due to human error......but hey maybe you are the worst programmer ever..but even then i would say your programs are more reliable then a human.
You'd be surprised how many manual processes there are in places like this. It's a combination of legacy systems / processes, and a general paranoia around automation going wrong. I wouldn't be surprised if they always have someone there to shepherd the system along.
We had hundreds of jobs and upgrades happening over each weekend. It definetly needed an eye casting over it regardless of the automation.
And the AU orders going through is a good sign, but it's far from guaranteeing a free monday, as Japan, Korea or Shanghai can fuck it up, each in their own little ways. Hong Kong is the best, low regulatory crap, invested regulator, high volume low latency traffic everyday (relative to the region), I cant recall a time it broke.
Once, someone fat fingered an excel import at close, and we lost our trading license for that entire country for 18 months. And we're not small. But the amount mismatched at settlement was super tiny. High attack surface, low holistic understanding (it works despite us, we honestly have no clue sometimes), heavy consequences on screwup.
Basically at the market open all the requests to buy and sell get matched at the same "open" auction price, then (a second later) the orders get sent to the order book, where the price can go up and down based on size. Because the system didn't think there was an opening, there wasn't the opening auction, the prices went straight to the book and there were large swings in price.
You'd have to have more than 1 person involved to forget that DR is still active when completing these failover exercises and tests off-hours.
Well, they are a tenent at the Cermak data center. It’s a truly massive building with huge amounts of connectivity and colo opportunities. Probably also the only 100+ year old data center building on the US register of historic places, lol (it’s a former catalog printing facility, built to hold insanely heavy printing presses on 8 or 9 really tall floors, so it has no problems with densely-packed server racks)
It seems to have multiple purposes - DR, customer software testing, etc.
Edit: I also note that this piece is lacking the traditional "The NYSE did not respond to a request for comment".
Anonymous means they aren't revealing the source, not that Bloomberg doesn't know who the sources is, or what they do.
There's usually a link to it in the middle or end of any story it publishes using an anonymous source.
The Times isn't Bloomberg, but it might give you some insight into how these things work.
Still this report (and the previous statement) does not give enough detail on why a backup system misoperation resulted this. Also, critical large systems like exchange rarely have a single point failure. Usually there will be a sequence of issues along the event chains leading to this. Thus one "failed to properly shutdown" caused all this is a bit incredible. We will need more explanation.
Personally, I always do limit orders, but I would consider market on open/close as reasonable options. But I don't think this is typical, a lot of orders are market orders against whatever limit order is at the top of the book. Normally, that's ok, but it gets weird when things get weird, as seen here.
From what I'm reading, this NYSE error seems a bit more complex where the presence of a backup system confused the current market state to skip the open auction.
If your system can be hosed by a single person the system is at fault. Start with the scapegoat's manager.
Which is also why exchanges are very reluctant to mass cancel trades. The knock on effect goes beyond just market data feeds
Smells like horse shit. Most of what comes out of the profession of "journalism" does too, lately; but this smells strongly.
I know the world is held together with duct tape, but it's embarrassing when you see the tape fall off.
The NYSE DR guide [1] says that if DR is active, production is not. It's not a distant reach to consider that some of these clients have a deadman switch doing a healthcheck poll on DR and switching to it when it see's that it is "up". If they've built their systems in such a way that when it detects the DR site active it uses that, then it makes sense that having both "online" would cause some havoc. I'm sure the complexity of the entire exchange is fairly significant, and having "two" copies of it running in parallel with both able to accept and execute trades would be a scenario that can cause some unintended consequences. Fundamentally, an exchange is "atomic" and transactional and cannot be meaningfully distributed to two sites that are that far away. The replication in place is likely master/slave with a switch to make the slave primary. Anyone who has toyed with master-master replication on less complicated databases knows the issues that can come up with split writes. Imagine that at the scale of a system as large as the NYSE.
[1] https://www.nyse.com/publicdocs/support/DisasterRecoveryFAQs...
https://www.amazon.com/Phoenix-Project-DevOps-Helping-Busine...
tl;dr: The brave knight implemented devops and everyone lived happily ever after!
Some people need a story.
I didn't even know about this process. I don't know much about trading, but it surprises me that there is a separate process for setting prices at the start of trading, and that if it's missed, chaotic prices result.
Is this related to how stock markets aren't really ever open 24 hours? Do they need that reset to function in stable way?
Part of it is legacy from when trading was done by actual humans being at the exchange physically to trade during those times and part of it (I would guess is still the case) is to allow plenty of non-trading hours for back-office jobs and settlement.
So yes there's a special start of day process that runs at 9:30 that runs through all the orders on the books at that time and determines a price at which some optimal set of those orders can trade, trades them at that price, and also posts that price as the Open price for the day.
The process is different during continuous trading since orders are one by one matched against the order book.
Source: ran one of the world's largest equity platforms for 5 years.
The price is usually calculated algorithmically by the DMM firm and sent to the person at NYSE to approve. Pretty arcane. Also somewhat shady, as the DMM firm can be and is part of the auction themselves. DMM firms can analyze the order book to see what the imbalance is in the overlapping region, and place an order of their own to correct the imbalance and then set the opening price. I can see how one can profit from this in certain situations
Now, as a result, there needs to be a way to set the opening price and closing price, like a bootstrap process. A smaller version of this process actually happens every time a stock gets halted and resumed.
An exchange has an order book - orders of things people want to buy and sell at different prices. During normal operation the buy and sell orders don't overlap in the order book - if two people want to buy and sell at the same overlapping price, they just get matched by the exchange at that moment. Unmatched orders stay in the order book data structure until a matching order comes along. The "price" you see in charts is just the midpoint between the highest buy and lowest sell price in the order book.
Now, if the order book is empty, what the heck is the price? That's what the opening auction needs to solve. The way it works is that people can start placing orders ahead of the opening bell, but they won't get matched until the open. So before the open, the order book is getting filled with orders, but crucially the _orders will overlap_. This "crossed" order book is a no no during normal trading, but ok before the opening auction. When the auction comes, a price is picked which maximizes the amount of orders filled (it's more nuanced than that, but bear with me). Imagine you pick a price in the overlapping region of the order book - every buy order that has a higher price than that will match with every sell orders that has a price lower than that. They will get matched and executed at the opening price, and BAM, you have an uncrossed order book, full of orders.
If the auction doesn't happen, and you just open the stock, then all hell breaks loose. Many things can go wrong here. Firms connected to the exchange may have code that assumes a book is not crossed (or at least not as crossed as it would be during an auction) causing wild behavior. The exchange itself could start matching orders haphazardly in the overlapping region, causing those "price swings" that the article talked about.
Can't imagine the panic that day haha.
Does that track with your understanding?
A buy market order would try to match with the "best price" which in a deeply crossed book would mean matching with a really low priced sell order. Exchanges match orders in price-time priority. Similar is true for a market sell order - would match at an extreme high price.
Besides the midpoint of the order book, another metric for a "current price of the stock" people use, is the "last trade price". In the situation above you would get "swings" in the price because market orders would be trading very high and very low if they alternate between buying and selling. The data structure on the exchange itself isn't "swinging", it's just the overlapping region being slowly eroded by market orders. The "last trade price" metric looks really insane in this situation.
> Now, as a result, there needs to be a way to set the opening price and closing price, like a bootstrap process. A smaller version of this process actually happens every time a stock gets halted and resumed.
So this suggests that if you did have a hypothetical exchange that ran 24/7... and something unusual happened to make trading halt completely (which always is going to happen occasionally, whether 9/11 level or more frequently)... you would still need to have that "bootstrap" process in place to re-start trading.
But if you normally ran 24/7, you'd have a process that you maybe had never used, or hadn't used in years!
This maybe provides another justification that isn't just historical for having exchanges shut down every day. So you are at least testing the bootstrap process daily, you don't have a bootstrap process you're going to need in an emergency (the worst time to have further problems) that has actually just been sitting around unused for years!
(Reminding me of making sure you test your backup and continuity processes regularly, right? And the irony here is that it's the backup/continuity processes which are alleged to have caused the issue here! but still, you need the backup/continuity processes...)
> Question: Can I connect to both the production and the DR site at the same time?
Answer: No, only one site is available at a time. When the primary site is up, the DR site is down; and when the DR site is activated, the primary site is down.
I think they need to update these docs to say /should/ be down
Like the S3 being blown away with a simple change in the early days, or GitHub running a test suite with production settings. It's like the FIRST thing I think about when starting a project.
That said I find the US market structure is unfair Charles Schwab does protest too much. Retail orders never seem to get near the central order book. there is no direct market access. brokers just sell your order to whomever MM pays them for the spread in return for a kickback. this should be a fantastic fair multiplayer game, but instead its pay to win mobile crap with vested interests milking their customers.
> Meanwhile, market professionals and day traders are rattled and waiting for the exchange to elaborate on what it publicly called a “manual error” involving its “disaster recovery configuration”.
Oh, I love it -- a disaster caused by "disaster recovery configuration" :-)
People install failover configurations to minimise time-to-repair or time-to-resume service (and some customers' contracts will demand this). This is at the expense of another layer of stuff to go wrong, and raising the possibility that it fails over when it shouldn't, causing brief but embarrassing outages.
It's possible in some such situations that, on the balance of probabilities, introducing mechanisms like this cause more disruption over time than they were intended to protect against, and that this is more widespread than often considered. Still, their operational cost must be borne in order to satisfy the clause in the customers' contracts.
The fact that these systems do not exist is an exchange problem, not a "staffer".
Retail traders realistically have only luck to rely on to beat hedge funds and banks. What they do is akin to gambling, which is on net quite negative for those who participate in it and heavily regulated. Retail traders don't serve any purpose in our society. They don't help with efficient allocation of capital and anyone who might be an actual savant in trading can join or start a firm rather than staying independent and unlicensed.
What about people who designed it this way?
In reality what probably happened is previous market day and post-trading data encountered some kind of error, which triggered a cascade of problems overnight that they were unable to properly rectify. This caused delays up until market open. They were unable to fully resolve the issue, and forced with either delaying opening the market (which is a HUGE no-no) or opening with wrong data as is, they chose wrong data.
All in all a lot of people didn't get much sleep Monday. More than likely they implemented some changes or updates over the weekend that were not properly done, or they encountered some errors, and didn't have adequate controls/time to roll-back Monday night. They made the right calls too late and there was a controls process up the chain that seriously fucked up. These are the kinds of problems that get the CEO woken up in the middle of the night.