Thoughts on low latency trading if exchanges went full cloud
blog.abctaylor.com
blog.abctaylor.com
Knowing you can saturate your entire network with 10G traffic and every participant will get the same market data packets at the same time[0], and there will be zero queuing or bottlenecks is very hard to do otherwise. There is a pretty good podcast episode about it out of Jane Street[1].
I know AWS have 'multicast support' but last time I tested it, it was clearly just uni-cast traffic with a software switch doing fan-out/copying, I assume using the same tech as their transit gateway, I think it was called hyperplane or something.
[0]: for some definition of the same time, at least low enough that you can't measure it without equidistant optical splitters or White Rabbit synced devices.
[1]: https://signalsandthreads.com/multicast-and-the-markets/
Yup, this is a problem for us in GCP today even outside of trading. I don't know how Pub/Sub works for them.
We lost a lot by thinking of HTTP as the one true level of network abstraction.
A reliable multi-tenant multicast network also appears to be a rare beast. I’ve only heard of it in finance, and that’s only because it’s private and expensive and all the participants need to be generally nice to each other because it’s a repeated game and the operator can literally pull the plug if the rules are broken.
Also, yeah, you have to do some engineering around your multicast distribution to make a pub/sub system, but multicast pretty much solves the data rate scaling problem - you are now basically O(1) in the number of connected subscribers.
You had more control over your services talking to each other and control plane tech, thus could make some more guarantees than with inbound data. Don’t cross the streams.
My point was that other advantages/disadvantages are not being cared about, not that we should provide milisecond access to Average Joe.
If Average Joe wants returns comparable to these hedge funds, then they should stop trying to time to market and instead stick to diversified ETFs and stop worrying about millisecond differences in the stock market.
Believe it or not, if Average Joe does that they can actually beat most hedge funds over a long time horizon [1].
https://www.cnbc.com/2018/02/16/warren-buffett-won-2-point-2...
No, expertise is not the difference. If you're a private person with 100 years experience in trading, you still can't do HFT. You need to be an instutition, have lots of capital to invest in servers, software development maintainance etc. As a private person I think you don't even get access to the API.
"Of course whoever has more capital has advantages in the market"? Of course they do, but I don't think "of course they should".
For this discussion, that funds don't beat the market average over a long term is irrelevant. Why not say "who cares you get more latency than the other bank, if you want money just invest in S&P500 and long term you'll beat them". But you don't apply that to banks against banks, only to banks against Joe. Why?
Can do better at what? Can get their trade in the order book faster? Yes they can. But does that automatically mean they will make more money? No it does not.
>If you're a private person with 100 years experience in trading, you still can't do HFT.
Of course not, 100 years ago there were no computers. Having 100 years of experience in trading on the pit would not give you any expertise in software development.
Someone with 10 thousand years of experience plowing can't compete against someone with a tractor. That's kind of the point of the tractor...
I'm sorry that it disturbs you that the Average Joe sitting at home with his discount online brokerage account is unable to gain the same kind of benefits putting out individual orders here and there on speculative stocks that he likely knows nothing about, that hedge funds, institutions, and other highly specialized and skilled professionals are able to gain by doing this for a living.
The Average Joe does have access to highly diversified and low fee ETFs, and as I said the Average Joe can reap almost all of the rewards that the best hedge funds and banks do by sticking to those instead of trying to play the market.
Giving retail traders access to the "actual market" would most likely result in worse execution on average, according to some studies.
Hold on a second. Multicast is nifty, but it does not perform miracles. If you operate a 10G multicast network and actually saturate it, you will experience drops and buffering-induced delays. Perhaps you can play games with time-synchronous networking, but as far as I know the exchanges don’t do this, and it likely needs special hardware.
The point of 10G multicast is to use simple, standard (but complex to configure!) equipment to distribute much less than 10Gbps simultaneously.
If there are other data flows also going through the switch, that could obviously change things, and the sending computer could drop packets if there's jitter in how quickly the application produces them, but it seems impossible for the sending computer to burst packets into the switch any faster than it can handle because all the incoming packets are coming over the same 10G link.
Not an expert here, legitimately curious.
You can’t run a queue at 100% and have any expectations of latency. In fact the rule of thumb from queueing theory is 50% to avoid latency spikes.
Must be nice to open Wireshark and see _nothing_.
That being said, the ASIC can typically handle line rate on all the ports. You could have 10G input and fan it out to 10G output on all the ports with no drops, but if there is other cross port traffic, something could get dropped.
When 800G nic's come out next year (along with PCIe6), we will start buying those as well so that it is 800G everywhere. Let's see how long that lasts before it is considered slow... heh.
Most markets can disseminate their feeds on 10G effectively. This isn’t true of the major US exchanges.
Queuing theory has many many bad things to say about actual saturation.
[1] At least in my time in the front office.
[2] For example a very common pattern at the very low level for a marketdata subscription is when you subscribe to marketdata for some symbol the system will actually have a double buffer where it writes into one slot and you read from another slot and every time you read it switches the slots around. This means you can generally accept marketdata as fast as it arrives and process it when you can and you will always get the most recent packet when you ask for the next packet.
One thing to realise about marketdata specifically is it's really different from other low-latency situations most people are familiar with (netcode in a game for example). As I mentioned before, it's not that big a deal generally to miss a few packets - the thing that is a big deal is to make decisions based on stale data. So you're not generally trying to reconstruct the full state after a drop- you just want the freshest current packet as fast as possible. If/when you need to reconstruct state you can make specific requests if needed.
Moving to our own multicast hardware not only greatly improved performance, but also greatly simplified the design of the system. We required specialized expertise, but the overall project was reasonably straightforward. The biggest issue was that now we had a really efficient packet-machine-gun which we could accidentally point at ourselves, or worse, can be pointed at a target by a malicious attacker.
This 1-N behavior of multicast is both a benefit and a significant risk. I really think there is opportunity for cloud providers to step in and provide a packaged solution which mitigates the downsides (i.e. makes it very difficult to misconfigure where the packet-machine-gun is pointing). My guess is that this hasn't happened yet because there aren't enough use-cases for this to be a priority (the aforementioned video use case might be better served by a more specialized offering), but exchanges could be a really interesting market for such a product.
It would be pretty efficient to multi-cast market state in an unreliable way, and have a fallback mechanism to "fill in" gaps where packets are dropped that is out-of-band (and potentially distributed, i.e. asking your neighbors if they got that packet)
I would love some feedback!
This is a lot harder to do when a server is virtualized somewhere on some rack on EC2. Exactly as mentioned, people will try to optimize by spinning up/down instances as close to the exchange server as possible. Customers will be unhappy because they can’t prove that it’s fair, even if they have the closest server.
Overall great, thought provoking writing btw
There are bare metal EC2 instances.
which sounds like what AWS Transit Gateway is
IMO, if you have a problem with limiting it to 5 seconds long quanta, you are doing something wrong.
I don't think the reason has anything to do with price discovery, it's just because exchanges want to maximise their trading fees. Continuous order book trading leads to more trades and hence more profit for the exchange.
If you want to kill HFT you can do it directly via very very small transaction fees. But guess how popular that is...
You mean you can never completely eliminate the advantage? But mostly eliminating it might still be useful?
Suppose the rule is that if you get your request in by 01:23:45 then it gets handled in the following 5-second period and the response is sent out at 01:23:50. Does someone (A) who finalises their request at 01:23:44.9999 and gets the result back at 01:23:50.0001 have an advantage over someone else (B) who has to finalise their request by 01:23:44.8 and gets their result back at 01:23:50.2? Yes, certainly, but it doesn't seem to be much of an advantage ... So person A can take account of exciting news that arrives at 01:23:44.9, while person B can't, true, but when it comes to reacting to other trades, person A has 4.9998 seconds to think about the news, while person B has 4.6 seconds to think about it, which doesn't seem like a huge difference. Compared to how things work today.
A per-message would probably significantly affect existing strategies and greatly increase spreads, but I don't think it would prevent all forms of ULL trading.
[1] But even there exchanges offer rebates, if not outright incentives, for market makers to provide liquidity.
You could match what you can distributed equally and leave the rest unsettled.
You could let people decide whether to roll-over the partial bid into a new bid on the next clock or to cancel unsettled.
You could clock to something both very fast on a human scale (50ms), quick enough it'd still feel instant but slow enough that it could reduce HFT silliness and need for extreme low latencies.
Equally per market participant? Do large participant like banks trade same amount as retail investor one trade at a time? Per quantity? HFT will time the end of the interval and decide to place a large order or not.
I'm not sure I understand the problem with "waiting" for the end of the clock. The pool wouldn't be public so you couldn't get knowledge inspecting the pool. All bids and offers would be published on the clock and settled by weighing all the bids and offers against each other and matching by volume.
The trickier issue is what happens in this scenario (assuming limit orders):
Person A bids for 500 units @62
Person B offers 100 units @61 Person C offers 400 units @60
Clearly there needs to be full settlement, we have a bidder who wants to buy 500 units at a price which sellers are happy to sell at.
Correct me if I'm wrong, but in a traditional market it would depend on the order they came in.
Here we would need a formula to work out the correct settlement price. Intuitively this ought to be somewhere just above 61. ( If it were just two people, a bid at 62 and an offer at 60, you could intuit a fair settlement would be 61. )
I'm sure fair formulae can be derived however.
I guess it is possible that there are remaining marketable orders that never fill because of an imbalance one way or the other, but I doubt that ever happens in practice.
Quite the opposite, thanks to the tough competition the market makers are setting the bid/asks spreads as minimal as possible. Which leads to less costs for human investors, pension funds, insurance companies etc.
I used to be a market maker in the 90's before HFT took off. The margins we kept sometimes felt like a rip off but customers had no other choice but to accept them.
People who ask for transaction fees, forced delays in executing or whatever, tend to forget that these force market makers to increase their spreads, which means customers eventually pay the price.
It's not automatically the case that the disappeared margins & thinning of bid/asks have been shared equitably between the trading firms and customers.
Take two exaggerated markets for example:
1) No HFTs: The customer wants 100 shares in Company A. The shares are available on two exchanges, one at $100, and another at $105. A market maker charges the customer $5 to access the 100 shares at $1 each. The customer pays $105. The market maker earns $5.
2) With HFTs: The customer wants 100 shares in Company A. The shares are available on two exchanges, one at $100, and another at $105. The customer clicks "buy" on their trading platform, the HFT races to the $100 shares, and purchases them, then fulfills the order at $105. The customer pays $105. The HFT firm earns $5.
For the end-customer, all that's happened is the margin goes to another firm. The consumer still has no other choice but to accept these transaction fees. There was arguably a need for HFTs to reduce the market-makers exorbitant fees in the 2000's, but that requirement has been served, and the technology now exists to remove both from the market entirely.
HFTs are a rent-seeking entity interjecting in a market which, at least in theory, exists to most efficiently allocate capital to the productive benefit of all.
How the price improvement gets allocated is complicated. Some of the price improvement goes to the broker (in the form a payment-for-order-flow) and some goes to the actual investor (you). But in either case the retail investors are strictly better off.
The UK's Financial Conduct Authority:
>We use stock exchange message data to quantify the negative aspect of high-frequency trading, known as “latency arbitrage.” The key difference between message data and widely-familiar limit order book data is that message data contain attempts to trade or cancel that fail. This allows the researcher to observe both winners and losers in a race, whereas in limit order book data you cannot see the losers, so you cannot directly see the races. We find that latency-arbitrage races are very frequent (one per minute for FTSE 100 stocks), extremely fast (the modal race lasts 5-10 millionths of a second), and account for a large portion of overall trading volume (about 20%). Race participation is concentrated, with the top-3 firms accounting for over half of all race wins and losses. Our main estimates suggest that eliminating latency arbitrage would reduce the cost of trading by 17% and that the total sums at stake are on the order of $5 billion annually in global equity markets
https://www.fca.org.uk/publication/occasional-papers/occasio...
The University of Michigan's Economics department:
>We illustrate this process and the potential for latency arbitrage in Figure 1. Given order information from exchanges, the SIP takes some finite time, say δ milliseconds, to compute and disseminate the NBBO. A computationally advantaged trader who can process the order stream in less than δ milliseconds can simply out-compute the SIP to derive NBBO,a projection of the future NBBO that will be seen by the public. By anticipating future NBBO, an HFT algorithm can capitalize on cross-market disparities before they are reflected in the public price quote, in effect jumping ahead of incoming orders to pocket a small but sure profit. Naturally this precipitates an arms race, as an even faster trader can calculate an NBBO* to see the future of NBBO, and so on.
http://strategicreasoning.org/wp-content/uploads/2013/02/ec3...
The Bank for International Settlements:
>Conservative estimates suggest that at least 4% of dark trading occurs at stale reference prices. High-frequency trading firms (HFTs) almost always benefit from such stale prices, being on the profitable side of the trades between 96 and 99% of the time. Furthermore, stale trading does not happen at random but is driven by the behaviour of HFTs. HFTs as a group almost never provide marketable liquidity in the dark and rather behave strategically to exploit their speed advantage by submitting marketable orders to execute against stale quotes.
The specific strategy currently employed by HFTs is somewhat immaterial in the broader context of a discussion about front-running. For as long as a firm can legally front-run the market with any strategy, it can undermine the market and risklessly extract profits.
Bids and offers are collected for auctions that happen at regular known intervals, for example every 15min.
I think it’s pretty uncommon to do them every N seconds.
A common pattern is to collect quotes before the market open, do an “opening auction” to set the opening price, and then switch to continuous trading for the rest of the day. If trading in a stock ever pauses (which can happen for a variety of reasons) then another auction occurs when trading is restarted.
Doesn't seem like a profitable move for the exchanges. They make more $$$ with on prem setups.
I suspect the provider would end up with a plan where traders can get servers that all have the same network distance from the exchange’s nics (down to the same length of fiber).
Alternatively they could just stick a bunch of Outpost racks in the NYSE/NASDAQ data center and create an 'eXn.8xlarge' instance type and charge 100$ an hour.
NYSE machines hosting the trading server fail: presumably they have hot backups they're ready to switch to but that takes time and will interrupt trading during the cut over. Not to mention that not all failures are hard failures, what if the NIC is downtrained to a lower speed, RAM is slower than it should be, or a single hard drive storing important data crashes? Lots of interesting failure modes. When the NYSE owns their own machines they can handle these cases directly. When they don't and Amazon is responsible for repairing these machines it might take a lot longer to get things fixed. I hope NYSE is thinking about hardware failures and building a system to check performance of their trading servers before letting them become the active host.
Thinking about failures on the side of the traders: basically if they get unlucky then there could be delays as Amazon rerprovisions them replacement servers in the case of failures. This likely impacts what trading strategies are viable, and could cause them to lose money if machines fail at unlucky times.
ULL and currently HFT seems to be very useful for market making (buying the ask and selling the bid and profiting from the bid-ask spread making parts of a cent per transaction, done a few million times a day), but there are other uses for HFT. One of them would be to execute very big orders over time to instead of drastically rising the price of the security they can get a better cost basis by performing a set of trades, letting the market absorb the impact and continuing with the order.
The thought of having the market in a cloud provider like AWS scares me! Although I’m sure that AWS might have pitched the idea already. If the markets could be controlled by a private company that could schedule “maintenance” at convenient times for them, that sounds like a recipe for market manipulation bay trillion dollar company. Sounds like something the SEC wouldn’t stand for.
If you list a buy or sell order it just has to be in force for some period of time, say a minute or something.
HFT shops will say this would reduce liquidity, but it would only make clear what real liquidity was in the first place.
Also, what's wrong with hft?
If no, knowing that there will be no liquidity in the very last minute before financial statements, will you put an order 2 minutes before?
This won't propagate forever, but is going to bring very complicated changes.
> Hogging instances might pay, if this stops competitors getting good hosts. This would eat into PnL and is also wasteful on energy.
Aren’t reserved instances cheaper than spot?
> Bad players could do a ping test to many thousands of EC2 instances, find those which are also at very low latency to their good boxes (assuming these are competitors), and DDoS them during trading hours to hammer the hypervisor’s NIC. This would result in critical overhead occurring for competitors sending orders out.
Leaving out the logistics of how someone could do this (why are your instances reachable from the internet?), wouldn’t you have a good case with your exchange to get them kicked out?
Would you know who is doing the pinging? The cloud provider would know what account was pinging, but somebody doing this as a trading tactic would have the resources to churn through aws accounts as quickly as they are banned.
Front running the market in any manner should be illegal.
Markets already conduct an opening and closing auctions and conduct an auction to resume after a volatility break (what people often call a "circuit breaker" in the press although it's a volatility break) so this would not be as much of a technological lift to implement this as it may appear.
How it works from a practical perspective is the exchange suspends matching for a period (so say 30mins) but order placement still works. Then when the market comes out of suspension a single print runs to uncross the order book, and everyone who submitted an order which matched gets executed at a single price. So as you say timing arbitrages of the current kind are effectively impossible. So in the case of a rolling auction you would do that print and then immediately suspend matching again and do another auction.
Here's some background on how auctions work in financial markets in general but it's not the specific paper I was referring to https://www.princeton.edu/~jkastl/auctions_finance.pdf
*someone who is not a market maker
I “get” cloud in a lot of circumstances but it doesn’t seem to make much sense here.
That's essentially what you're buying from a cloud provider. Most of the time its not so much renting the hardware as renting their labor in maintenance.
That is assuming your hardware needs don't have a wide enough variance from time to time (scale up/scale down)
Net-net, it still benefits startup quant shops and sophisticated independents. Most retail isn't doing HFT or really any quant. But for people wanting to have their own shops, this is a better version than having to build hardware and colo.
Sometimes this may be making pennies (x N shares), other times it may be quite substantial.
These traders do it at high frequency. Imagine $0.01 x millions at a time.
How would you make money off that?
Quicker can be on the scale of years to milliseconds, depending on what is being exchanged and amongst whom.
If you're a market maker, you (usually) want your orders to be selected. A common market making strategy is to issue a buy order a bit lower than the last executed price and a sell order a bit higher than the last executed price with the assumption that there's a lot of random and small price motion up and down. If you can consistently process order fills and update/replace your orders in the book faster than the other traders, you'll get more of the trading volume, and other traders will have to compete with you on price. For some stocks where the minimum price increment is large relative to share price, most market making traders will converge on the same buy and sell prices, so latency is it.
There's also value in responding to filled orders in one venue at other venues. Many stocks have a 'home' exchange, but trade at many exchanges, if there's a significant price movement at one exchange, other venues will quickly follow, but if you can follow quicker than most, you can execute against the now mispriced orders on the book, etc.
This is a common sentiment, but the reality is that increasing market participation is good for everyone. Yes, even retirement funds benefit from the presence of market-makers. Liquid markets allow for better price discovery and cheaper transaction costs.
Nobody was helped by the 200 nanosecond thing that the machines did when the marked opened (except the owners of said machines, of course)
No. NY4 is in Secaucus. NYSE operates out of an ICE (NYSE parent co) owned facility in Mahwah about 25 miles north of there. They managed to pick out the one big US equities exchange operator _not_ running in an equinix facility.
Sorry but this whole post sounds like someone who is sort of HFT adjacent but doesn't really know what they are talking about. Sending orders at "09:29:59.9999971 at the hope your order arrives at 100ns past 9.30am." What?
This literally does happen, though. One of the things the hyperscalers have convinced the world is that precise time is hard. Precise time is easy if you are willing to pay extra for your hardware. Sub-10-ns precision is unremarkable when you use PTP.
However, sending things just a hair early for scheduled events to catch an exact time is a pretty well-known trick at this point. I remember complaining to the exchange that their clocks weren't precise enough for this to be reliable.
Also: how often have you guys seen the stock market being down? What's the "x nines" availability of, say, the US stock market and US options feed?
Now: do we wanna talk about the various cloud outages that made the news? Sometimes lasting hours?
Also what I've seen with the cloud is websites are now displaying spinners everywhere, for the myriad of not-low-latency-at-all microservices often taking seconds to respond. And that'd be on an ultra low latency fiber to the home setup, with 2 Gb/s down (and the ISP really supporting that).
Why the heck do I have to wait seconds for oh-so-many things to display in my browser, on a last gen Ryzen ultra-speedy machine, with a super fat and low-latency Internet pipe? The worst offenders being all those banking websites showing a balance of 0 instead of "-" or "n/a" while fetching my info: nearly gives me heart attack every single time.
I take it it has to do with micro-services all contacting shitloads of other micro-services, all living in the not-low-latency cloud. The problem being compounded by an army, a generation, of programmers who have never learned anything about optimization or latency and who solve every problem they have with the only hammer they have: the cloud. All these programmers know are JSON (or, worse, XML)...
I mean: JSON vs 40 Gbit/s of interrupted bit-packed binary feeds? How could these two world ever reconcile?
Now I don't do HFT but I do trade options and I do it through a desktop app and that app also offers an API through which I can fetch prices, send orders, etc. It's a good old Java app. And it's more advanced than any website I've ever used.
Can we please not enshittify everything with countless micro-services and JSON files in the cloud?
100 Gb/s is possible on AWS via Direct Connect.
Well, it goes down every afternoon :p and it's down on the weekends. More like seven sevens than nine nines.
It's been a while since I noticed a story about a significant stock exchange disruption, but there's a lot of things going on there. Tickers are largely independent, so it's easy to shard, and exchanges do shard them; (operational) trading outages often affect only a single stock, or rarely a set of stocks. There are procedures for administrative trading halts on individual stocks or the whole market and procedures to resume trading during the market day. There are also procedures for resuming trading after an operational error halted trading; I'm sure most traders don't like brief outages, and exchanges do their best to avoid them, but they can happen and be resolved with out a lot of confusion because the procedures are known and manageable.
There's also redundancy. If one exchange is having difficulty, there are many others that likely still work. Whole system events are usually not operations issues, but trading issues --- one or several participants placed weird orders, the exchanges processed them, and weird things resulted. 'Circuit breakers' have been designed in to pause trading when this happens. This is an intentional outage, and it's ok because it's intentional and the parameters are known.
Exchanges do provide very limited special conditional execution instructions such as peg orders or stop orders, but it seems like a hard problem for them to support anything more sophisticated and general.
> Also if two people want to make the same trade, who gets it?
Whoever pays more currently, right? So it can be uniform random and folks can pay for better than random chance, etc.
It's the same hypothetical "EC2 model" without the meta-game of trying to get closer to the cores the exchange runs at a given time.
Can you elaborate on the downsides besides security?
If a third party provider (there are several which offer this service) manages the rack, router, host infra, and provides you with the virtual guest OS, or going one step further with the application into which you load the strategies, then it's much easier to "right size" the infrastructure you actually need.
As others have pointed out, the pain point is multicast into the VM. Drop a packet = you lose money.
All that said, some trading is going the other way, RFQ based flows with 30 second quote lifetime, which is perfectly suited to the cloud and it could well be the nature of what is traded moves away from volatility based products (options) to value driven (bonds, funds).
Is the stock market for investing in a company, or for extracting money from momentary fluctuations? Those seem to be mutually exclusive.
With the exception of the IPO and buy-backs, the stock market is NOT for/about investing in the company.
When you buy IBM on NYSE, IBM doesn't get anything. Similarly for selling stock. If IBM doesn't see any money from a trade, how is that trade an "investment" in IBM?
Stock trading is trading partial ownership. That's very different from investing.
In other news, there's no money in the stock market. The money that you pay for IBM stock does not go to the "stock market". It goes to whomever owned the stock that you bought.
The exchanges should add random amounts of latency to every trade.
Even large funds like Vanguard have said that the current system has reduced costs in net.
^ excellent way to put it