Case study: Algorithmic trading with Go
polygon.io
polygon.io
"This aspect, the platform itself, seems to be often overlooked in most discussions. Many conversations revolve around strategies (mean reversion, trend following, linear regression, etc.), and backtesting, without fully addressing the practical mechanics or logistics of strategy implementation, particularly in the context of live, intraday trading."
I'm glad you had fun, OP, but also I think I can shed some light on why most people discuss strategy.
Trading is a perfect storm of ridiculously high tech, ridiculously complicated, ridiculously regulated (Not over-regulated, mind you, this isn't a value judgement. But the amount of regulation is extremely high.), and ridiculously competitive.
But that said, it's the last bit that drives it all. Since it's so competitive, even though building an order entry system, and a risk system, and a position-tracking system, etc is a huge accomplishment (again, congrats OP!), it's table-stakes to even dip your toes in the pool here. Trading shops can attract top talent and robust, bespoke trading systems are basically cost of entry.
So people talk about strategy because everyone already has the table-stakes stuff and are now trying to make money with it.
It doesn't help, too, that lots of market participants aren't even playing the same game. In HFT, we operated on trades with alphas that lasted a few seconds, where races to entry/exit were battled in shaving nanoseconds off FPGAs being able to shoot out orders and microseconds off wireless networks flying market data around new jersey. Meanwhile, banks are more concerned with elections and geopolitics than they are about the weather in Carteret. (Rain = no microwave network for the day). And then there's a million strategies in the middle with alphas that last from hours to weeks.
So it makes it really hard to even speak the same language to each other when talking in common forums.
It's a fun world. I miss it sometimes.
HFTs are definitely playing a completely different game. I was reading about the exchange architectures and how things are actually wired. I'm getting my data from SIPs while HTFs are directly connected to the exchanges [1]. I'm transacting in seconds and they, like you said, are transacting in microseconds, so there is no comparison. Which, in a way is actually nice in that I'm not really competing with them. Or, maybe I am but I can still make some money. haha.
Cheers and thanks for the awesome comment!
[1] https://www.researchgate.net/figure/Latencies-in-the-Electro...
The strategies though are where the discovery is. There are a few strategies that are well known and still profitable but those are largely consolidated to the biggest firms. For everything else it’s a discovery process. And done strategies are only profitable for very short regimes.
I miss it sometimes too, but so much has been consolidated it’s largely a big firm world now.
More straightforward, heh, sure. But still damnably complicated. Which just goes to show how much money and how much engineering talent is invested in this world that these things are so taken for granted.
What did bother me, and was acknowledged by my coworkers, was how much top talent was being pulled away from productive tasks to essentially wank around. Not just technologists either, but all sorts of mathematicians and scientists were drawn to the flame.
We as a society have managed to allocate so many of the “best and brightest” to either fintech wankery or placing ads in front of eyeballs. It’s enough to make one want to give up on capitalism, except that everything else appears to be worse.
There is 0 social impact. That’s the downside of course - but hey, how many jobs out there are really having any kind of positive social impact ? Not 0, but close to it.
Intraday financial games are zero-sum. What HFTs gain, they leech away from mutual funds and pension funds and retail investors and market makers who operate over a longer horizon.
Regarding social impact, the world does have some demand for liquidity and price discovery. Providing those services is both essential and extremely difficult. It's definitely not the most social good I could be doing with my talents, but I think it's weakly positive.
Those things have driven a load of proprietary and open source tech that helps everyone else.
That depends on how you measure it.
Obviously this would mean that countries that previously relied on subsidised labour from those third world countries would have to start paying up. I'm not suggesting that it is a zero sum game however.
The problem with ads is not that they're useless, it's that it is an industry prone to scams and grift. (Because doing advertising right requires all sorts of actual science, and ain't nobody got time for that when there's money to be made.)
t. Worked for 20 years and adtech.
It's nothing to do with "we as a society". I'm a quant trader and know many others, and the vast majority are in the industry because we care about making money not some leftist save the world crap. Even if socialists managed to completely destroy the financial market, we'd just find another way to make money without trying to save the world (e.g. like the mostly corrupt officials in Russia and China when they were communist). "Society" can't change human nature; even mass brainwashing on the scale attempted by Maoist China failed.
I.e. the best and brightest will always work where they want to work, not where leftists like you want to "allocate" them.
Cut the fascist bullshit. Everyone can comment about how society should be run, even people that are very successful financially.
If you try to exclude greed from the design of your societal system, it will immediately fail. In large numbers, economics shows us that altruistic people wash out of the model and everyone operates in their own self-interest (greed).
> Cut the fascist bullshit. Everyone can comment about how society should be run, even people that are very successful financially.
It's not financial success that precludes someone from commenting on how society is run, it's blatant, deliberate anti-social actions. This is the foundation of a social contract. It's the same reason you'll get locked up if you go around punching people in the face.
Also, I don't think you know what "fascist" means.
It doesn’t seem like fascism to you because you’ve done some mental gymnastics to pretend you’re not just crushing intellectual opposition, but that’s all it is in the end. Op is not punching anyone in the face. He/she is just blatantly unsupportive of socialism.
Please, I want to know; what benefits do the actions of the op have for "societal planning"? The person he was replying to was at least honest in admitting that this behavior "served little social good".
Once again, I think you'd do well to actually read what fascism actually stood for. The propaganda of fascism had a lot more to say about nationalism, "shared history" and the good of corporations than it did to say about things being "anti-social" or even mentioning "selfish greed".
Largely, though I receive 1 or 2 job specs every week for start ups with the keywords 'hft' and 'low latency'. Admittedly there's going to be duplication there if you read them closely.
I think it's a bit of a myth that (ignoring FPGAs) that writing a low-latency software trading system is a time/cost expensive process. Anecdata = I worked at two firms where we did a rewrite from scratch with teams of 5-6 people and traded in the market within 3 months. I'd argue a senior dev that's been around the block a few times could achieve similar when you remove corporate politics, and bikeshedding over design.
The big firm part is paying for multiple quants at $200k++ to come up with strategies and historic market data access for trading models. Small firms are getting backing as long as the co-founders are 'ex-CxO from MegaCorp'.
This depends a lot on the complexity of the trading system and the trading venue specifics. A system to trade single stocks or futures can be built, certified and running in 3 months. A system for options market making will take a lot longer.
The big costs for small firms are historic data (if you don't have any), colo, distance to exchange, and number of connections.
From the number of job specs I see, it feels like the HFT/low latency market place is healthy enough that there are always new firms appearing. It's competitive, so it's hardly surprising that if someone has new ideas they'll find a backer.
In the old days before HFT, you weren't sure you'd get the best price. You'd have to rely on a broker to make sure that happens, but as a retail trader you generally got a worse price/out of date price.
Nowadays with HFT you can get pretty much the best price anywhere. Those <1ms HFT traders make that happen. It's the efficient market hypothesis in effect, made possible by HFT.
Now is it pointless to shave off even more ns in the all out war to grab a piece of the order flow? For retail and institutional investors, at a certain point, yes it's completely pointless.
But it's also just pure capitalism at work. HFT firms compete against each other, and the competition is about speed to provide the best price and volume. If you try to regulate with something like enforced delays, then what do you compete on instead?
Capitalism necessitates efficient markets (and efficient markets necessitate HFTs) so any criticism of HFTs is a direct criticism of capitalism as well. I mean this isn't really a problem that's specific to HFTs -- there are just a lot of jobs that we can perceive as providing no value or even negative value (jobs that are possible specifically due to capitalism e.g. payday loans, 2008 style trading, certain scams or predatory practices etc.)
I think we can all agree that this is a flaw, and we're not criticizing capitalism to replace it with something else, but rather just recognizing this as a problem. At the end of the day, a useless job is a useless job even if it exists solely because of capitalism.
People are saying this because, HFT sounds similar to 'crypto mining'. That's people with best infrastructure, the 'big-guys' -- win. While leaving out the retail investors as broiler chicken, pumped with 'drugs' (by influencers) to spend more on imaginary assets, so that they can be used for 'food' by these 'big-guys'.
There are different influencers for retail investors vs crypto. In retail investing there are promises of 'retirement paradise', actual tax deductions, the Jim Cramer-like people (at least what I heard in US)
For crypto investing the influencer are different, the geography is wider. A promise to participate in markets if you do not live the country that has adopted US/UK-based financial services.
- - - By the way, I think the markets will still have liquidity if there is a rule to wait, for say, 30 min before a stock that was just recently bought -- can be sold (unless by a clear fat finger mistake)
This rule will cause the HFTs to stop existing in the current form.
Do any regular pedestrian middle class retail investors EVER actually earn enough to retire on? Or do anything meaningful with? If so.. what is THEIR secret sauce, since it's not HFT...
a)Members of political elite that get insider trading stock tips. (illegal of course). The number of folks in usa congress and senate that become 'very lucky' investors after they join the rank, has to be amazing
b)Lucky
c) everybody else -- that looses.
Overtime, I would say last 30 years, the amount of 'influencing' retail investors to trap them into unreasonable actions had gradually increased.
So the percentage of folks going becoming the victims of the charade, will become higher.
Certainly if you concentrate on ( b ) you create the plausible deniability defense for the manipulators
You are profiting off workers as a middle man in the economy by doing HFT, and trying the justify it by some vague concept of the "correct price". You are producing nothing of value, merely taking away value before someone else notices it is there.
Our entire economy runs on "middlemen". Convenience has huge value to most people.
It's a bit like stock brokers - and why wouldn't we want stock brokers to operate at drastically faster-than-human timescales, because we all know the value of a company changes every nanosecond! And "flash crashes" create opportunities for investors to make huge amounts of money!
And just as crypto has poured money into GPU companies (providing opportunities for enterprising secondary market resellers of same) HFT has poured money into networking companies.
Short of creating the great firewall or helping governments slurp up all the traffic on the internet, what could be a more beneficial application of network technology?
HFT (like finance in general) has also gobbled up lots of tech grads, making it easier to find jobs at tech companies.
Now when I say the same thing about index funds people get all huffy
Regarding liquidity, I imagine it's possible that the index fund might allow you to withdraw money faster than you could sell on your own, but realistically I kind of doubt it.
Likewise with crypto. Personally I don't think that if you're calling an API over an Internet it matters if your trading bot is written in go or python (mine was in python). Use the language you're most comfortable in. The network and trade submission/execution at your broker will be 10x slower than your bot anyway. Unless the size of your operation approaches the size where you can get direct market access which seems to be reserved for big forms only.
That's just it. The whole premise is pretty absurd. The market, the actors, everything. It's so far removed from literally anything remotely human. It's the financial equivalent of an infinite sea of AI bots producing CVs and research papers which are only being evaluated and read by other bots.
If you step away from it all for a second, what the hell is the endgame of this whole hustle.
Anytime you create technology that sufficiently replicates the creators, you end up with the spirit of the creator embodied in the technology. So, of course digital brains are going to do weird things like crossword puzzles and sodoku at scale, because it's the same kind of useless shit we do to entertain ourselves.
You get into this conversation whenever you dive hard into cyberpunk, which is so many, many things. Even the Internet itself started out that way. The endgame started as a game, and it will end up being a game, played by our creations as odd mirrors of their creators.
I think there's a lot of people who subscribe to doing the same thing to save the planet, of which I have a keen interest. Solarpunk is the name of that movement, and it also has similar crazy ideas.
We're inventing digital brains. It's literally an architecture designed to be removed from being anything remotely human as it's a mimic or replacement technology for intellectual capacity.
If you really take a step back, the endgame is crazier than just bots producing content merely for other bots to consume (which basically describes the vast majority of scientific papers these days too, ironically). When you take that absurdity and multiply it by tens of thousands in terms of efficiency, the whole system we're building looks WILD and almost inconceivably strange to the way we do things now.
I know this is only tangentially related, so I appreciate your understanding that I already understood that and wrote this anyways. :)
Is it basically "do what MFT does, but faster", or is there any specific advantage like getting into an order queue with priority?
I've been really wanting to use Go, but as you say, much of the community is Python due to the data analysis strengths. To the detriment of the other things Python does do poorly.
Can you give some thoughts with your experimentation on the following from a Go perspective.
1. Supported TA libraries in Go. I'm familiar with TAlib (python), bloom, etc. - certain forks tailored to real time rather than historical (eg: no re-compute on ticks)
2. Data storage (article mentioned you're all in memory). I've been using S3 & ArticDB
3. If your in-data memory is treating you well for multiple TA calculations (example: in Python, you can compute & save pickled dataframes - and re-read those over longer time periods)
I had not looked at Mojo... thanks for the pointer. In the past, I've had issues with library compatibility with compiled derivative languages of interpreted ones (eg: Crystal of Ruby, etc, etc).
Know if Mojo directly uses existing Python ecosystem?
I've been using Polars, etc.
They say they do, including libs with C bindings.
I've been basically, just manually coding the algorithms from python into Go. ChatGPT is amazing at this. I really only just about 4 so it was a one time thing.
> Data storage (article mentioned you're all in memory). I've been using S3 & ArticDB
Yeah, I ran into issues and then was like what would be the fastest, then just went in-memory. I download all raw trades/quotes each night and store then into gob+lz4 compressed files. Then for backtesting and stuff I can load these in and build the aggrogate bars on the fly.
> If your in-data memory is treating you well for multiple TA calculations (example: in Python, you can compute & save pickled dataframes - and re-read those over longer time periods)
Yeah, I have a historical lookup table that I build nightly too. This gives me a reference point when I'm doing % change calculations and stuff. I should probably have mentioned that.
Can you share your approach for plugging in various strategies? I quickly learned that having a pluggable strategy system is tricky as it could span across multiple layers of the system.
Also, with backtesting, are you storing and replaying all the quote/tick data? or just using the historical aggregates?
For back testing, I download all raw trades and quotes, and put them into 1 file sorted by time (1 file per day). Then, I compress them using lz4. This allows me to sort of replay the entire market and build up all my intraday backtesting from the source. This took me a long time to figure out and build but has been so worth it. So, I have an off-line script that basically, loops through these files, and replays the market, and makes simulated trades, and then spits out what would have happened. There is a GUI for that too so you can go in an explore the trades and see what triggered the buy and sell. I have seen nothing that goes backtesting for intraday like this.
This is super inefficient but I'm just building the aggregates on the fly. I could probably cache them somewhere but it takes maybe 4-5 minutes to replay a days worth of trades/quotes and build all this so I haven't bothered yet.
I’d like to play with something like this but I couldn’t be bothered to build it all from scratch.
I've also got a few questions regarding the way you manage memory, so I might ask you a couple !
Excerpt from his site:
"Tech Trader is a fully autonomous trading system live with no human intervention or updates, now for over 10 years. It is unique from conventional algorithmic systems, not only because it actually is fully automated, but because it takes a "human" approach to markets. It is not quant. It is not stat-arb. It is not high frequency. It is a program that looks at stocks the same way a person does but with the cold discipline and infinite attention span of a machine. It is analogous to having a thousand independent traders each focusing on a single stock, as opposed to a single quant manager trying to make sense of a thousand datapoints. A person doesn't think through stats, correlations, or complex math models when trading, and neither does Tech Trader. Tech Trader leverages technology to do what human traders do at scale rather than approach markets from the point of view of an academic, mathematician, or scientist.
Since its launch in Dec 2012, Tech Trader has been trading live capital completely on its own, fully automated in the truest sense with no human input, no tweaking, no updates. It is, for all intents and purposes, an autonomous hedge fund, one of the first to truly trade unsupervised for years on end. Whereas many "automated" or "AI" funds may have a hundred scientists providing the actual intelligence behind the curtain, the creator of Tech Trader consists of just one person - a self-taught individual going by the gaming moniker pftq, who created the system at age 21 and has long moved on to other interests. "
I call bullshit. Reads more like a scam for trading signals.
https://www.sec.gov/Archives/edgar/data/1653903/000165390318...
I can't say any more about its validity. Ask your financial planner if techtrader is right for you :)
I have no idea: does IB send the price feed in real time? (they certainly send the data in real-time to TWS as it constantly updates right? But is the order book available through their API?)
Basically and even though I know this was published on polygon.io's blog, would that work by only using IB / TWS's API?
[1] https://www.interactivebrokers.com/en/?f=%2Fen%2Fgeneral%2Fe...
I've been going through this journey myself. I started learning on Tradingview, then bought Build Alpha to discover how to test strategies. I chose Portfolio 123 for my automated factor trading but had been working towards creating a program/basket trading system like yours that can act on intraday data.
I moved to long-term investment until I could build a simulator capable of verifying the correctness of my investment strategies using fuzzy testing ideas stolen from TiggerBettle.
I have almost two years of polygon quotes and trades for the whole market captured with a monotonic timestamp to be able to replay the data –and test the handling of polygon socket glitches.
I'm focusing initially on capturing the data in a way that allows fast replay and aggregation, similar to what Kafka can do with topics but in-process using zig and custom memory-mapped data structures. My idea is to be able to generate signals like VIX (once I add options data), ETFs, and indexes and hopefully be faster at doing so than others :), please HN folks, call me out here if I'm being too naive.
This has been a three-year learning process for me. I have been a retail investor for +10 years, but over the last three years, I've gone deep into learning algo-trading, drank del Prado Kool-aid, and read numerous trading and investment books.
I'm now focusing on my technical chops to build the engine to build order books for individual stocks, baskets, and indexes with realistic market prices. I aim to develop a system that can get as close to the market price in the next dollar bar as possible.
This has been a very lonely journey, and after reading the responses to this post, I'd love to connect with others on a similar path. Sending you an email!
I've used ATR bots for years, and would love to hear your thoughts on how what you're building delineates itself as a distinct product strategy beyond just the programming language GO.
I've used wonderbit, zigz, and some of the others.
Cool project.
I was curious in getting a websocket setup with Polygon after seeing your post, but I noticed it's $29 for first package with the websocket feature. Then I noticed that the package with real-time data is the Advanced package at $200 a month. I'll admit, real-time with unlimited API calls sounds pretty sweet, as my strats rely on to the second data. However, I don't know if I can justify $200 as my algos run no more than 78% win rates with a normal bankroll funding it all. I'm looking to save wherever I can while I build these things and get a passive income stream rolling in.
I would love to know which package you were using, as I didn't see it in the time I quickly read your post? Also, any pro/cons to that specific package? Any and all other details are welcome, and if you would rather respond to my email, it's in my profile. Thanks!
The last 18 months have been weaker than the first given the enormous structural shift in the market in this high inflation and rapidly rising interest rates environment, but we've still managed to deliver a return of +14.11% since the site launched in Jan 2022 compared to -7.83% for the SPX. We've managed to do it without any use of leverage and also with lower drawdowns as well of -16.48% vs. -27.57% for the SPX over that time frame.
It's also worth noting that every additional data source adds some risk of that data source being down or publishing inaccurate data during real-time signal calculation which can cause inaccurate signals, so in order to justify that risk, the external source must meaningfully contribute to alpha or better risk-adjusted returns.
None of them involve a human element in real-time. However, they are occasionally updated as new data comes in, but any updates only apply going forward so as to preserve the live trading history accurately (live trading start date varies by model from mid 2020 to jan 2022 with 2009 - 2020 being purely backtest for all models).
Am I right to doubt that something this simple generates any alpha whatsoever?
Kudos to you if you really sit on an untapped gold mine, but imho, there are some red flags that makes me not buy in.
However, the April 2009 start date is not actually random--it's the first start date for which intraday futures data is available for more than just front month contract. Several derivative indicators of the VIX futures curve are the most foundational to all the VIX-based models, and they simply cannot be processed without it. The VIX futures were only created in 2004, and I've scoured the internet for intraday data for more than just front month (can't create the curve if you only have front month data), and the earliest it can be found is April 2009.
It's like the minor leagues for algorithmic trading. It's fascinating.
The system keeps track of the gains/losses, so no cheating on the reporting.
You can authorize Collective2 to access your Interactive Brokers account, so that the trade signals are managed on your behalf.
It's been around for at least a decade, so you can see some longstanding performance numbers, but most systems just don't last that long.
https://collective2.com/leader-board
None of the leaders have been there for very long. One to two years. Showing that most system's alpha disappears fairly rapidly.
The equity curves are sporadic ~a few of the trades accounting for a majority of the gains.
- Data feeds and data ingestion. It can be a fairly independent component which collects data from different sources (might be even discussion forums) making it available to other components in some uniform format
- Feature generation. The source data is rarely used in its original form for decision making and having good (informative) features is frequently the primary factor of success. Moving averages is an example but nowadays this will hardly work
- Signal generation. Here some logic should be applied in order to emit discrete decisions and such models are heavily parameterized with thresholds.
- Real trading and order management as well as coordination of all activities.
The article sheds some light on the technological aspects and the general pipeline used to process the data and manage orders. Although it might be interesting indeed, I would expect more details about how to scale the solution and how to implement it asynchronously. Especially if it uses Go which has a special construct for that purpose - channels.
I understand that it is not the focus of the article, but having some general information about its trading logic and how to plug new and parameterize existing strategies would help. Some links at the end are quite interesting for me because I am developing an intelligent trading bot based on ML and feature engineering (https://github.com/asavinov/intelligent-trading-bot) for which such articles might be quite important
I'm just using go routines and channels to talk between them and then a giant mutex for locking. That's basically it. So, as new data comes in, it builds aggregates (tick based candlesticks) as needed, this then triggers the the BUY logic loop on that new data, if something is detected, that triggers a IB API order. It is dead simple and nothing complex in here. I've had upwards of 100 positions being tracked at anyone time and seems to just work. So, I haven't messed around with complex async logic too much.
I'm actually just hard coding the parameters right into the BUY loop. This probably sounds crazy but for a small setup like this they don't change that much. So, I can run some trades, tweak things, restart, and then test some more. I imagine if you were doing that in an enterprise setting you've have some formal language and hot loading and stuff. But, for me hard coding seems to work well enough.
If you process many positions and then choose top 5 candidates then do you choose one of them for buying or you can allocate resources between them (depending on some score)?
That logic is happening for 5500+ stocks all in real-time which is pretty insane. But, it works really really well. Go is amazing. At market open and close there is like 60k+ events per seconds across trades and quotes. So, that loop is processing like 60k events at times, and building each stock out, and then looking to see if we should buy/sell.
For example:
1. Do you have anything that limits the size of a single trade or position?
2. How do you measure and manage overall volatility/risk/VAR to your portfolio?
3. What kind of safeguards do you have to avoid catastrophic bugs? (For reference, see Knight capital and how a single bug brought down the entire company: https://www.henricodolfing.com/2019/06/project-failure-case-...) Given how fast your system trades, I imagine it must be difficult to visually spot these errors. (You did mention paper trading, but I wonder if you have anything else you want to mention)
Thanks so much for sharing your knowledge publicly. It's very much appreciated!
- No single bet can be more than 5% of all money. The small bet sizes are what really saves your bacon in that even if a few lose 10+% you're still fine overall. - I'm also limiting the amount of shares I bet and try to keep it in the low hundreds so that I get really fast fills. - I have something that stops everything if I lose more than 1k in a day.
I don't do anything to measure overall risk or volatility. Almost everything I'm in is highly volatile. I'm basically betting as things go up and then try to cash out. Sometimes you hit the top.
Yeah, I inspect all the trades at the end of the day, well and during the day, and try to feed anything new back into the system. This is 100% manual. But, like if a single trade loses $100 or something definitely I'm in there looking at what happened.
What is your 'win' rate? (Percentage of trades that make money as a fraction of overall trades)?
How many trades does your system make on average in a given day?
Do you hold positions overnight?
I could use Visual Basic, and it would be better than Go, Rust, or whatever it is out there, given the algo and strategy are flawless. Language is just a tool. It's great you used Go, but I think the title is a bit misleading - people think of it as some kind of advantage. It isn't.
And for HFT trading a language with Garbage Collector is not a great choice IMO.
I used Go to write trading algos that would find small windows of triangle arbitrages in crypto exchanges. Made me some money but the risk of a big loss made me stop pursing crypto trade and it required too much time and attention. It's a full time job from my experience. Reasons I could lose big at any given time if I scaled up the stakes:
- Exchanges temporarily pausing some specific crypto trading for N reasons (happened very often) while I'm in the middle of the arbitrage.
- Getting caught in the middle of a pump and dump event (also frequent)
- Any algo mistake that would perform excessives trades in succession would incur huge losses because crypto exchanges charge %.
Also, most cryptos have too low volatility. Every time I tried to scale my bot would start interferring oo much with the market. And it wasn't even much money.
I used Go because that's what I knew and for the non-professional trading I was aiming, C++ wouldn't have made much difference. My bottlenecks were network (1s+ per trade roundtrip) and chaotic unreliable crypto markets.
Just note that Java is used in HTF, but it's a different beast than our average CRUD Java. For example this article states:
"Essentially, we use a contrived form of Java that avoids all the Java constructs that make things go slow. We only use the constructs that are fast and efficient, and we avoid all the garbage."
https://www.efinancialcareers.co.uk/news/2020/11/low-latency...
I read like 5 or so books about trading, mostly the classic ones, and tried to toy around with backtracking some self-made strategies using MetaTrader: https://www.metatrader4.com/en/trading-platform
First week was funny. I kept thinking I found goldmines with fine-tunned stochastic lagging indicators, only to realize that I overfitted and an algo that made me rich when tested against a certain period of time would make me go broke when applied to another period.
I tried applying some defense against local maximum, but to no avail.
Best I could do is tie after paying exchange fees. And that took me endless weekends toying around with code, reading books and articles.
My hunch is that algos that don't rely on external signals (i.e. news) need extreme technical edge like High Frequency Trading have with their ultra low latency and premium exchange data feeds.
And I wasn't ready to dive into news sentiment analysis and other external indicators.
Therefore the title was actually on point.
At the end it is necessary to make a decision whether to buy or sell (and how much), which will compete with other decisions made based on some logic. Developing such a logic (strategy) manually is of course quite difficult. I developed an intelligent trading bot which derives its trading strategy from historic data:
https://github.com/asavinov/intelligent-trading-bot
Currently it works for cryptocurrencies but can be applied to other markets:
https://t.me/intelligent_trading_signals
> I've spent months on figuring out the best parameters for trading. Ended up this working only on historical data, while in reality it was totally different.
It is a typical situation. The whole problem is to develop a strategy which works for future (unseen) data. Even backtesting algorithms should be designed in such a way that future experience (data) does not leak to the past.
Depends. I've heard of places that use Java by creating a huge heap, minimising allocations to almost zero, and just restarting the app periodically before it can fill up. You can achieve very good performance doing this and you don't have to worry about memory bugs.
There are places that use Java and just preallocate all the memory they need at startup. Jane Street famously uses OCAML
It is an overwhelming lonely endeavor. With all my other projects, I've always worked on teams, although they have always been very small teams and most of my work was autonomous -- still there was the occasional meeting and stand-ups. I've been working on this for six months and thought about bringing a friend onboard for no other reason than to not be alone.
We all know how to load up OHLCV data and do basic math on it. Where does that gain us any edge, you know?
I would be interested in some kind of regular meetup as well to discuss new techniques and to analyze publicly disclosed strategies, etc.
I recently posted this article on HN... There is so much more knowledge to share:
The only way I can think of is to get data that they can't.
However if you start increasing scale to $1mm or $10mm, your buy or sell orders begin to actually move the stock price itself. You might not be able to successfully sell $10mm of stock without dropping the price, signaling others to sell, further dropping the price, cutting into your own profits.
I.e., if you're buying, the larger your buy order is (well, assuming a visible buy order) the more likely it is that liquidity-adding sellers will increase the price of their sell orders. Also makes it less likely that liquidity-taking sellers will want to trade against your large buy order, because (like everyone else) they'll tend to interpret your large buy order as a sign that the price is likely to increase, so not as good a time to sell.
You could of course use hidden orders to avoid some of those disadvantages, but hidden orders have their own set of tradeoffs too.
Obviously whatever trading bot you're running separate from the actual trading engine itself is somewhat proprietary, but it would be great for the community to get more of this type of software in the hands of other hackers.
Quantopian / Robinhood tried and failed, and the numerous clones since then have been somewhat sub par.
It works better for videos and code, IMO. You still have two huge portrait fields for documentation/webpages/etc, or 8 portrait-oriented quarters.
I usually run 8 portrait-shaped windows quartered on the sides, then two side by side on the middle, which are each approximately square.
I don't like my editor window getting too tall.
> 3 screens like that, put the middle (primary) one in landscape, so the array is like a big H.
Honestly my perfect screen for work would probably had something like 1:1 ratio, so I can have 2 nice columns of code that are tall enough that can be split horizontally if needed while still being useful.
Very interesting. We can facetiously say that ChatGPT is using you as a medium between setting up algorithmic trading in Go!
It did successfully grab the arbs but there wasn't enough juice to justify more work on it and I got a job in the meantime, so I open sourced the whole thing: https://github.com/abissell/cempaka
I do think that if Java can deliver on the combination of the foreign function/memory interface and value types, it might really start to look competitive for certain strategies which are just a bit too complex for the "do everything in the network card" approach. When the Aeron guys first implemented their protocols in Java, C#, and C++, C# was actually the fastest, which they attributed to the presence of both the optimizing runtime and value types.
Thing is, you will be fooled by bad data - you will find strategies that works in a backtest but not in production. And the reason will be because you missed some important finance concept (like taxes, dividends, stock splits etc).
In this field 99% of success in my opinion is knowing what you don’t know. And only model that small part that you know you know and are fairly sure about it.
Ps. By stop it now - I mean stop algo simulations and learn about those concepts, make sure you understand perfectly what data you are putting into your backtest.
2. Sometimes you get a cool api and think wow this would be fun, and next thing you know you've lost thousands on boneheaded trades.
I did something similar during the pandemic with Rust with the Polygon API (and instead of interactive brokers, I used tradier). Eventually I learned I actually had more fun building the thing than actually trying to beat the market.
Or do you believe there are more fundamental changes needed so your app can trade in shorting as well?
After a bit of digging, I found this Q&A from the 2019 shareholder meeting: https://www.youtube.com/watch?v=geRIJQJXRVo&t=17980s
And the meeting minutes in PDF form (see page 120, question #32): https://s3.amazonaws.com/static.contentres.com/media/documen...
The text from that document...
32. It’s easy to make 50% on a million, but much more difficult on larger amounts
WARREN BUFFETT: Station 9. We’re just about — yeah, we’ve got time for a couple more.
AUDIENCE MEMBER: My name is John Dorso (phonetic), and I’m from New York. Mr. Buffett, you’ve said that you could return 50 percent per annum if you were managing a one-million-dollar portfolio. What type of strategy would you use? Would you invest in cigar butts, i.e., average businesses at very cheap prices? Or would it be some type of arbitrage strategy? Thank you.
WARREN BUFFETT: It might well be the arbitrage strategy, but in a very different, perhaps, way than customary arbitrages, a lot of it. One way or another, I can assure you, if Charlie was working with a million, or I was working with a million, we would find a way to make that with essentially no risk, not using a lot of leverage or anything of the sort. But you change the one million to a hundred million and that 50 goes down like a rock. There are little fringe inefficiencies that people don’t spot and you do get opportunities occasionally to do, but they don’t really have any applicability to Berkshire. Charlie?
CHARLIE MUNGER: Well, I agree totally. It’s just you used to say that large amounts of money, they develop their own anchors. It gets harder and harder. I’ve just seen genius after genius with a great record and pretty soon they’ve got 30 billion and two floors of young men and away goes the good record. That’s just the way it works. It’s hard as the money goes up.
WARREN BUFFETT: When Charlie was a lawyer, initially, I mean, you were developing a couple of real estate projects. I mean, if you really want to make a million dollars — or 50 percent on a million — and you’re willing to work at it — that’s doable. But it just has no applicability to managing huge sums. Wish it did, but it doesn’t.
CHARLIE MUNGER: Yeah. Lee Louley (phonetic), using nothing but the float on his student loans, had a million dollars, practically, shortly after he graduated as a total scholarship student. He found just a few things to do and did them.
It seems that this is the key to your approach. How is this part achieved?
Gaining 0.5%/day 60% of the time and breaking even 39% of the time looks great until you run into the 1% of the time where you lose 50%.
1.005^260 = 366%
and
1.01^260 = 1329%
In case anyone sees "0.5%" _daily_ and thinks low risk.
[0]: Link to my (non-monetized and WIP) blog where I keep a collection of excerpts: https://sileret.com/projects/warren-buffet-shareholder-lette...
You might want returns that aren't correlated just to an index - this is a major reason to look to invest in (say) a hedge fund.
Mostly because you'll never be able to respond to each tick as by the time the tick gets to you the market has moved.
Use the language that you know and can work with the fastest. For retail trading GC vs non GC will never matter at all.
Also I wonder how can there be changes in the price of a stock after market if the exchange has closed? Isn’t the whole point that trades need to happen for stocks to get a certain value?
You don't have to use the system I am building, but it's worth thinking about that design.
This is a general question, I'm wondering is there any good framework/wrappers out there that one can learn from to code up a complex trading application?
Like dealing with all the asynchronous nature of process/submitting trading and quotes messages.
THE GIG ECONOMY COMES FOR HEDGE FUNDS
Platforms that offer money managers the freedom to build a business and maximize their return on performance while removing the hurdles of launching independently could change financial markets.
If the trend continues it could have a big effect on financial markets by making it easier for a wider assortment of unconventional managers to rise in the industry and offering investors better and cheaper access to them. …
Well-received start-up ClearAlpha Technologies has moved closest to the gig model. Its first offering is a commingled fund apportioned among its managers, but it has the platform to act as an exchange, matching investors to individual managers or customized portfolios of managers, cutting out all the expense of intermediaries.
Source:
-- https://www.bloomberg.com/opinion/articles/2023-06-09/the-gi...
Reprints in case paywalled:
-- https://www.washingtonpost.com/business/2023/06/09/the-gig-e...
-- https://www.garp.org/risk-intelligence/technology/brave-new-...
You no longer have to have graduated as finance into finance. (In fact, we prefer if you graduated with some other math modeling heavy emphasis, think turbulence and flow, or actuarial modeling.)
A bit more about us, although this article is about our first fund mentioned above, not about the firm co-founded by Brian and I that owns the fund and built the platform it runs on: https://www.bloomberg.com/news/articles/2023-06-01/goldman-a...
If you're into this, we've come out of two years' stealth and are now hiring and remote work friendly.
A couple of pointers. One on data, one on that RAM usage.
First I’ll go with the RAM usage because this is hackernews and everyone loves algorithms.
—-
There are a lot of libraries out there that do technical analysis, and most of them are designed for batch processing. TALib is an example - it works on large data sets but is not appropriate for live trading because it repeats calculations over and over and over again. If you have 10000 datapoints and calculate indicators, it’ll calculate 10000 of them. Add one more bar, now you have a dataset of 10001 items, which TALib will calculate the indicators on from scratch. Or maybe you just feed the last 10000, and still perform that calculation over all of those, but save that 10001st oldest one. Either way, it’s bad. Same goes for every library I’ve seen, presumably because no one would open source a production grade indicator generation system.
The production approach to this is somewhat different. Turn your features into state engines. Most features are just running calculations that are very easy to perform once per bar.
Moving averages are a perfect example - for a moving average of N bars, store N items in a fixed size array on the stack and keep track of the latest index to be written to. When a new value (X) comes in, increment that index, and grab the value (Y). Modify your old mean by adding (X-Y)/N. Then write X over the old value Y (It’s a ring buffer).
For EMAs, it’s even easier, because you just need to keep track of a numerator and a denominator - nothing else is needed. On a new value X with the scaling factor A (such that A^halflife = 0.5), the numerator N becomes NA+X and the denominator D becomes DA+1. Divide the two and you’ve got your new value.
If an indicator depends on another, don’t recalculate it. Break the indicators down into fundamental calculations and you’ll often find a lot of redundant calculations being done. Rearrange it all so it fits. Automate that process if you enjoy that kind of thing like I do.
Most indicators can be handled this way. The indicators that can’t - are rarely useful. After all, what you’re tracking is the evolving state of the market, and if you’re doing gymnastics over an indefinite number of bars, it probably doesn’t mean much.
The end result is that you don’t end up accumulating memory throughout the day. You receive a bar, you throw it through your indicator generators (all of which using a fixed size of memory), then you discard the bar and wait for the next. Save state at the end of the day and load that on the following market day.
The result will be a speed up like you couldn’t imagine. I promise. I run an indicator generation engine in a container capped at 40MB ram on one CPU, and it generates hundreds in much less than a millisecond after the bar arrives.
—-
Now, onto data. I recommend cleaning your data. You have a screenshot of a Tesla chart in there with some funky highs/lows every so often. It has been a long time since I’ve worked with US equities (and gladly so, it’s a mess of a system!) but the following is the best of my recollection.
The trades you receive will come from several sources. For US stocks, there are several different venues that operate their own order books. These will operate as typical markets between the open and close of the day. By typical I mean that the bid and ask represent what you’d get if you market order instantly (which you can’t), and the market trades on them have to take from the bid and ask side of the order book (formed by people placing limit orders).
However. There’s also the ADF - the Alternative Display Facility. This is the DIY of trade reporting (and quote reporting, but no one does). If someone sells some shares to their grandmother for a low price in exchange for the recipe to her famous apple pie, the ADF is where they can tell other participants about that trade, manually, subject to fat finger errors, prices of weird fractions of cents, and very relaxed constraints on timing.
It’s also where dark pools post trades.
The problem with this is that this data has no direct relationship to the rest of the market. This is why, every so often, you’ll see those blips.
If you don’t clean them, they’ll play havoc with any indicator that uses highs and lows.
The other problem is that ADF trades - at least when I last analysed this very issue - are not rare. They make up a large fraction of trades. So my approach was to clean ADF trades more rigorously than the venues by matching them against prior prices from an ADF-free background.
Are you using number of trades or sum of trade sizes?
Because if the former, there is little distinction between me firing off two market buys of 100 shares each in quick succession vs me firing off one market buy of 200 shares, so your sampling shouldn’t be impacted by the difference.
I’m not altogether convinced by volume sampling. It’s an idea popularised by De Prado, but I’ve never seen it actually work in practice. It makes you trade more when the market is going through turmoil, it makes you trade more over time (as volume per day generally increases), and I haven’t seen any evidence of the importance of information content.
If you’re trading based on patterns in the market, it’s easy to lead yourself to believe that you’re predicting the market. That isn’t the case, though - the market is formed of many many independent people making guesses.
The thing you’re actually doing is predicting what other people are going to predict. If there’s an established pattern, like some moving average crossing another or a wedge or anything else like that, the reason it tends to complete is not mystical - it completes because other people see the pattern, think it’s going to go up (or down), buy (or sell), and then that has the effect of pushing the market in that direction (it also means that anyone late to the game can’t benefit from the movement).
As such, the strongest strategy when trying to use technical analysis to determine possible market moves is to use the resolution that everyone else is using. This is overwhelmingly time-based. There are people trading in the 1s regime, the 1m regime, the 15m regime, etc, and they’ll often stick to that and execute trades with a proportional rollout time and aim for a proportional profit.
If you pick just a random number of trades or amount of volume that suits you, you can easily find that you are out of sync, competing against no one in particular, and you’ll see that it’s impossible to find a pattern.
Many other people use price levels. So they’ll have their limits and stops at round numbers, or at percentage changes on the day, week, month, etc. So there is an argument for price bars too.
Thank you for your amazing comments. I really like your idea about indicators and saving state. I'll give that a try! Yeah, Marcos López de Prado is actually where I read about the tick bars and sampling at higher rates. You need like 2 phd's to read his books though. haha. I am doing this based on X number of ticks and not even looking at volume. I was using tick count as an indicator in that you can really see patterns when activity picks up. This sync issue is really really interesting and I'll explore this.
"The thing you’re actually doing is predicting what other people are going to predict." I think I've actually seen this in the data. In that you can see this mini-cycles almost when you really zoom into a fast moving stock. I'll check out. Both your comments are amazing and it 100% shows you know what you're talking about.
What's cool about this is that you can look at the IB TWS client and see things happening in real-time. So, it acts as sort of a sanity check. I know they have that gateway too but personally I like to look at the client all the time too. My workflow is to run the bot and the TWS client side-by-side and watch it make trades, see how things are moving, etc.
(It's also kind of off-topic, my apologies for that. To me it seems semi-related but would agree with mods' assessment if it doesn't match mine.)
---
I recently "broke up" with some extremely toxic "investors" that wanted me to do a fully parallel trading demo -- meaning it receives ticks for N instruments (I successfully got to little less than 300) and trades with each of them depending on strategy. All in real-time.
I got very far but the open-source libraries for Interactive Brokers are quite low quality in general and it was very hard and slow to progress (one of them couldn't even post orders, another used Mutex-es for "parallelism" which was of course not parallel at all, another one seemed to work well but only worked on servers of older versions compared to those I have access to, etc). I also had to gather code from separate places and assemble my own Frankenstein as I went along.
Eventually I muscled through but by that time I have drained all my savings, my tax fund and even got a new loan. And the liaison + the investors of course refused to acknowledge the demo was basically 90% done (couldn't do full parallel trading due to defects of the IBKR libraries I have used and I used like 5 of them, and was in the process of repairing 2 of them to unlock the said full parallelism). They refused to send a small pre-funding wire (we're talking something small like $30k - $50k, not millions; for the work that was done, namely months of professional Rust programming work, that's a -75% discount if we look at market rates).
They had all the proof and paper trail they needed to see that I was very close but I had to stop because I was literally about to be unable to pay rent and bills. They still did not concede. Obviously I picked a job and dumped them but I still have some regrets because the door is technically still open (they have not cut my access to the IB Gateway servers that they own), but other factors like now-ruined health are seriously getting in the way as well. Not to mention completely shattered trust.
They insist "if you just finish the demo we'll give you money" and used every gaslighting technique I knew about (and many I didn't know about, so I learned a lot about gaslighting from them, lol) to try and coerce me to keep working for free with zero guarantees of funding -- but I no longer trust such rich investors to fulfill a promise without a binding and legally enforceable contract; I am in Eastern Europe, they are in the USA, even if they sign contract and violate it I practically cannot do anything to them i.e. I can't afford to travel and sue. Also they insist to get access to the source code but swear to everything that's holy that they will not run away with it and never give me a penny. Which is exactly what I think would happen.
I had to draw the line at one point. To me it was red flags all around.
---
I made many concessions and I will not make another one until they do. I'll be finding a new job soon again since the previous contract was agreed upon to be for several months, I helped a team accelerate bootstrapping a business-critical product (and we did that successfully). And then maybe, just maybe, after I settle a bit, I might work an hour or two on this again during some evenings. Maybe. And that won't be used to provide the demo to these toxic investors -- I'll use it to have a proven implementation that I can pitch to other people. These guys don't deserve the time of day from me.
I have quite a lot of good code (mostly Rust, but also some Elixir and Golang -- I experimented a lot) that interfaces with Interactive Brokers and I have many building blocks of a good trading bot framework. That kind of creative and thorough work needs funding and peace of mind to be finished properly, of course. And I can't afford to invest so much energy on this without a stable income so I am shifting focus to that for however long it might take for me to feel comfortable to invest time and energy in this again.
If any business person wants to partner up -- give me a shout. Mail is in profile. I also wouldn't refuse fellow techies (or anyone else really) stopping by and telling me how stupid and naive I was -- I'll agree right away and say that they are right. Because truth is truth.
I admit I still can't get over the fact how close I got but I can't afford to count the stars while my livelihood is being endangered the longer I coast on savings. Dreams don't pay bills. And some dreams need more than one person to be fulfilled.
Rant + story over. Thanks for reading if you stuck this far.
And yes I was tempted. I also don't have any trading credentials. I'm simply a senior programmer with an eye for details who always wanted to try his hand at algorithmic trading. So yeah, I got hooked. :(
IBKR is kind of an elitistic VIP club, not just anyone can gain access. Thus you won't find a lot of libraries. There's a good number of them but the quality is not great.
I always wanted to learn OCaml by the way but in my current life and career phase I still can't justify the time and energy expenditure, and I know it will be significant.
Plus, some of it is unavoidable if you don't have a single unified exchange. Where there's latency, there's inefficiencies, and where there's inefficiency, there's profit to be made. And competition among exchanges is healthy for the ecosystem, so I don't think we'd want to consolidate.
And lastly, HFT has consolidated so much that I don't think it's worth worrying about. Virtu literally switched sides and make most of their money on order execution. Industry-wide, HFT revenues are down like 80% over the last 5-8 years. Between wholesaling/PFOF taking non-toxic order flow off the lit exchanges, and banks finally wising up on not being pants-on-head about their order execution, it's literally just sharks in the pool now, there's not even any water.
Shitty situation either way.
Shares are not held at an exchange.
Besides, when has an exchange gone done for hours, let alone days? Absent intentional breaks in trading as speedbumps.
Why would you assume that.
Every exchange is a valid place to trade and the sip ensures you always have the correct NBBO
Also, "exchange competition" is a thing in US stocks, but not, for example, in the futures' market. And there's HFTs in futures too, so having a single exchange wouldn't "remove" HFTs (not that you'd want to).
That's not true according to their filings. In their filing for 2022 Q4, I see 185m from market making, and 89m from execution.
It could be all run once a day if it was just to give ability to raise funds via investment
But really, the key thing is that liquidity provision actually is a service -- market makers are basically selling insurance. There's money to be made doing his, so people will compete to do so. That's what the millisecond race is about. Faster trading -> less risk for the MM -> less capital needed -> lower profits acceptable. Contra what you're claiming, the race is cutting into profits, not raising them.
If you want to reduce market-maker profits, crack down on payment for order flow, and let everyone compete for a chance to trade against it.
instead of order entry + continuous matching: order entry is allowed, but no matching (like pre-open phase)
then after a random period of time the auction algo runs on the entered orders
then repeat the entire thing every second