Show HN: I discovered a trading algorithm that returns ~24.85% annually
github.com
github.com
Here's the calculation used in main.js line 77 applied to a very extreme unrealistic example. I simulated 253 days of return percentages from a uniform distribution between -5.5% and 5.6%, and then the actual total return percent, calculated in R
set.seed(2020)
n <- 253
daily_gain <- runif(n, -.055, .056)
total_gain <- sum(daily_gain)
avg <- total_gain/n
annualizedReturn <- (1 + avg)^n -1
annualizedReturn
# [1] 0.2933685
prod(1 + daily_gain) - 1
# [1] 0.1324846
Edit:In reality the actual numbers are likely to be not nearly as different as this example. I chose uniformly distributed returns with a wide range to make the reason against this calculation very obvious. Here's an example return distribution where there's hardly any difference. Normal returns with average of 0.085% and standard deviation of .05 i.e. daily_gain <- rnorm(n, .085/100, .05/100) gives
annualizedReturn
# 1] 0.2414539
prod(1 + daily_gain) - 1
# [1] 0.2414051
For good measure here's one in the middle where your returns are normally distributed with an average of 0.35% and a sd of .2%, but then you have on average 10 bad days a year where returns are 5 percentage points lower than that distribution i.e. daily_gain <- rnorm(n, .35/100, .2/100) - rbinom(n, 1, 10/n)*.05 gives annualizedReturn
# [1] 0.2712024
prod(1 + daily_gain) - 1
# [1] 0.2490317The uniform distribution is a pedagogical choice, to explain why OP's average return calculation is misleading.
Backtesting is more likely to be meaningful. Am I missing something here?
It helps if you measure in the right units[0], namely bits, orders of magnitude, or fractions thereof. Up 50% is log(1.5) = +0.58 bits, down 50% is log(0.5) = -1 bits, and indeed 0.58-1 = -0.42, or 1.5*0.5=0.75, down 25%.
0: Well, strictly speaking the problem is that up/down X% isn't even in units at all.
Can you try it with the Laplace distribution? It's a bell curve like the normal distribution, but has fat tails. Extreme events aren't common, but much more common than with a normal distribution.
First of all, if you're shorting US equities and making 25% annually, that would be awesome. Heck, even being flat would be great because a strategy that is long SP500 could also short your equities and be delta-neutral and likely have a much lower volatility for the same return.
Second, so many people are mentioning commissions, trading fees, taxes and so on. Commissions and trading fees are much less than 1 basis point per trade if you use reputable brokerages. That would, at most, amount to a 1-2% in fees per year. Market impact matters but opening and closing auctions are very liquid and represent respectively more than 1% and 5% of the daily volume, probably even more for these kind of ETFs. Shorting fees are also quite small, in the range of 0-2% for liquid ETFs. If you don't hold positions overnight which is your case, you also don't pay to short!
Finally, here's what really matters. Returns by themselves don't matter. If you want a very high return strategy, you can short a long VIX ETF like VXX but every once in a while, you will be down more than a 100% ; it will bankrupt you if your available capital is less than the value of your short. You also need to look at your Sharpe ratio and maximum drawdown. Anyone somewhat experienced could tell you if the strategy is valid by having a look at plot of returns. If it's not too volatile, it could be a good strat.
Edit: addressing shorting fees
It's true there are risks associated with trading. But keep track of the money you have at risk and there is no reason not to have a go.
I know more than enough people who have made consistent returns doing algorithmic trading in under-serviced market segments.
Learning about trading on youtube is different from learning about, say, Python on youtube. If a video on programming is incorrect, your program doesn't work. If a video on equity trading is incorrect, the author of the video can take all your money.
Video views, upvotes and subscribers can all be purchased. If you have profitable trading strategy that reaps newbies who implement a bad algorithm that you publish in a public video, then you have created a perpetual motion machine.
Here's a search for "how to win slots" on youtube: https://www.youtube.com/results?search_query=how+to+win+slot... Thousands of results, millions of views each. Every video is either fake or wrong, by definition. These videos make money for the authors, and the casinos, by taking it from the marks dumb enough to watch and believe them.
- Trading Evolved, Andreas F. Clenow
- Systematic Trading, Robert Carver
- Developing & Backtesting Systematic Trading Strategies, Brian Peterson
- Algorithmic Trading, Ernest P. Chan
- Algorithmic Trading and DMA, Barry Johnson
- Trading Systems, Emilio Tomasini & Urban Jaekle
- Evidence-Based Technical Analysis, David R. Aronson
- Machine Learning for Algorithmic Trading, Stefan Jansen
You can get good at eeking out those advantages and exploiting them, but you're talking about making it your job to find financial "security holes" where the reward is printing cash. And there's a lot of really smart people spending a lot of time doing that. And you have to consistently find new holes as each one gets closed by market participants as you reveal your hand. And then you have to ask yourself if that's how you really want to contribute your time to the world.
If you don't think it's your calling, best to find products/companies you really believe in and make calculated risk-taking investments. As Carnegie would say "put all your eggs in one basket, and then watch that basket"
http://lucylabs.gatech.edu/ml4t/
The former instructor / creator & author of one of the books eventually joined JP Morgan as a ML research director (cant recall exact title).
* This is a simple strategy, which is fine, but also means you are not the only person who has noticed this. Why do you think this makes money? Is there some risk you are being compensated for, or is there some forced trading you're picking up the other side of, or something else?
* Which of the common equity factors (https://mba.tuck.dartmouth.edu/pages/faculty/ken.french/data...) is your strategy exposed to, and by how much?
* What are the basic return statistics of the strategy, like Sharpe and drawdown?
* How sensitive is the strategy to parameter variations? What if you sell the second best ETF instead of the best? What if you sell on day n+2 instead of n+1? What if you buy the worst ETF?
And about a dozen other things that you should look into before you actually try trading.
Some other metrics to measure your performance are drawdown, best month/worst month (to see if a small number of events account for the majority of returns), and Sharpe ratio. As other commenters said, try backtesting with fees/slippage. Even if there aren't fees now, you should include fees at points in history when there were higher fees. HFT's have been forced to tighten their spreads as retail traders have become more liquid with lower/nonexistent fees, so that will affect any mean reversion strategy being tested across fee change periods.
I do like the idea of spot mean reverting on large indices. Takes a lot of risk out of it (while a company can tank overnight, any decently weighted index will lack that volatility).
## The Algorithm
This algorithm is really simple.
1. On day `n`, determine which ETF gave the highest return
2. On day `n+1`, short sell the previous day's highest performing ETF at market open and close your short position at market close.
Because this algorithm operates on Vanguard's 11 Sector ETFs, it is resilient against the volatility of individual stocks.
### Caution
Hindsight is 20/20 and because this is a backdating algorithm, similar results are not guaranteed in the future. Use at your own risk.
.067 sharpe ratio, .11 sortino, largest drawdown was ~41%
Not very good numbers. Fun stuff though!
Larry Connors, Buy the Fear Sell the Greed
What was the volatility of this strategy? When backtested on the historical data, what was the maximum drawdown? What happens when trading costs or slippage (buying at the ask, selling at the bid) are modeled additionally?
252 trading days times two trades per day (short sell at open, buy to close at close) is a lot of trades, and execution quality will be very important.
Does this strategy hold up with week-long holding times?
And I'm not worried about "getting blown up" by a rise in volatility (although since this is fundamentally a short position it would be blown up by a theoretical 100% rise in a sector ETF), but I'm more interested in the risk-adjusted return or Sharpe ratio.
Since individual sectors are less diverse than the market as a whole, and since this strategy invests (shorts) one sector at a time, I would expect it to see greater day-to-day variability even before the reversion to mean comes into play. I'm curious about how much greater.
The way I discovered this algorithm was initially I wanted to buy the previous day's best performing sector ETF, with a hypothesis that the momentum would continue. But I learned that it actually ended up losing money.
So I decided to inverse the algorithm. There are still a few optimizations I can test out, e.g. Buying the previous day's worst performing sector.
I'm sure that in the ~60 seconds I wrote this post, some brilliant future software engineer just reinvented binary search.
It is also clear from other posts by the author that this has not been tested in the real world, only back tested.
Practitioners largely consider back testing to be somewhat irrelevant as it’s hard to achieve real world results that compare favorably to the back test. It is essentially over fitting the curve.
There is a good reason successful hedge funds like RenTech are not hiring finance people but mathematicians who have no idea about "mean reversion" and other pseudo-scientific terms.
It’s definitely not the case that they want people to reinvent the wheel. Options pricing for example is a Nobel prize winning discovery.
–Richard P. Feynman
Wise words for anyone wishing to try out an algorithm on Wall Street.
There is so much wrong with this I don't know where to start. But this seems common these days, I think it's due to the influx of inexperienced traders who have no proper statistical background.
I'll let ryanmonroe point out the first glaring problem with these types of simple "algorithms":
I found a bug[1] in a popular (4000 star) stock forecasting model on github where future knowledge subtly leaks into the training data. People seem to keep using the project though!
[1] https://github.com/huseinzol05/Stock-Prediction-Models/issue...
I'm not familiar with Yahoo Finance so I'm not sure how feasible this is.
Here's one I found just with a google search: https://alpaca.markets/
Erm - edit - scratch that: But what if one of the ETFs goes to the moon? Then the algorithm will continously sell that all the way up.
The algorithm of just buying the worst performing etf of yesterday.
Of course, this might be a rationalization of me not wanting to spend the time and effort to construct an effective trading strategy.
i’m curious what the backdated return would be for different time periods
You can try sampling different time periods by downloading data from Yahoo Finance. I may need to make the code more robust because there is some data that I hard-coded just for the dates I selected.
for (var i = 0; i < performance.length; i++) {
total += performance[i];
}
var avg = total / performance.length;
var tradingDays = performance.length;
var annualizedReturn = (1 + avg) \* 253 - 1;
}I think this math is wrong. You cant just add the daily performances up to get the total return. unless im missing something.
Though seems like your "cash" variable is correctly calculated.
If so, it would be clear that buying anything at the opening price is not easy.
Check it out: https://www.golfforecast.co.uk/profitgraphs
And that’s the problem with any successful algorithm except buy and hold value investing. The latter being hard because it requires doing little, and nobody believes that which requires the least effort to be the best.
This needs to be re-run with the borrowing costs added.
The fact that EFT short selling costs have gone up indicates that others are active in this space, which means the profit opportunities for a simple algorithm have probably been already taken.
[1] https://www.reuters.com/article/us-usa-stocks-shorts-etfs-id...
You don't say...
You massaged an algorithm until it produced a 24.85% annual return training on historical data.
Come back when you are ready to claim you have made ~24.85% per year with an algorithm you created 5-10 years ago.
Deny it, Downvote it: Destiny still arrives.
I fully understand the idea of historical algorithms being no true indicator of the future. With that said.......
If an algorithm consistently performs over 20+ years of data (through multiple black swan events, multiple major events), then why is it not safe to assume it likely will continue going forward? Wouldn't 20 years of "evidence" be a huge amount, such that future events likely wouldn't deviate much from that...?
My ML knowledge is somewhat rusty... does overfitting occur more often on models with many input parameters (ie.. neural networks).
His algorithm seems very simple, without really using ML at all, it's more of just a procedural 1-2-3 step thing, with no actual learning.
Can you explain how overfitting works into his algorithm?
I specifically asked how overfitting applies to a simple procedural technique, rather than a multi-dimensional method like a neural network.
Trying myself, and losing money, doesn't explain how overfitting applies to procedural steps (as I said, my ML is rusty)
For momentum, Jegadeesh-Titman paper is much of the foundation for these types of strategies, if you're interested. But as others have pointed out, the "smart money" saturated this trade decades ago.
> Implementing this on any kind of scale would be expensive to trade since it rebalance's daily.
On {day/week/month} n, determine which {n} product(s) of {product brand}'s {product type} of {underlying asset type} gave the {highest/lowest/some of each} return.
At n + {a number} {days/weeks/months} go {short/long/some of each} at {market price/limit price} at {market open/market close/time in day} the previously identified securities. Close your position on day n + {a number} afterward at {market price/limit price} at {market open/market close/time of day}.
This particular attempt seemed to have succeeded. But how many were tried that didn't? If you torture the data long enough it will confess.
People invent many algorithms all the time. You only hear about the ones which turned out exceptionally well.
In the extreme: If you take completely random price fluctuations and plot them on a graph then review them, they are no longer “random” - they become fixed because they’ve been recorded. It becomes possible to look for patterns, of which there will always be some (that’s the nature of randomness). It doesn’t matter the time scale. So your system that works over 20, 30, 50 years of past data is just fitting to the randomly-generated but not-actually-random once recorded set of data. From now into the future, the randomness will re-emerge and the system will fail.
Another way of thinking about it: imagine you click a button to generate a random time series over 20 years, then build a trading algorithm based on that single click. Great - it works! Now click again and regenerate the entire series and see if it still works :)
Obviously, sitting down and creating that "strategy" would be silly, nobody would think that would work. But if you're using ML or just test a million random strategies, you could end up with something along those lines, a strategy that is optimised for the particular paths stock prices took up to know.
The best strategy doesn't seem to me to be straightforward to calculate, even with perfect hindsight.
It might be to continually switch in the short-term most profitable asset, after taking into account the transaction costs and risks of influencing the market.
especially since market paradigms shift.
ie the last 20 years have been low interest rates, low inflation, with strong secular growth in tech and stagnation elsewhere, and globalization
this varies a lot from the 60s or 70s where there was massive inflation, high in interest, etc..
we may be entering a new paradigm with the changes in fiscal, monetary policy and globalism running its course
If anyone is interested in trading algorithms there are entire communities that filter out random backtests that havent been run live.
Check out QuantConnect