Show HN: Bateman, a stock trading system I'm working on
github.com
github.com
This is a huge red flag with respect to the simulation results. You show some trades like this one
> 2013-03-05 6:41,2013-03-05 7:15,14.47,14.48,LONG,5192,75145.75,52.85,100237.47
where you enter the trade and exit a penny higher. It sounds like you're just looking at the trade print and assuming you can execute at that price with a market order (or marketable limit order). Consider a stock at 14.47 bid x 14.48 ask. If I cross the spread to sell at 14.47 and then someone else crosses the spread to buy at 14.48, you will see two trades at the two different prices, without the prices on the inside having changed, this is why the midpoint between bid and ask is considered a more useful value than the last trade price.
With the system you propose, you are a price taker. You are crossing the spread with both your entering and exiting trades. Most of the trades you show are for around 5000 shares. Assuming the spread is $0.01, you are going to spend $100 just to get in and out of the position.
I don't know what kind of data Google offers about the intraday state of the order book, but I think you'll need to incorporate it into your backtesting in order to get a better picture about the profitability of your strategy.
I understand this is an intellectual exercise, but for those considering going into algorithmic trading, those words are dangerous:
- what transaction fee model is being used? Almost all profitable day trading strategies trade too often that the profits and adverse selection reserve are decimated by commissions and taxes.
- have you tried to approximate the presence of your own trade? For example, if you sell a boatload of google shares, the price will start falling. Even with more liquid issues like AA it doesn't take much (5K shares) to rock the boat.
- Have you considered the spread? It's unprofitable to quote a penny spread on google or other high-dollar names (the SEC tax alone, roughly $25 per $1M sold, doesn't allow for really profitable market making without at least 3 cent spreads.
There are many more questions, and for each question there are hedge funds and prop shops that have lost significant amounts of money, or were driven out of business, due to an oversight.
However, from experience developing and testing algorithmic trading systems I can tell you that your strategy probably has some issues in its current form. I haven't looked into the code but from your description it appears you (correct me if I'm wrong): 1.) Pick a stock 2.) Use PSO to figure out the parameters 3.) If profitable, run the strategy on the stock with the optimised parameters
This means you're making a well known error in the system development community which is curve fitting parameters to historical data. This'll look very good in the simulations, but there is a high probablity that it will break down when trading it forward with real money, because it is optimised for the past. This is why there are a couple of widely accepted best practices when it comes to developing and testing trading systems.
First of all, your system should not have or need too many parameters. As a rule of thumb a robust system shouldn't have more than a handful of parameters and it should ideally show profits in simulations without a great deal of optimisation on those. When optimising make sure that the optimised parameter values are robust. This means that changing the value by a small increment only changes the resulting performance of your system by a small margin (somewhat analagous to numerical stability). If the performance changes by a big margin, then those values aren't robust and should be discarded. Furthermore, don't run optimisation on all of your historical data. Instead, optimise on portion of that data (the 'in-sample' data) and then test the optimised parameter values on the more recent data your didn't optimise on (the 'out-of-sample' data) and see if the performance of your system stays the same or breaks down. Another popular approach is 'Walk forward optimisation' [1] which takes the above one step further by repeatedly optimising and forward-testing on your historical data to find robust parameter values.
Some other things to consider: You need to factor in transaction costs, spread and slippage (the difference between the price you enter the order at and the price at which you get the fill). Transaction costs are easy to determine. Spread and slippage only apply when using market orders and can be reduced by trading with limit order if your system isn't negatively affected by this. Trading with market orders in a fast-moving market may incur siginificant slippage and there are predatory HF algos out there making money from screwing you on your execution. To get a better sense of this, it is considered a best practice to run your simulations on a lower timeframe than the one your system is supposed to work on in order to eliminate inaccuracies in the results. Ideally, you run simulations against unfiltered tick-by-tick data and additionally used bid and ask data series to factor in the spread. This may, however, be overkill and not needed for a system that runs on a daily timeframe, but it may make all the difference for a faster system.
Particularly running the optimise & test routine on chunks of past data to test for robustness & over-fitting.
The the thing about backtesting a strategy is that it is very easy to make a mistake in your backtester. Look ahead bias is the most common mistake.
Another challenge is the data. Are you testing against a history of stocks that includes bankruptcies? If not you have survivorship bias.
I suggest you take a look at my website, www.quantopian.com. Look at our open-sourced backtester, https://github.com/quantopian/zipline. Between the two we can help you get past those two sources of error.
I wrote my own fairly pessimistic backtester. Also, I currently use the generated models for buy signals, but I tend to sell earlier than they dictate, because I find it hard to turn down even a modest profit.
Assuming you did put the money in in January you would have probably made more profit by simply investing in an index tracker.
My gains are realized. And fairly predictable.
Also, I only have about $55k in capital. But my rapid buy/sell cycle (not HFT) allows me to keep 'reusing' it.
- trading commissions & fees - capital gains taxes - currency exchange (for those of us "unamericans")
then all of a sudden blammo, I decided I'm better off putting "investment" cash into my mortgage
I don't know what country you're in, but if you want to invest in the US stock market, I'm sure your broker sells an index fund with very low management fees. That may be a better investment than putting everything into a single home in the current market.
(In the US, mortgage interest is tax-deductible, making this strategy even less of a good idea.)
All jokes aside, this sounds promising. Will your framework be flexible enough to allow for composing and running any ML/statistical algorithm?
It's easy to come up with an algorithm, but where the real pro's win is with risk management.
EDIT: This isn't meant to be negative at all. Keep up the good work, but know that you're just starting down the rabbit hole now.
I think we just generally label the higher latency, lower frequency strategies old fashioned 'stat-arb' -- this better connotes that the 'edge' from these trades is derived out of superior mathematical modeling and not better execution.
I highly recommend decoupling the simulation/back-testing framework from your strategy work.
Building a reasonable simulator is no small task -- market data is often NOT an accurate depiction of what is executable and there's a fair number of corrections/assumptions that must be made to reflect this.
Either way, I've always liked PSO -- I highly recommend playing with DE (differential evolution) as well. I've used DE to tune parameters for many strategies with great success where SGD would have been, well, painful.
Their point is moot. DC comics will have to go after people who own 'Patrick Bateman'. This isn't infringement(nobody is going to confuse a trading software named Bateman with a book/movie character); even if it were, the owners of 'Patrick Bateman' and not DC comics can claim infringement. At most, he will have to remove the image of Christian Bale as Patrick Bateman, but I doubt it will come to it.
Also, batman.js is in business, and DC comics didn't go after them. At what grounds did they went after your friend? Naming your software batman or superman isn't infringing DC's trademarks.
How do you define and track the "online" world though?
Wink wink, nudge nudge, that's the secret sauce. ;-)
But more seriously, the plan of attack right now is just to scrape newsfeeds from news sources via rss.
While of course this is susceptible to bias in the media, my hypothesis is that so are the stocks. A rising tide lifts all ships as has been said: so long as you buy lowish and sell highish you'll do all right[1]
Other routes might be:
- stock trading forums
- industry news sites
[1] bear markets demand some kind of shorting strategy I believe.
http://www.idsnews.com/news/story.aspx?id=80469 http://venturebeat.com/2012/05/28/twitter-fueled-hedge-fund-...
http://www.theatlantic.com/technology/archive/2011/03/does-a...
While far from a silver bullet these can help you avoid buying positions when the market is hopeless, or trying to play a statistically losing game. It might be worth testing.
Make sure you include all trading fees and software license costs in your models.
I would suggest trying out-
1. market neutral positions. Eg- If you go long AAPL, go short XLK ( The tech stock ETF). 2. going long slightly out of the money call options instead of stock. ( If the options are liquid)
I mostly do stat arb, and for back-testing, even I tried Metatrader, Quantopian and several other platforms and didn't think any of them were suitable. FXCM's Strategy Trader is worth taking a look at. It can only trade forex live through FXCM, but you can import CSV data and back test on whatever you import.
>buy stocks that are going up intraday and sell them higer
to be blunt it seems interesting, and I would encourage you to work on it further because complex systems rarely work right. but also bear in mind that its not any different than a candle stick trading strategy, or ichimoku clouds. The key is how fast you converge to usable parameters (time,money,trades), if you can do this fast the underlying strategy can be a multitude of things, and thats my 2 cents...
Anyone not using limit orders in an automated trading system deserves whatever they get.
Sometimes a trading system just needs to get an order done. Having too many busted pairs can really cause your system to become ineffective fast.
or put another way, limit orders are your best bet, unless they aren't:)