Machine learning predictions for QM eMini Crude Oil
dutchess.ai
dutchess.ai
I use historical open, high, change, last, settle, prev. day open interest, plus several other fields I use to help recognize patterns and properly weight time-series data.
"..It contains information that is confidential and privileged. If you have received this document in error, please notify the sender and delete this file."
I use a 70/30 split of training/eval
Aside from pure P&L, you should be looking at how much risk your system is taking, and under what conditions it's doing badly. All backtests are overfit: their use is mostly in identifying problems with your strategy, rather than predicting how much money you'll make.
One question you'd get asked if you were proposing this in a real trading environment is this: what is it about the QM emini contract that makes this work? Does it work for other energy contracts? For other commodities? For bonds, or equities? If not, why not?
It took some doing to get this model to perform well. I did this by adding features that help recognize patterns in the time series data.
The features I created are not specific to QM as they are technical (eg. numbers, not news), and time-series related. So the models should work with any historical dataset with the same fields.
My goal is to add another future at some point.
I feel like you're talking past me a little. The first thing you need to do is generate all the positions your system would have taken over as many years as possible, and figure out at what times you make and lose money. Otherwise you don't have a backtest.
Right now I have residual data from the AWS machine learning data that tells me weather there is any structure to the times it does guess wrong. And a value below baseline is a better than 50/50 guess according to what I have learned about how AWS does its ML. Knowing that I use this personally as a supporting indicator to my trade decisions. Since its so new and I really don't want people to think I'm scamming or something. I'm just releasing my results free for now, not trying to be a douche ;)
AWS defines the baseline as follows
Baseline RMSE Amazon ML provides a baseline metric for regression models. It is the RMSE for a hypothetical regression model that would always predict the mean of the target as the answer. For example, if you were predicting the age of a house buyer and the mean age for the observations in your training data was 35, the baseline model would always predict the answer as 35. You would compare your ML model against this baseline to validate if your ML model is better than a ML model that predicts this constant answer.
Still, this seems like a model that will likely work about 95% then fail on outliers ie "picking pennies in front of a bulldozer" model.
Maybe because the capacity of day hold futures (especially e-micro crude) is so small that there is almost no way that the revenues earned from your 365 a year subscribers is going to be greater than the decreased capacity of your strategy from having all those people trading it.
As someone who works at a quant fund, this kind of shit pisses me off to no end. It makes a legit industry look like Herbalife
Also the point of this project is to get machine models that evaluate below baseline by improving the analysis of time-series data. I spent a lot of time improving the quality until they became better than baseline. Like I said I do not have a crystal ball or strategy for sale. Just insights from pattern recognition of historical time-series data that has helped me trade so I'm putting it out there.
It is _extremely_ likely that your model's guesses are just as bad as coin flips.
The best thing about this is that it looks like OP prototyped and released v1 for sale in under a month. That's respectable.
Thanks for the one compliment :)
Given what I see, it's really hard (for me) to understand what's going on and what's novel. I'm not very active in the financial or machine learning area. All I can understand is that someone wrote an investment program that makes money.
Even with my limited knowledge of machine learning, I know that it's very easy to confuse luck and success. How do you know that you're not just lucky? How do you know that your computer program is really investing, and not just "good timing"?
My models evaluations are performing better than baseline when trained with 70% of the data and evaluated against the remaining 30% so I take that as value. As someone else put it a potentially "favorable guess". At this point I'm using the predictions regularly. And I guess I'll know more the longer I keep track of daily results.
Edit: looking through your trade history, it does appear to have a long bias. It's predicted a rise in price on all but two days.
1. Why QM?
2. Was this tried on any other future and if yes what was the outcome?