Twitter Can Predict The Stock Market
wired.com
wired.com
So I did some research, and most people who have written them will tell you that in cases like this, training on past data doesn't correlate well with current & future data.
The stock market of 2011 is not the market of 2008.
But what do I know, not like I've actually done it :)
You can rent your trading strategies to others, or rent someone else's strategy.
Also some good info/tools regarding automation.
If you have a good strategy, you won't rent it out, you'll trade it. Why risk others frontrunning you? Of course, you might post a historically good strategy and front run it. Or they might just be risky strategies, which look good for a short time (encouraging people to rent them), but which carry catastrophic risks the creator doesn't want to take on.
I can't see a single reason why someone would post a good strategy here.
It's hard to see how those economics would apply to collective2.com.
Isn't there a whole industry that revolves around paying people to provide you with their trading strategies? How is that considerably different than what's happening on Collective2?
Having a good strategy doesn't mean you have the money to actually trade on it. Or that you can slowly build up your trading bankroll using the same strategy.
Then there's strategies that only yield modest returns. Why not make some money on top of that by renting it out? If you let a dozen people use your strategy, does acting on that information give you much of an advantage? I would guess that depends on how much money those people are trading on your strategy.
I would also guess there aren't many big players renting strategies on Collective2. It's an interesting concept and I think the fact that they've been active since 2003, somewhat validates the idea.
The best thing about it seems to be the ease of using an automated trading agent. I don't know how easy it is to do that elsewhere, but one reason to put your strategy on Collective2 (I'm guessing you can keep it private) would be to use their automation facilities.
A hedge fund requires capital to operate, and the owners can't necessarily cover fixed costs (salaries, etc) with their own personal capital. I'd be surprised if many of the strategies on collective2 fit this model
Investment advisers often fine tune a strategy to match your personal risks - i.e., help Southwest Airlines hedge their exposure to gas prices, or Apple to hedge their exposure to the RMB. Since Southwest is already short oil due to being an airline, the trading strategy of going long evens them out. It wouldn't make sense for me to trade this strategy, since I don't have an intrinsic short position in oil (plus the alpha in Southwest's strategy comes from selling flights, not oil).
If you let a dozen people use your strategy, does acting on that information give you much of an advantage?
Buy $10k of some low volume stock. Have a few other people pile on and buy the same stock (after you). The price will go up a few cents. Then you sell, probably to the same people buying from you. This is called frontrunning. If you didn't frontrun, you bear the risk that one of your renters would buy the shares before you do, thereby driving up the price before you purchase it. Less of an issue with GOOG, admittedly.
Anyone doing serious research into this, will first partition the past data into training and test sets. (And sometimes other validation sets).
So the idea would be to fit a model on one set of past data (the 'training' set), check it works, and then, in the final evaluation, run it on the never seen before, never used, never thought about, 'test' data.
If you have a model trained on 2009, and it also does a great job the first time you run it on the Q1 2010 data that you've never looked at before, I'm now interested, even though every data point is in the past.
I imagine they had to do something like this to pass review.
Sure, the data is historical, but your model doesn't distinguish between "old" and "new" data. If your model predicts test data (in-sample forecasting, right?) well then you have something interesting.
So, you have a model trained on 2009. You try it on the Q1 2010 data and...it doesn't work. Damn. So you throw it out, go back to the drawing board, and try again. And again. And ag...hey, this one works! Trained on 2009 data, and it predicts Q1 2010 perfectly!
Do you trust this model to predict Q2 2010?
That's why I mentioned the validation sets, and that the test set must never have been looked at, or used before.
But the point stands - if the method works on a clean test set, even if the test set is in the past, then it should be taken seriously.
Would I trust such a model to predict the stock market in Q2 2010? No, because my prior belief is that the stock market is very hard to predict, so I would need very strong evidence to the contrary. But that has nothing to do with having confidence in models that have been tested on historical data.
The Dow Industrial Average over the last 10 years
http://www.google.com/finance?chddm=997050&q=INDEXDJX:.D...
* Notice that the end of 2008 was unusual for the index. 2008 had the most herded and fearful stock market in recent history. If at anytime the stock market was correlated to mood, it would be then. I am not sure if a 2008 analysis can be generalized to any year but 2008.
* They have not done an analysis on 2009 or 2010, and they chose to split the analysis and pick December 2008 based on a qualitative assumption from the "stabilization of DJIA values after considerable volatility in previous months and the absence of any unusual or significant sociocultural events". December 2008 was very much in the midst of the crisis still.
* For their December "stable" data set, they only used 30 days. That is limited in sample size. There is a big pool to draw from since 2009 as the market has been relatively stable.
Also, some food for thought: it would be interesting to see someone testing Twitter moods as an instrumental variable for a project.
But overfitting is definitely still a concern. Looking at the overall trend for the Dow Jones in 2008, I wonder what the success rate of an indicator that always said 'down' would be.
If that were true, there would be an exceptionally easy way to make money: Buy a future or option today based on yesterday's move. Leverage ad infinitum.
Random coin flips also tend to run in streaks, btw - in a few thousand throws, you'll probably have several 10 "head" streaks and several 10 "tail" streaks.
This kind of co-correlation and general market direction is already baked into the option and futures prices. Also, they don't tend to fluctuate as much from day to day, since their prices reflect what the value will be on the contract delivery date, not what the price will be tomorrow. I'm not clear on what profit opportunity you're seeing.
Random coin flips also tend to run in streaks, btw - in a few thousand throws, you'll probably have several 10 "head" streaks and several 10 "tail" streaks.
And in a flat market, that's often the behavior you see. When the market starts trending, though, the coin starts acting 'rigged', and streaks in the prevailing market direction tend to become longer.
I don't know what futures you were thinking of, but financial futures (single stock, index futures, currency futures) track the base value EXACTLY (but also taking into account interest rates, dividends, etc). If this weren't the case, there would be an immediate arbitrage opportunity.
Specifically, once you factor the interest rate out, the DJIA future and the DJIA index are in sync within seconds. The HFT traders take care of that.
And while it is true that the market does trend occasionally (more than a coin flip), timing the start and end of the trade is empirically very hard.
If you know the market is trending, why don't you buy a future betting on the trending direction, with a stop at 2 ticks above and below your entry price? If the market is trending, you have positive winning expectation.
Except the market doesn't work that way - and if you think the market is trending when it isn't, you lose money with this scheme.
Consider a set of random signals; arbitrarily select one as the benchmark. Then from among the rest take the signal that best predicts the daily direction of the benchmark. That signal will likely have much better than 50% accuracy because by definition the worst signal will be around 50% accurate (if it were any less it would have an equally useful inverse correlation).
Then: predicting up or down movement of a stock is very vague. At what time scale, what sort of trades are required and what sort of response times to execute. What are its drawdowns like, does it account for taxes, commission fees etc. Next, use of a complex nonlinear learning model with lots of parameters - raises alarm bells - these tend to be very susceptible to noise, trading data is highly correlated and typical regularization methods often do not suffice. Then there is the whole issue of over-fitting in general, data used to train on (size, survivor bias, accounting for splits and what not) which makes the whole thing very hand wavy. Without additional info as basic as rate of return, the stated 83% accuracy is meaningless. Like with all things, its easy to get results that work within the limited and safe confines of academic testing but actually shipping a working product is another story.
There has always been a draw to beating the stock market. And these days there is nothing more romantic than doing so using Artificial Intelligence! But I think the most important part of any trading strategy is to be made up of parts that are constantly being swapped out and replaced based on research. you can't just throw a machine learning algorithm at it and think job done. The thing will likely only profit for a couple microseconds. however, as an aside, I would not be surprised if one of [anti]spam/virus/botnets or HFT wars will one day produce AI.
An even better question: is the relationship causal? The researchers use Granger causality analysis to test their hypothesis. Wikipedia tells me this analysis "may produce misleading results when the true relationship involves three or more variables." [2] By definition, Twitter and the DJIA are macro aggregates of a number of factors. How could the researchers apply Granger here?
[1] See Table 1 at http://www.sca.isr.umich.edu/documents.php?c=c
[2] http://en.wikipedia.org/wiki/Granger_causality#Limitations
Really? There are some true believers out there under this impression, but I didn't think anyone credible was. It wasn't so long ago that someone showed efficient markets were an P=NP problem.
EDIT: I'm not the one who downvoted you.
The NP completeness of efficient markets has been known for quite a while.
http://dpennock.com/papers/pennock-ijcai-workshop-2001-np-ma...
It wasn't so long ago that some jerks at Princeton wrote a paper along the same lines, completely ignored all the existing literature to make their paper appear more novel, and got a lot of publicity for themselves (hint: prediction markets are unsexy, CDO's are sexy).
Furthermore, the paper you reference (while an excellent and fascinating paper) does not directly bear on the NP completeness of the stock market:
> In Section 3, I discuss the prospect of opening securities markets that pay off contingent on the discovery of solutions to particular instances of an NP-complete problems. Such NP markets would provide direct monetary incentives for developers to test and improve their algorithms, and allow funding agents to target rewards to the designers of the best algorithms for the most interesting problems. In Sections 4 and 5, I discuss markets in #P-complete problems, where prices serve as collective approximate bounds on the number of solutions, and bid-ask spreads may indicate problem difficulty
is his summary of what the paper does (sections 1 and 2 are introductory material). I claim that this does not at all show the NP completeness of markets, and further that it's a claim irrelevant to the discussion here.
In what sense are you claiming that he proves the "NP completeness of markets"? What does that mean? Why is it relevant to whether or not to invest money in the stock market?
(sidenote: I don't think the question of the NP-completeness of some questions related to stock pricing is irrelevant or uninteresting; indeed I just applied to grad school to study problems like these. I just don't think they bear on what you're implying they do)
That said, I voted you up because of your first sentence.
However, I made a mistake and linked to the wrong paper. Here is the correct one:
http://dpennock.com/papers/fortnow-dss-2004-compound-markets...
Basically, the result says that if you have a market in derivatives which pays off when certain formulas in propositional logic are true (e.g., a derivative which pays off if A && (!B || C) is true, for specific events A,B,C), then the auctioneer's matching problem is NP complete. The auctioneer's matching problem is simply market making, and if the market were efficient, this problem would already be solved (by looking at prices).
I don't think that loewenskind's claim is true, for the most part, I was just providing a more detailed source on NP completeness of some markets.
I think people citing the fact the market isn't totally efficient aren't always proving what they think they are proving. The practical difference for the vast bulk of us between a market that is totally efficient, and one that is only mostly efficient but it is very very hard to find the inefficiencies, is pretty much zero. I often see people try to leverage this into the idea that markets are efficient as somehow being "touchingly naive" to score political points in various political fights, but, well, there's a reason you have to reach for that emotional trick, because the facts don't really support the idea of some grossly inefficient market in practice. (Instead, the problem is that efficient doesn't mean what you think it does; it certainly doesn't mean "good" or "moral" or "stable" or anything like that.)