Predicting Stock Performance with Natural Language Deep Learning
microsoft.com
microsoft.com
But in that case, it’s very unlikely a heavy CNN text model would be better suited than simpler methods like a classifier using LSA vectors, or even just extracting the gloVe vectors and doing no convolution. Especially given all the metaparameters they mention needing to tweak.
For example, an ablation study to figure out that you need Lecun initialization seems both overkill and requiring way more specialization in deep learning than a typical firm would be looking to have, all just to squeeze out what might be some slight efficacy in one industry group.
When I previously worked in quant finance, I used to be very passionate about trying to apply the latest & greatest methods.
But over time my feeling is most of it is inapplicable to finance, because it ends up being a lot of work for pretty much no efficacy, like in this post. I found much more value in simpler regression and tree models, and using simple bag of words models with text. The only “advanced” stuff that ever seemed to add significant efficacy was Bayesian hierarchical regression, but only for helping overcome limitations of classical random effects models, not for adding any greater complexity.
The post author deserves plenty of credit for skills in deep learning, but the analysis seems unconvincing if it’s supposed to market for using MSFT cloud for deep learning in finance.
1. They applied pandas.qcut over the entire returns (i.e. training and validation set) when generating the target performance labels, which could compromise the validation.
2. The ordering of news is actually important when it comes to forecasting the movements of the market; their training/validation split was done after shuffling the entire dataset, which means the model will have access to information that shouldn't be available to it during training.
3. It makes more sense to use a training-validation-testing setting rather than a training-validation split when reporting the model performance in order to avoid inflated results due to all the hyper-parameter tuning.
This can lead to artificially increased test set accuracy since by optimizing the model during training, you are partially optimizing based on information you gained by knowing about the test set in advance.
Training-validation-testing generally means keeping a portion of the training set held out for evaluating accuracy and convergence after one full iteration of training updates. So you get some information per each training update regarding whether the model had overfitting, or performs unrealistically well on the validation set (which it’s not trained on), or for early stopping criteria.
Then, after all that, you move on to checking the final accuracy and metrics on a fully held out test set.
This validation approach is common in deep learning because of the many diagnostics you need to get information about during training. Without some information outside the training set itself, it can be hard tonunderstand how the learning rate is affecting you, how likely overfitting is, whether there is a vanishing gradient problem.
Waiting until all the way to the end of training to get any feedback on these is sometimes just too inefficient or too risky that you’ll discover an issue only after a huge time sink.
1. How did they split the data for train/validation/test sets? If all of the sets had the data from the same time period, or the same companies end up in multiple sets, it is a major flaw. For example, if outperformance was persistent for the companies over the time period considered, the model may simply learn to identify specific companies by their filings.
2. What is the variance of the out-of-sample performance? Given that their dataset is very small, and the model performed badly at predicting high returns and reasonably well at predicting good returns, what are the chances of getting those results by luck alone?
3. How has the model performed since then, on the most recent filings?
4. Why use a convnet? Would gradient boosted trees not perform just as well/better but be more interpretable? Methods like the ones in the eli5 package can help get an idea of why the model makes a particular prediction, which could help sanity check the model.
Here is an article that explains the problem well: http://zacharydavid.com/2017/08/06/fitting-to-noise-or-nothi...
When wouldnt it be?
Take handwritten digit recognition task like MNIST, for example - the way we write the number '5' does not change much, and the distribution of the different ways people write that number stays pretty much the same over time. The labels we are given are very accurate and everything we need to classify the image is in the data.
All these assumptions mean that picking the right methodology is fairly easy - anyone can read a short tutorial and get it right.
None of the assumptions I listed hold in the world of finance. There is no standard methodology that everyone can agree should be followed. You really have to be very, very careful in order to produce useful results.
ALl of that means that when talking about applying ML to finance, discussing the methodology in detail is a must. If it is not talked about or only briefly mentioned, from my experience that usually means that the methodology used is rubbish. Whereas you don't need to focus as much on the methodology when talking about MNIST-like problems - one can usually assume a reasonable one is used.
When wouldn't it be? And why does this sound exactly like the other post I just read?
As for the second part, I don't know. Probably because we had the same thought? Not sure which post you refer to.
There are two mutually exclusive conclusions that you could draw from this: one, the informational advantage you get from reading company reports is small, and TFA is overstating their case. Two, convnets have discovered a way of reading reports that conveys a big advantage, and further, human readers can't replicate the process.
My money is on the first hypothesis. The second is an enormous claim, and so demands an enormous amount of evidence.
Unless it's a Moneyball-like situation, where managers theoretically have access to all the same metrics, but actually use inferior squishy human judgement instead of hard data.
Due to the amount of data across the entire industry, it's totally conceivable that "human readers can't replicate the process".
That said, I'm a bit of a pessimist, and concur with your suspicion.
With GBT, you can check why a particular prediction was made - roughly - by navigating down each tree for a particular sample and summing up the influence of each feature on the final score. Then you can see if something weird is having a large effect on your score. Can you do something similar with ANNs?
I want to posit a general hypothesis (perhaps it's already been said); better-than-chance performance by a classifier on some dataset is not evidence that a similar (or any) classifier can perform better on the same data.
Now climate is changing you see that all the models are becoming less and less acurate.
So yeah some models might be performing better but overall they are all performing worse.
That's not true.
"Last year, the National Weather Service 5-day forecasts were within 4 degrees of the high temperature. That's as accurate as 2005's 3-day forecasts and a full degree better than the 5-day forecasts of 11 years ago"[1]
[1] http://www.mcall.com/business/tech/mc-weather-forecasts-impr...
https://www.washingtonpost.com/news/capital-weather-gang/wp/...
Sounds really interesting.
Did all the true 'low return' predictions happen during a bear market, or is there a good temporal mix between predictions?
Did they all happen for the same company? If so, it it possible the report was revised after publication?
What horizon are they measuring performance over? What happens if the change that horizon? What happens if they lag the prediction by a day?
What happens if they use a finer granularity than high / medium / low?
How?
How to tell that it's good news: there are numbers near the top.
How to tell that it's bad news: the figures are buried as deeply as legally permissible.
I built system to determine which fundamental or technical factors have the most influence over a stock price. It required a fair amount of manual tinkering and sample runs to eliminate variance. The system itself is not that complex but it works. When I have demo'd it in a interview usually the interviewers are not that impressed... they are expecting that the system must be very complicated to have any results.
If they mention the weather during the last quarter - run!
And in the end you want to invest based on the model. Arbitrary bins aren't that helpful if some of them combine desirable and undesirable (e.g. positive and negative) results.
well thats not random or arbitrary just wrong.
>And in the end you want to invest based on the model.
I'm not quite sure what you mean here. Either you determine a model and hyperparameters (which includes the bins) is correct enough of the time via testing on out of sample data or even synthetic data or you are talking about a determining what to do given a single observation, which I would assume you give it to the model and ask it what it thinks and the bins are part of the determination (as done here in the article with softmax output) of what to do and given you've done the testing you should have a level of confidence about the outcome of acting upon the models output. The bins aren't a post processing step, the post processing you might do to trade recall for accuracy might be to expect the bin to have a stronger signal (class > 0.5 or something maybe and otherwise ignore), all of this is "the model".
>Arbitrary bins aren't that helpful if some of them combine desirable and undesirable results
Correct me if I'm wrong but it sounds like you are making the case that there are models and bins which would have only good outcomes (and those are the only useful ones)? Am I misunderstanding something here?
The arms race effect you mention absolutely exists. Corporate officers are often coached extensively on key phrases to repeat in earnings calls, and phrases to avoid, and how to steer caller questions away from SEO-like speech patterns they don’t want.
I bet someone is going to make a fortune based on the pitch alone.
Where it can help is to determine the state of the economy. Same as with image recognition for satellite pictures (traffic patterns, footfall at stores), those indicators can give you a faster idea if something's going wrong with the economy than classical indicators. But I have yet to see an accurate one.
I wouldn't train a AI on stakes- i would train it on history books/ old newspapers to find hidden needs and gaps. People do not use half of the day, due to darkness? People fighting, due to starvation, or reduced chances to mate and procreate?
A NN could find these needs, and hand it to a diffrent NN, that is trained to find solutions/ companys likely to develop solutions.
Financial forecasting using character n-gram analysis and readability scores of annual reports. by Matthew Butler and Vlado Keselj
There was a model to predict up votes on Hacker News based on title that made its rounds a way back. Fortunately we don't have a bunch of submissions with the title "YC YC YC YC YC YC golang is better than Rust" [1]
I don't really think we can predict the stock market because the market fluctuates every day based on news and things like that.
Stock market is like a legal gamble, in the words of Wrren Buffet, "The intelligent earn money off the fools".
He did not spit out " 100 stocks which will go up tomorrow using my fancy algorithm"
He researched the people behind the company the finances etx
Wildly different from what algorithms do
* You may not be eligible to trade that instrument because of where you live or who you are
* You may have found an edge based on historical data in one instrument and expect it could work for another but lack the data necessary to confirm
* You may not have the capital necessary to trade at sufficient scale
* You may not have the programming skills to turn a mathematical model into an build and maintain a high-frequency trading bot with adequate risk controls (waaaaaaaaaay different skillset)
* You may have found a prediction who's error rate is within the transaction costs and thus needs to be traded at higher volume (lower tx costs) to confirm
And before anyone thinks "why not just negotiate to sell it to someone who does" they get hundreds of random people doing that every day so your chances of getting noticed are between slim and none.
Now, what is an effective way to benefit from a market prediction model is to publish it. Those people may hire you to continue to develop it further or to refine it for other markets you didn't think of. No model lasts forever, so having demonstrated the ability to find one makes it more likely you will do so in the future, thus should attract big money salaries from people with the capital for your next one.
Not to say that these are legit, just pointing out its not hand-wavvy dismissive is all. I wish authors would start a paper with something like this just to address this inevitable comment. "We aren't trading this because we can't trade SLURM futures" would really help readers out if they also can't.
I think this is a fair and in depth comparison.
redpixie is a MS partner, so I don’t expect it to be fair.
https://www.redpixie.com/blog/iaas-paas-saas
“As a Microsoft Partner, we focus on Microsoft Azure IaaS solutions”
Microsoft: $21.2 Amazon: $20.4B
From [1]. That includes Office365 revenue[2]. Excluding that, Azure is about 1/3 the size of AWS[3]
[1] https://www.zdnet.com/article/cloud-providers-ranking-2018-h...
[2] https://techcrunch.com/2017/10/30/aws-continues-to-rule-the-...
[3] https://www.srgresearch.com/articles/cloud-market-keeps-grow...
[1]https://www.theguardian.com/books/2011/sep/30/fear-index-rob...