This isn’t ML. It’s cargo-cult performance of words and ideas that ML people use.
This isn’t ML. It’s cargo-cult performance of words and ideas that ML people use.
See Hyndman's fpp2 — https://otexts.org/fpp2/accuracy.html
Also, his description of rolling window validation: https://robjhyndman.com/hyndsight/rolling-forecasts/
What I've seen from companies marketing to Higher Education is - we have a lot of data, you set arbitrary flags to the data that you believe indicate 'x' (or even better, they have pre-built data expectations) and you will get 'y' outcome.
And none of it is actually based on anything real. It's all anecdotal applied to extreme amounts of actual data.
And when I read about ML on here, it seems to confirm my experiences.
1) work a LOT better for specific problems than statistics or statistical learning ever has (and at this point, I think we can safely say: ever will)
2) a lot of methods either can't be explained, or outright shouldn't work, according to statistical theory.
The use of statistics in machine learning is limited to evaluating performance and individual element performance (and even that is tenuous at best in many cases). If you ask, say, why would an autoencoder, with an LSTM on it's compressed representation and Q-learning evaluation have somewhat decent performance on half the computer games humans ever designed ? Statistics will not be useful in formulating an answer.
If you ask extremely valid questions, like "why would an LSTM predict anything ?". Statistics draws a blank. There is no good reason to assume an LSTM will ever converge (and on a truly random dataset, it won't, whereas statistical methods will still allow you to say something).
I think there's 2 reasons for this
1) the "upper limit" of complexity a human can understand in a statistical model is lower than the upper limit a neural network can "understand". In statistics the human understanding is critical to getting to a valid model, in machine learning ... it is not. Meaning machine learning can learn relationships a human mind cannot.
2) There must be some fundamental property of the world we live in that matches neural network architecture. In order for backprop to work on real-world problems, it has to be the case that almost all real world phenomena are continuous, both "raw" and in the frequency domain. If this wasn't the case, machine learning would never be able to learn anything.
When people invented a steam powered engine, some other people had probably said: "Modern physics can't explain how it works. It shouldn't work. It's too complicated". Then a few decades later physicists discover laws of thermodynamics.
You have to be very selective about what you consider "ML" to come to that conclusion. There has been a constant parade of incredible, mind-blowing results out of ML over the past decade, advancing the state of the art by leaps and bounds both in research and in real applications.
Do you not remember how terrible speech recognition and speech synthesis were just a few short years ago? Did you not see DeepMind finally crack Go? Check out BigGAN [1]. Try search in Google Photos. See translation getting better every year.
Yes, there are quacks and charlatans and people who are just plain wrong. But ML is real, and it solves real problems that people failed to solve any other way despite decades of concerted effort.
[1] https://medium.com/syncedreview/biggan-a-new-state-of-the-ar...
That's because they have the platforms and applications that people are using at scale. So ML is a force multiplier if you already have a consistent and strong user base for a good product.
If you're trying to get a product or company started, unless you're a pure ML research company like Clarifai (and arguably even then), ML is probably going to cost you more than you gain.
You don't have to invent or implement the algorithm yourself. You can use AI/ML intelligent/cognitive services provided by the big-3 cloud companies to reduce your starting cost significantly.
To have positive ROI, your problem complexity should have crossed the threshold where common-sense traditional solutions don't work any more.
If you do embark on ML research yourself, then be sure to walk the path from simpler models to complex ones while carefully establishing performance metrics.
50$ increase from 50$ is 100% increase 50$ decrease from 100$ is 50% decrease
e.g., if the model finds 50$ increase/decreases, that actually corresponds to very different wealth changes
- edit: miscounted number of LSTMs