Data trumps everything
timpark.io
timpark.io
The shift that we are seeing in lean startups reflects this - it is not a guru based process but a data driven one to start a company. Obviously, execution counts as well, but given two otherwise equal companies: One with a guru based approach (ala Steve Jobs) and one with a data driven process, I believe the inflection point has been reached where the data driven company wins consistently.
That the world would be such a better place, if we just could ban intuition and let everything get run by machines?
Even if that would be hypothetically be possible, I would neither want to work nor life in a world like that. Predictability kills any excitement.
For a fact to exist without massive supporting evidence via test data is outside the realm of science which doesn't make it much of a fact in my opinion.
which perhaps was roughly similar to what you were saying if we rephrase "the philosophy that data trumps everything" as "the ideas of the scientific method".
Deutsch argues that human progress is essentially due to the power of "good explanations", and the circumstances that allow good explanations to be generated, improved, error-corrected, and discarded in favour of better explanations.
Deutsch is critical of the idea that we should merely focus on predicting things rather than creating explanations for them.
There has been a lot of recent progress and excitement about the application of statistical/predictive ideas to business and industry - machine learning, "big data". However there there has been little (or no) progress in understanding how to generate explanations.
Deutsch argues that human creativity is a key component of the scientific method, that is often downplayed. We currently have no good explanation for how human creativity works, because if we did, he argues, we could go and program it tomorrow, and then we would have genuine AI.
(n.b. this is all hugely paraphrased and possibly misrepresents Deutsch's arguments. But another of Deutsch's arguments is that creativity on behalf of the human recipient allows ideas to survive transmission through awfully lossy mediums, such as this one)
Sometimes you are just better of with a solid theoretical abstraction, or a good intuition. If you are smart, you will understand that they make great priors that can get better with data. If you take the dogmatic route and only trust data, you will constantly feel like someone else had a head start on you.
Sounds like over-fit to me. I much prefer interpretable and plausible models to magically accurate predictions.
I am a programmer who works with a bunch of statisticians, doing "big data" stuff. What I observed is that most of them don't really spend very much time doing statistics. They spend all their time finding, collecting and massaging data. That generally involves a lot of programming. Once you get the data, the conclusions are fairly obvious without any statistics. Just make some plots and there are glaring orders of magnitude deficiencies.
I also wanted to learn more statistics... but basically with software, you are overflowing with data. The challenge in science usually is to gather data. In software the challenge is the oppoosite -- you have so much data and you need to make sense of it. To be concrete I'm talking about stuff like logs from web servers and various other systems.
So I learned a lot about sampling algorithms to cut down data, as well as various streaming algorithms. But I haven't actually learned that much about statistics. So I wonder if I am missing something.
I actually looked online for some references on streaming algorithms ... but somewhat surprisingly, I couldn't find anything really. I realized a lot of this knowledge has been hard-earned, I guess that is good :) But there really should be a reference.
Definitely look up "reservoir sampling", which gives you a fixed size sample of an infinite length stream. This algorithm has come up over and over for me, and I've implemented it in multiple contexts. There is a way to do it in MapReduce which is very useful.
You know probably the most condensed version I can think of is quite hidden: see the open source Sawzall code:
http://code.google.com/p/szl/source/browse/#svn%2Ftrunk%2Fsr...
All those functions aggregate some property of a stream. Someone (maybe me if I finish the 30 projects I've started...) really should write up some real documentation about all those algorithms, because they're not only fun, but useful.
A lot of them are (necessarily) approximations. You don't learn that many approximate algorithms in a traditional CS class. None of this will be in any stats class for sure.
Oh no, there's no risk to these mortgage backed securities at all! They're priced to be AAA by everyone on Wall Street!
September 2007:
Oh shit...