Nate Silver: What I need from statisticians
statisticsviews.com
statisticsviews.com
There's a fundamental difference between data scientist and statistician, I think. I see statistics as an academic discipline and data science as an applied discipline.
More concretely, the statistics approach is: formulate question --> formulate hypothesis --> collect data in a controlled environment under a specific set of assumptions (i.e., perform an experiment) --> determine probability of the data given the hypothesis (and assumptions).
While the data science approach is: hey look, we already have all this data --> generate predictions --> collect more data --> refine predictions.
Of course, that's an over-generalization. But I think the different emphasis on hypothesis testing vs. machine learning/data mining is fundamental.
As a statistician-and-engineer who is currently on the job market (my graduate program finishes this spring), I feel this pain.
I've been referred to as a "data scientist" multiple times (that's even been my official title at work before), though I do still cringe sometimes when I hear the word, for this exact reason.
That said, I don't usually present myself as a statistician, even though my degree is a statistics degree. Most people who hold statistics degrees are fairly lousy engineers[0], and I don't know of any other term that (concisely) expresses that I'm equally competent as a statistician and a (backend) engineer[1].
Of course, this is because many of these programs haven't yet caught up to the fact that computers exist and are still teaching statistics as if we're in a pre-computation era. The perfect solution is to fix this, and thereby fix the connotation of the word "statistician".
It's the same reason I dislike the term "growth hacker" - really, that's just the way marketing should be done (ie, based on numbers and verifiable statistics). In a perfect world, all (competent) marketers would be "growth hackers". But many marketers aren't, and so we have to make up another cringe-worthy term for it.
Unfortunately, that's a problem that's beyond my means to solve. So I bite my tongue and add the word "data scientist" to my resume anyway.
[0] Usually self-proclaimed, too [1] ie, "I could work as a backend engineer if I wanted to/needed to, but I'm looking for work involving both skillsets"
Citation: working in data science and machine learning engineer roles for nearly three years straight out of grad school in math.
This statement, given by Silver to the annual meeting of the Joint Statistics Meetings (the main cross-organization stats conference), was guaranteed to be a crowd-pleaser for that audience.
Unfortunately for them, it's not really true.
The problem is that much of conventional academic statistics consists of proving theorems about model classes. This requires a lot of sophisticated analysis, but has turned rather vacuous. And much conventional applied statistics consists of computing diagnostics based on dubious modeling assumptions. Under pressure in the last 20 or so years from computer science, machine learning, computer vision, Moore's law, and the data avalanche, the discipline has changed, but not fast enough.
As a result, a lot of what should be taught and researched in statistics departments has been co-opted by these other disciplines. And many people with a real problem would rather work with a "machine learning" person than a "statistics" person.
The best summary of this state of affairs is Leo Breiman's essay (http://projecteuclid.org/DPubS/Repository/1.0/Disseminate?ha...). The abstract of this essay is brutal:
"There are two cultures in the use of statistical modeling to reach conclusions from data. One assumes that the data are generated by a given stochastic data model. The other uses algorithmic models and treats the data mechanism as unknown. The statistical community has been committed to the almost exclusive use of data models. This commitment has led to irrelevant theory, questionable conclusions, and has kept statisticians from working on a large range of interesting current problems. Algorithmic modeling, both in theory and practice, has developed rapidly in fields outside statistics. It can be used both on large, complex data sets and as a more accurate and informative alternative to data modeling on smaller data sets. If our goal as a field is to use data to solve problems, then we need to move away from exclusive dependence on data models and adopt a more diverse set of tools."
Breiman was mathematically sophisticated, so it's not that he wasn't able to follow the theory he critiques, it's that he wasn't snowed by detail and could see its lack of relevance to real problems.
In my experience, improvements in performance tend to come from improvements in domain knowledge. This usually plays out in feature selection, but may involve radical changes in the classifier/model design. The secondary ones are more fun and can more easily be the basis of a PhD. The initial ones are usually swept under the rug when they're brutal and boring, though fundamental changes also form great PhDs.
I think it's fairly clear that the success of NLP almost entirely came down to a feature engineering change.
I don't know what the ideal Data Scientist/Statistician position is. I think it varies based on many of the implicit needs that only statistics savvy businesses can pick between clearly. The role may require strong programming skills, the ability to build and iterate massive multilevel models of social data, a familiarity on-line classifiers, the ability to graph things nicely in R and D3, the ability to from scratch write a non-linear optimizer over a completely novel space, a familiarity with NLTK, or, finally and perhaps most commonly, the ability to pander bullshit about the benefits of machine learning.
I love the field and that's why I tend to side with the statisticians who want to really figure out what's needed to do an objectively good job drawing inference from data. Many in the new generation of statisticians are quite successfully crossing the cultural divide. A smaller cohort are brushing up at least their R coding skills to a suitable degree. Some already have an outside competency in programming. I'd like these people to eventually form the core of what real "data science" is.
He points out a lot of ways the standard "supervised learning, iid feature vectors, moderate dimensionality, moderate n" problem setting is a theoretical construct that does not encompass enough of the real-world problem setting. And he highlights the fact that an epsilon improvement in error rate is not that remarkable.
So, I agree that his paper was smart as a critique of a certain culture within ML practice (and publications) of totally eliminating all problem context, and just saying "Give me a bag of labeled feature vectors -- I will build a classifier with an error rate better than your classifier, and publish the result."
But that is largely a straw man -- critiquing ML hackery is worthwhile and salutary, but it's going after small game. The random forest, support vector machine, and deep learning approaches should have been developed in Statistics. Why weren't they?
There's simply no term for someone who works with lots and lots of data (not necessarily "big data", I don't like that work) through a more general kind of lens that may or may not involve a statistical approach. When the term "data scientist" first really started to float around I got excited thinking it was describing the latter, when it just ended up describing the former.
i.e. A data scientist role is more likely to answer a question like "What percent of people in the White House visitor logs show up in the past 24 months of Reuters News reporting?"
vs.
"Get me a list of people on the White House visitor logs who've shown up in the past 24 months of Reuters News reporting and build a relationship graph from both the news reports and the visitor logs and merge them."
Both require the data munging part of grabbing the visitor log, grabbing the news reports and finding the names in the reports (controlling for various data quality issues). But the "data scientist" approach is to then jump from that onto a statistical answer, while the second approach is to build a social network that might be used for other purposes. These are very different kinds of approaches, with different kinds of outcomes.
It's frustrating that a person in the second kind of approach doesn't have a name, "data scientist" would have been perfect, but the field seems to have gone a different direction with the term.
Sometimes I don't need a mathematically proven process. I often find a use for something like a Monte Carlo simulation, without rigorous methodology, because it works for me, and for what I'm doing.
By analogy, not every website needs a multiple tier, doubly redundant architecture. Sometimes a "dumb" setup with a single webserver and database on a single server gets the job done.
In fact, when I worry less about scientific validity and focus more on what is necessary to get the job done... I make a lot more progress.
It feels a lot like scientific elitism to me.
Really, it comes down whether you need to confirm your assumptions and quantify your uncertainty, etc. Often, in applied situations, you don't. In science, you do.
By freeing myself of that level of effort, I can test a lot more stuff and build things a lot faster. In what I do, that's more important than having something be "scientifically accurate".
Similarly, when I am coding a prototype, I don't do unit tests. Don't need them, and it takes me longer to develop.
The function of a peer reviewed publication is to formalize and verify that some result is correct, reproducible, and relevant (and to advance tenure). One would hope that whether you publish or not is mostly orthogonal to how you conduct your work. If you're doing data analysis in any professional capacity, you should be using mathematically sound techniques, and you should understand those techniques. There's no other way about it. I'm not sure what your line of work is, but I can't think of any technical field where one's analysis techniques being "scientifically sound" is anything but paramount. Why else would the work have any value? Perhaps it helps to think about publishing your work to be more like deploying your code, rather than testing it.
That isn't to say that every step of the way you have to conduct yourself with the utmost rigor. It doesn't mean that you must prove every theorem each time you use it, nor that you have to be laying out strict tests and hypotheses every time you load a data set. But to do data analysis in a scientifically sound manner does mean that you have to understand the background of the techniques you're using, and how they apply to your data. It does mean that you have to periodically bring your analysis back to basics and make sure your ducks are still in a row; that you haven't fallen off the assumption wagon somewhere along the line. It does mean that you should know why something is working, or why something is broken, not just that it is.
"Statisticians" don't sit around all day laying out hypothesis tests in latex. They do the same thing you do; they fire up R, load the data and start playing with it. It's all about the context you have in your brain when you are playing with it. This, I think, is where the formal training becomes quite important. Because as with most things, it's the unknown unknowns that will destroy you, and having the broad formal foundation gives you the tools to protect yourself from walking into a minefield problems you didn't even know existed yet.
I find that higher standard requires a significant amount more time, effort, and ultimately prevents me from accomplishing what I need to.
It depends on the application and how you're using the result. Most of the time, for me, it just doesn't matter.
It may be worth noting that much of what i'm talking about is for personal understanding (e.g., baseball statistics) or is not mission critical and will never see the light of day.
Sometimes good enough is just that.
I do appreciate the thoughtful response you wrote. But often my decision is this: do I spend 2 weeks trying to understand a technique, or do I spend 1 day using it, understanding that my analysis has limitations. When you only have 1 day, my choice becomes "do what seems like it will work" or "do nothing." I agree that more complex techniques (neural nets, etc.) may need more understanding. But because they do, they are also inaccessible to me.
Maybe I can sum it up like this: it's not always necessary to do an analyses in a completely scientifically sound manner, if I can answer the questions to my satisfaction. Especially relevant when the other option is to not get anything done, and get stuck in a textbook trying to understand complex theory.
It's bit me once in a while, but that doesn't matter in what I'm doing. (And its usually pretty easy to recover from.)
Also worth a mention that I would not call myself a data scientist, nor do I operate in that role.
It's all tradeoffs, all the way down. I just happen to choose mine differently than someone who might call themselves a data scientist or statistician.
http://l2r.cs.uiuc.edu/~danr/Teaching/CS598-05/Papers/Collin...
I actually think the easy availability of such toolkits is a good thing; for many applications, rigorous statistical knowledge is overkill. It's just unfortunate that these people call themselves data "scientists" when they commonly lack the years of training, journal publications, and expertise of actual scientists.
Do you have any figures to back that up? I think we should run some tests.
The fact remains that people are unkind to uncertainty and statistical models (as opposed to "common wisdom" or rote learning) requires the audience to embrace and understand uncertainty.
I think data scientist is a sexed up term for a statistician.
fuck-yeah