Does kaggle publish how the models performed under normal business conditions?
But you have to interpret the data within the context of the business need/requirement.
Building a credit risk model is vastly different from building a (personal) insolvency/bankruptcy model though both may entail the same set of steps in developing the model. The variables that make it to the model depend on the business need.
In kaggle, one of the datasets that I messed around with had variable labels as Var_1, Var_2, ..., Var_X. So while fitting a model, I would not know why a particular variable made it into the model. You can see that this kind of variable labeling does not give me any insight into how that variable was generated. I need to know whether the variable was raw/aggregated/transformed etc. And that takes you back to understanding the data in the context of the business/domain.
1) a key company policy 2) specific business activity 3) input coming from another model
It is better to
a) define the problem, b) collect the data, c) build the variable library, d) and then fit the model
rather than jump to step (d) directly because the modeler/scientist has greater understanding of the entire set of data going into the model development. It is very likely that the modeler would uncover any/all of the three influencing factors I mentioned above, during the data collection stage.
While kaggle is an interesting concept, from a different perspective it looks like an "effort harvesting" operation. For a pittance, the companies/institutions that are sponsoring the contests are getting a steal. (I am not sure if the million dollar prize is still up for grabs.) However, for folks who do want to break into data sciences/statistics field, kaggle certainly is a good platform to get acquainted with data science/statistics related skills.
For example, early analysis of a "fresh" data set involves a lot more "'of course length of hospital stay' correlates with 'had disease'" or "'Phone number' is a unique identifier in this data so of course it shows up most often in random forest predictor" moments.
Putting together weighted ensembles in R of a clean data set is basically manual labor but doing the same analysis at scale involves a lot of of nuances that should be considered data science. In particular you need to be able to determine if a problem can be split into mostly independent parallelizable sub problems (which also often requires domain knowledge) or reduced/relaxed into something that has a well established optimized solution (like matrix decomposition). And finally you need to be able to determine if your prototype and the production code are converging to the same results and debug it if it isn't which is non trivial with stochastic algorithms.
That's irrelevant, data scientists don't do data mining contests for a living. In my experience finding the right question to answer is a large chunk of data science, and that is never spelled out for you like it is in a contest.
Yes. Your view of data science is extremely narrow, there is more to it than creating a predictive model to optimize a single metric. Reread the second paragraph of my first comment.
This splits the "domain expertise" and "predictive modeling" components into two separate chunks. While domain expertise is crucial for asking the right questions, we've found that it isn't as necessary for the predictive modeling component. For example, in the essay scoring contest we hosted, none of the winners had touched natural language processing prior to the contest. However, they beat out many experts with decades of experience in NLP.
For an internal data science team, the "domain expertise" component is at least as important, as they are charged with asking the right questions as well. However, this does not mean competition winners cannot develop and learn this - they have already demonstrated their creativity and tenacity in one domain (applied machine learning), and this carries over nicely to other domains from our experience.
edit - for clarity.
Strongly agreed. I actually got into the field through a company's data mining contest (pre-kaggle). I think people who are strong at building predictive models are great candidates for data sciences. But it took years of work experience after doing my graduate degree in machine learning to get to a point where I'm comfortable calling myself a decent data scientist.
It's easy to think model-building is the only important skill-set, but data and models don't exist in a vacuum; a more holistic view of where your data comes from and how your work will be used is essential. This excellent netflix blog entry illustrates what I'm saying quite well, I think. http://techblog.netflix.com/2012/04/netflix-recommendations-...
They make two points that illustrate the divide between a contest and the day-to-day work of a data scientist:
* The winning model was not usable in production. Netflix had to gut the 100+ model ensemble to a much simpler 2 model ensemble
* Business needs change, the question they were trying to answer changed from the start of the contest