But you have to interpret the data within the context of the business need/requirement.
Building a credit risk model is vastly different from building a (personal) insolvency/bankruptcy model though both may entail the same set of steps in developing the model. The variables that make it to the model depend on the business need.
In kaggle, one of the datasets that I messed around with had variable labels as Var_1, Var_2, ..., Var_X. So while fitting a model, I would not know why a particular variable made it into the model. You can see that this kind of variable labeling does not give me any insight into how that variable was generated. I need to know whether the variable was raw/aggregated/transformed etc. And that takes you back to understanding the data in the context of the business/domain.
1) a key company policy 2) specific business activity 3) input coming from another model
It is better to
a) define the problem, b) collect the data, c) build the variable library, d) and then fit the model
rather than jump to step (d) directly because the modeler/scientist has greater understanding of the entire set of data going into the model development. It is very likely that the modeler would uncover any/all of the three influencing factors I mentioned above, during the data collection stage.
While kaggle is an interesting concept, from a different perspective it looks like an "effort harvesting" operation. For a pittance, the companies/institutions that are sponsoring the contests are getting a steal. (I am not sure if the million dollar prize is still up for grabs.) However, for folks who do want to break into data sciences/statistics field, kaggle certainly is a good platform to get acquainted with data science/statistics related skills.