But you have to interpret the data within the context of the business need/requirement.
Building a credit risk model is vastly different from building a (personal) insolvency/bankruptcy model though both may entail the same set of steps in developing the model. The variables that make it to the model depend on the business need.
In kaggle, one of the datasets that I messed around with had variable labels as Var_1, Var_2, ..., Var_X. So while fitting a model, I would not know why a particular variable made it into the model. You can see that this kind of variable labeling does not give me any insight into how that variable was generated. I need to know whether the variable was raw/aggregated/transformed etc. And that takes you back to understanding the data in the context of the business/domain.