2. Cleaning/understanding data - nearly all data sets I've used have duplicates, missing data, highly anomalous distributions in some fields that indicate we aren't measuring what we think we are, etc. So a lot of my time is spent figuring out what's going on in the data, cleaning up the issues, figure out what subset of the data is reliable, and then dealing with the biases introduced by what is missing or wrong.
3. Dealing with people who don't understand or respect statistics and data science. For example, I've been brought in to "do the analysis" on an "A/B test" where a team didn't appropriately randomize their samples, and also hadn't done a statistical power test beforehand so had an underpowered test anyway, so there was just no hope of validating that their change was an improvement.
I want to know if there is a way to spread the load for this.
Work on the coq program proving the calculus is taking longer than expected though.
If you're interested in contributing to Gorgonia, lmk.
An expectation that data can eliminate the need for reason and thought is problematic.
I have tried to communicate the reasoning for things like judgemental forecasting but success is hard to achieve.
One of the things my team does is build data tools. The number of people who want to take data they have hardly looked at, put it through a tool they don't understand, and rely in important ways on the output, is astonishing to me.
If you have to interact with enough of these people your job as a Data Scientist will be miserable.
It's incredible how many "Data Science" problems can be solved with a better dashboard that enables people to look at data in a more useful way.