As the hiring manager for that kind of profile for more than ten years, the problem is rather simple: they want data scientist, but they don’t have good data. The position are still open but no one well-informed would apply.
Managers want to understand what’s happening in their company; they’ve been told data scientist can make that happen, or better, fix the problems automatically. They expect the solution to come without significant investment in data collection. One, even over-paid, data scientist is a far smaller budget than having your dozens of teams make an effort in fixing data collection, and setting up a reporting team. That kind of problem is easy to detect during the interview, if you’ve been there, and any data scientist has. What’s left are fresh graduates without diploma, who were taught statistics, but not how to get a project running -- so they don’t know what to say during an interview.
In more details:
Too much focus on stats. The actual skillset required should not be that demanding on statistical side: interviews generally focus on testing that exclusively. It takes the form of Whomever has any statistical background will grill on what they know best, their own MSc thesis. The thing is: no one can know the minutia of every model, certainly not a student’s interpretation of them, so those end up being a lot of “I would have to go back to the documentation.” Unless your interviewer is mature enough to either know that’s reasonable or stop doing that, it’s a lost cause. At best, you end up hiring a pile of people with the same focus, and every problem will look like a nail to a group of people carrying hammers. One alternative is to ask about something they did recently, but that’s rarely very instructive, either because the interviewer has no idea if that makes sense, or because the implementation has a lot of properties that make the whole process tedious.
The focus should be the ability to learn new techniques at speed. Because the vast majority of the time, the right model is:
a. a basic implementation of a better one, just to get things started (fix the data collection, the engineering of the input to the serving model, think about scaling the serving model, improve the actions taken from the decision, improve the decision format and the actual model objective)
b. a model that you haven’t heard about, and you need to learn about it within a calendar that is conditioned by a. not being sufficient.
Be more project-focused Description like “codes and does statistics better than either” fail to mention that the bulk of the time, a data scientist is actually more of a project manager, trying to identify blockers, get the integration of their tool prioritised, compare the impact and risk of different approaches; there’s also a big chunk (if not all of it) of being an analyst, generally to clarify what problem is being asked, or checking that the data is consistent. That’s if all is reasonably good: far too often, the time is spent trying to get management to grasp that Machine learning can only predict something, and that what to do with that, what to expose to the users, how, when, is more the remit of the product team i.e. you don’t have a real direction. Many large companies actively select against people willing to say that the kind is naked.
Some company test for “business sensitivity” but that too often means checking that aspiring data scientist can walk through a P&L. Reality is closer to: Can you listen to a project manager, frustratingly oblivious to his own contradiction and general ass-holery, rant for 45 minutes without strangling him? Because you need the support of his team to get through.
Fight for good data I have yet to see a company that had data you could use for data science on day 1. The vast majority of the time, it was inconsistent. Not a little bit, or hard to find: big, glaring, obvious contradictions like 20% of the revenue unaccounted for. If there are analysts, they generally have found a way around it, rather than address the issue. In more mature companies, there were some problem that would make basic model stumble, and more nuanced one go to very stupid conclusions. The classic case of a trip-up on good, but not well-structured: What drives retention? Completely failing to delivery your service: people will re-order within the hour, a larger ratio than any other case. Executives pay for the servers so they think that data quality if measured in peta-bytes. They get fed a continuous streamed of over-analysed data, filtered by hand from any inconsistency. They have no idea how inconsistent it really is. Any improvement is seen as a cost centre, rather than the ability for the company to be better informed.
Don’t get me started on technical debt, short-term focus, insensitivity.
What data scientists need, before they even join: clear objectives, clear breakdowns to understand what effects are correlated, clear plans to address some of the concern, pilot tests to make sure that those solutions would work, and collect some data, audited data pipeline with clear documentation to avoid misinterpretation, dedicated engineering to help with data logging, hosting computation, serving model. I am willing to bet that hardly any of those dozen of thousands of jobs looking for candidates don’t have most of that.
There is a small piece of good news: Agrawal & al. summarised their theory on how ML will make predictions cheaper in a all-access book, Prediction Machines. The predictable success of that book will help. Basically, it tells about what ML does, without ever doing any ML at all: it just looks at how it’s a lynchpin technology, like search engine, digital photo, etc.