Analysis of the data job market using HN job posts
emiruz.com
emiruz.com
The largest group of DS was non-ML/CS/Math PhDs who started panicking once they realized their future job prospects in academia were very slim and so they signed up for bootcamps and got jobs at places hiring DS by the hundreds. Many of the people in this latter group had no idea how to write Python outside of a notebook, generally just structured problems to fit into XGBoost, and when not doing that tried to squeeze resume-boosting-complexity into any problem the could find. They also tended to have a hilariously poor understanding of creating business value.
Nearly everyone I know in the first group has switched back to just being an engineer of some sort, typically ML or AI engineer. I suspect the small set of talented people from the second group will end up in lesser paying product analytics type roles or closer to product management roles, while the majority that don't bring much to the table other than a PhD will be slowly attritioned out of the field as companies start looking for the value different skillsets bring to the table.
To be fair, for 90% of business problems that require ML, I’d rather take the guy who throws XGBOOST at everything instead of the one trying to be fancy with neural networks. You get an explainable output and good results without deep subject matter expertise that the person likely won’t have. It also runs at a fraction of the cost.
To be fair though, most value from data science is driven by the analysis, not the models.
Is this different than your average SWE?
I've been a DS for 10+ years, and I feel the exact opposite. The worst "Data Scientists" I've worked with are all ex Software Engineers who seem to assume that business problems are really computation problems. So they find convenient ways to ignore the human aspects (e.g. trying to figure out why the data is a mess) and gravitate to using more complex algorithms and breaking down the problem to an achievable programming pipeline that runs in production, but the results are of low value. But it looks awesome on a resume.
Are you right or am I right about SWEs turned DS? I have no idea. But one quality that IMHO is important is the interest in actually looking at data and asking questions, which is much rarer than most people realize.
It doesn't sound much like your worst and their best is the same kind of person. I don't see necessarily conflicting views.
@OP - mind rerunning this analysis for "AI Engineer" titles? https://www.latent.space/p/ai-engineer anecdotally i saw 8 of these in the last Who's Hiring and wanted to tease out the emerging difference between ML and AI Engineer
So for newly minted Math PhDs, sure go out and learn how to do some ML coding in notebooks, but if you can't get a decent dataset together to train it on it'll be all for nought. Anyone with AI/ML coding only, and no SQL, is a no hire in my book.
I'm already struggling to parse the first paragraph -- is this a typo? Or do they mean a role that is both ML Engineer and Data Scientist?
Companies think they need a "data scientist" but they actually want a "software engineer with enough stats/data background to implement data science algorithms in production."
The result is that a lot of "data scientist" jobs are mostly data engineering or ML engineering and that is reflected in the list of requirements. It makes finding a good job extremely difficult, and it's for the better that the "data scientist" title is being eroded, because it makes it easier to tell when a "data scientist" job is actually data science as opposed to something else.
Not that "data science" is a good name anyway, but that's another story.
Missing is the Data Analyst component, and (as is normally typical in the discussion) the statistical experimenter (A/B, MAB, etc.)
Took me for a bit of a spin because when I was first interviewing this year, I'd apply for DS roles and only get interviews that were very stats heavy with a leetcode easy, but started getting further when I basically stopped applying for DS roles and went straight for MLE roles.
[1] https://medium.com/@rchang/my-two-year-journey-as-a-data-sci...
Systematic MLOps helped to decrease _some_ of that, but not nearly enough, and certainly not with the recent explosion of LLM-induced hype.
I view MLOps engineers and ML engineers as tasked with scaling the problems, and research scientists as the scoping and pioneering of solutions. All three fall under the larger umbrella we call "Data Science" IMHO.
Thus, we end up with a significantly weakened analysis plan and execution. Much to the disappointment of everyone involved.
Moreover the people analyzing it don't need to be data scientists -- they can as easily be statisticians, economists, geneticists, etc.
While data science is a newer academic field, most of its practitioners come from statistics, economics, genetics, etc.! (I say this as a 10-year data scientist who is an economist).
I'm not sure this is accurate. To the extent that a data project is underspecified (which, let's be honest, all projects are) then the engineers will end up making some decision somewhere that may have an impact on what's available for analysis.
If the engineers have some understanding of project motivation and hypotheses then they'll make better decisions.
I believe you should edit the text to read "Data Engineer" in the first numbered item: "I argue from data that the Data Scientist role is poorly differentiated and I speculate that its responsibilities are being eroded by better specified roles such as ML Engineer and Data [Engineer]."
I'd love to see an analysis on this thesis!
Or more generally, how various sources of new-hire job descriptions correlate with each other: HN Who's Hiring and other job-advertisement boards; LinkedIn profiles; etc.
And same thing for programming languages: appearance in job postings vs. TIOBE index vs. ...
If former then maybe the thesis holds true, but data scientist in more established companies do vastly different things than in startups.
It’d be good to extend the analysis by accounting for company size.
(I wrote a blog [1] making a call that Data Engineering was likely also to see something of a relative demand decline, or be better defined into Software Engineer and Analytics Engineer, so I am quite interested in your analysis here)
https://groupby1.substack.com/p/data-engineering
> Most businesses' data engineering needs have been solved or will shortly be solved by managed services that 10 years ago would require endless and extensive self-built ETL pipelines, databases and tools. For the exceeding majority of businesses, this means they can and should focus on building capacity for business logic, analysis and predictions instead of data engineering.
My thoughts: the tech job market as a whole has been in decline, as unanimously observed. There may be some signs that the slowdown is abating though. Next 6-12 months will be key to see how DS and DE rebound (or not).
> Most businesses' data engineering needs have been solved or will shortly be solved by managed services that 10 years ago would require endless and extensive self-built ETL pipelines, databases and tools. For the exceeding majority of businesses, this means they can and should focus on building capacity for business logic, analysis and predictions instead of data engineering.
Could not disagree more with your take of "DE demand will decline due to DE needs being already solved for most businesses". Apologies, but have you ever worked as a data engineer or even close to one? Pipelines break, requirements change, businesses expand, and infrastructure needs to be managed and optimized, etc. ETL processes, in the wild, are decidedly not one-off affairs.
Maybe the analysis as to _why_ is wrong, but that is what I'm trying to unpack.
HN Guidelines: "When disagreeing, please reply to the argument instead of calling names"
A lot of what modern data engineering has turned into is connecting various tools and software together. So adding another managed service doesn't feel like it is going to magically solve the problem. It is going to be just one more tool that the DE's will be managing. And indeed that has been my experience. For every tool that we added to the stack, we ended up spending just as much time fighting the tool as we did maintaining the self-built solution that the tool replaced. The two main advantages of using tools over DIY solutions is that they have an opinionated way of doing things, and they usually come with extensive documentation. So on boarding a new team member is easier. But engineering hours saved is pretty much a wash compared to DIY once you hit that first edge case that the tool does not handle elegantly.
This is probably the exact crux of the argument, and I wonder if it reduces the overall demand relative to the growth in the industry. Do you need 10 engineers (like a decade ago) when 1 (great engineer) can deal with the bulk of the work with "better tools" and then DIY on the edge cases. This has been my experience with the modern ETL tools.
But I don't think this is a long term trend. The DE role was originally tied heavily to the rise of data science, but it has turned into more of an operations role since then. You can probably get away with slowing down hiring for a little while, but just like with janitorial services, if you cut back too much things are going to get messy.
Edit: why disagree?
>It is likely that the Data Scientist role is in a long term decline...
Also
> Data science is in decline and vaguely defined
Reading this, you can think that "Data Science" jobs are decreasing. But I don't think that's true.
Let's just say that it's 2017 and I hire a team of 3 people with the job title of Data Scientist. One ends up focusing on the data side, one on modeling+analysis, and one on building the infrastructure. In 2023, I decide to change the job titles so one of them is now a Data engineer, one is now a Data Scientist, and one is now a ML Engineer to match what is happening in the job market.
It's still 3 jobs with 3 people doing the same thing. So the number of jobs aren't decreasing, but their titles are more specific. Overall, the number of "Data Science" jobs are still doing up.
Somebody will say "But that's exactly what the author said." But I think people who are new(ish) to this field might read it as "Data Science Jobs are decreasing." So I'm making this comment.
> skills such as data mining and visualisation are also out of favour.
Honestly, I just don't believe this. It's possible that as job descriptions are filled with different buzzwords, people just leave these out. For visualization it's also possible that there is a bigger focus on keywords of an established BI tool (e.g. PowerBI) instead of ad-hoc charts in matplotlib or ggplot. But some degree of data mining and visualization is useful, even to Data Engineers.
I think it really comes down to a lot of marketing, which you touch upon a bit. AI is in a hype cycle right now and people want it on their products and in their companies, so they want people that are capable of bringing those skills to the table.
Nah, recent times have been "different" for everyone. If the author entered the job market in 2011, then this is the first economic recession they have seen. Economic downturns happen roughly every 10-12 years or so and generally cause a fair bit of turmoil.
Forest fires are a necessary part of a healthy natural wooded ecosystem. I tend to think of economic downturns the same way. Every once in a while, companies (or entire markets) have to look carefully at what really adds value to their businesses and figure out how to focus on that when cash flow dwindles and investors clam up. If a business doesn't survive a recession, then it was on shaky ground well before the economy went south.
The overall economy is not doing well right now, at least in the USA. There is significant amount of extra talent (supply) on the market due to big tech hiring freezes and layoffs, which also contributes to lower demand of new roles. There is the return to office debacles occurring all the while housing is becoming even more unaffordable due to high interest rates (compared to the recent 2 to 3% during covid) and low supply a consequence of the rates where owners aren't going to want to trade a 2 to 4% for 7+ nearly 8% right now. I don't cite rto debacles to have a debate on the specifics of if rto is good/bad, but I would speculate that it's pushing more talent onto the job market to escape working environments forcing any style (rto or forced remote) that an employee disagrees with.
So, in my mind the only thing my speculation doesn't really cover is how that would contribute to lower HN responses to which I don't have an answer, maybe it truly is shrinking. However, my gut says it's the economic factors. I think (and hope) that the shift in conditions occurs in the next 6 months and hiring ramps back up as companies recover and adapt.
It's almost impossible to find a genuinely useful job there, regardless of experience.
In recent years, I’ve always gone via agents who contacted me on LinkedIn - which has become something like a naffer Facebook.
Perhaps I ought to pay a little more attention in future.
It the description given here (data mining and visualization), and the fact that for a long time this was (I’m pretty sure) advertised as a bootcamp-appropriate sort of role seems to indicate that make this is not a role in and of itself?
A little coding and the ability to think about data seems like a generally useful add-on skill for most roles? Maybe a we’re seeing unsatisfied need for, like, office workers with some technical proficiency?
I think the need for a broader umbrella term is still present and I think that explains the wide adoption of the word "data science" in the first place. But the current meaning has been stretched way too far.
> Maybe a we’re seeing unsatisfied need for, like, office workers with some technical proficiency?
Specifically, data analysts who also know Python.
1. Is this a true decline, or in line with general tightening of the tech economy?
2. Where are the "analysis" conventions going generally -- HN is going to be a weird subsample of the economy as a whole, given that BI, DA, BA roles still exist and overlap with DS -- on top of that, many industries still haven't adopted Research Scientists, MLE, MLOps Eng, etc. into their lexicon of roles
Few months ago I've created a platform to analyze jobs based on Google for jobs data in real-time.
You can search by job title and location, and it will give an indication about the job market.
I did it as I wanted to understand which publishers are appearing more on Google For Jobs, and how many jobs are remote in certain locations.
However where I draw the line (and where I think most data scientists should draw the line) is actually putting that stuff into production code. Maybe they're good enough to write the prototype, but you need somebody else on hand to help with test coverage, make sure it meets performance requirements, triage bug reports, etc. if you make your data scientist responsible for that, they are going to spend all of their time doing that, instead of doing the things that they are actually trained to do and that you are paying them to do. This is true even if they are a perfectly competent software developer.
1000%
I no longer hire data scientists. I hire recent CS graduates then they cut their teeth on that work because it is very easy, typically low risk, and they can learn the basics of software engineering.
In fact, this is what I've done at several startups now: take prototypes and productionize them. At this point, I'm almost a specialist in transitioning from early stage to growth stage. I can do prototyping, but I recognize that productionizing is my stronger suit.
The technology for the prototype doesn't matter — Jupyter is fine. In fact, it's better if the prototype is shite because then it makes it easier to make it the "one to throw away".
I have worked with "data scientists" that are more rebranded statisticians. Some are overpaid SASS users, but some are true statistical wizards who can also program (R/Python/SQL), and when you need them they're great. I don't need them to write production code. I'd rather call them statisticians but I'm glad to see them getting paid better!