Unpopular Opinion – Data Scientists Should Be More End-to-End
eugeneyan.com
eugeneyan.com
Then, the data engineering team will ask the generalist to write code to print the path from one leaf node to another in a binary tree at the whiteboard in 45 minutes, and conclude that the candidate just isn't going to be able to figure out how to identify matching terms in different parts of a JSON tree.
Eventually, both groups withdraw into their own silos and confuse each other. They may decry the lack of generalists, but when it comes time to hire, they will resolutely not hire candidtes with 80%ile skills in both areas. They will hire only people with 95%ile skills in one area. They may get some people who have well rounded skill sets through sheer chance, but their process selects against this outcome.
The goal of tight iteration with good alignment is of course a good one, but verbal sleight of hand doesn’t mean the way to achieve it is with more end to end responsibilities for data science.
The huge cost of course is that data scientists and ML engineers have a hugely asymmetric comparative advantage when spending their time on model training and statistical solutions. You want them at full utilization for this set of tasks because nobody else you employ can do that same statistical work, and that work is often hugely valuable whereas most of the end to end work is frankly grunt work and fighting through errors that anyone can do.
If you hire Michael Jordan for your basketball team, why would you make him spend his (expensive) time cleaning up soda bottles or checking the elevators for maintenance issues? It utterly makes no sense and wastes the comparative advantage - all your Michael Jordans will be heading for the door.
Assuming you could have correctly anticipated those design needs earlier, rather than going back to iterate only after you get concrete evidence of the specific changes you need, is a big trap and time waster.
I don’t agree that you lose because the data scientist failed to consider production SRE patterns early enough. You lose when those SREs fail to take a YAGNI philosophy, let the unusual ML system’s chips fall where they may, and then intentionally go back and iterate on improving it.
Especially in the case where you can't deploy meaningfully because your "data scientists" are working in a silo with poor communication.
“You can’t deploy correctly because of data scientists” is a failure of SRE organizations to provide support, tooling and training.
In fact, most data science and ML engineers are quite skilled in systems engineering, because you have to do so much work with GPU hardware issues, underlying scientific package management, efficient data transportation, etc.
Editing some kubernetes config, hardening a high traffic web service or optimizing a query based on an index are trivial by comparison, they are just boilerplate timewasters that need to be a different team’s job to automate.
Are you seriously suggesting that this applies 80% of the companies out there trying to apply ML methods?
If so, it really doesn't match my experience in this and related fields. Of course we'll both have sampling bias here, so I could be missing the big picture; it's not like I've studied it industry(ies) wide.
Failure or non-existence of effective SRE organizations is just one of many common failure modes ime.
This does not align with my experience working with many data/applied scientists at very large companies. A lot of them cannot even write basic code or use git commands. The engineers who are capable are often silo'd away from the scientists, and the organizations struggle to produce any real value.
Unless your data scientists are expected to build the machines they use, they won't be dealing with any hardware issues at all.
Literally every data scientist at big companies use pre-configured vms/notebooks in cloud.
As an ML engineer you spend a lot of your time dealing with Cuda installations, custom compiler flags and then build/compilations of things like Tensorflow, deep internals of Docker image builds to make these environments reproducible, image processing software with opencv and tons of cross platform & software packaging headaches, writing efficient queries and understanding data structure implications for spark, arrow, hdfs, presto, postgres, etc etc, and standing up things like tensorboard for telemetry of ML training systems, deploying mlflow or kubeflow in kubernetes, and so on.
The myth of data scientists as notebook jockeys is just one more symptom of the denial of SRE orgs to admit ML engineers are great system engineers, to try to control them with parochial devops requirements coming from outside specializations.
It's most obvious with this statement, but overall you seem to think "ML engineer == data scientist", which just isn't the case.
I manage teams of both ML engineers and Data Scientists and have designed hiring processes for both within large ecommerce companies for years.
I'm not the other person, and that is definitely not what I'm saying. There can be overlap, but the distinction is important: Data scientists use tools for analysis, ML engineers are capable of building those tools.
For example, the tool I'm aware of that some of our data scientists use is SPSS - but they have no programming experience, and could not remotely be grouped in with "ML engineers".
If people interpret it as “elitist” to refer to basic comparative advantage economics, they need to get over it and stop being so sensitive.