I am glad to see machine learning, ai, "data science", whatever, grow as a separate field. The statistics programs had their chance.
Title inflation exists, but there is a real-world role here that isn't really captured by "statistician" at all.
To me, data science is more than understanding statistics, it's been essential to know how to scale them up and out.
If you're a domain scientist, you won't necessarily learn how to write reusable tools that are performant (or runnable) on data that is different from your initial model data. I once worked with a group whose model had grown so unwieldy that their config file was in NetCDF.
I found my niche was often in doing things that were slightly (or completely) outside the comfort zone of most domain scientists who were competent coders themselves, but who didn't have the funded time nor the inclination to learn things like database, visualization, and networking technologies that became necessary either to share their work with other research groups or to operate on larger datasets.
One project had me take a big model that was normally run twice a day and on a 4km grid and help write something that could run and visualize the results of the same thing on a 0.5km grid over a larger area and hourly. And then devise something that could help them visually explore the timeseries as it evolved, sometimes over months.
Designing the pipeline that can handle that is outside the scope of most scientists, even the ones who are good coders.
A traditionally trained statistician would evoke negative result and decide not to use the model and support to maintain the pre-existing approach. A machine learning expert might not care, apply the coefficient out of the model as is because they are presumably closer than a guess and is more likely to be openly skeptical of human expertise.
That has lead to some frustrating situation for me: me arguing we should censor things like negative speeds, while I was told that there was no problem because the results were regularised anyway. Building and picking proper factors to use in regression is something that you can partially get away with when having larger databases, and back-propagation can take over; before that, insights still do matter.
I have not meet many who can articulate that transition effectively.
It seems that you’ve met mostly the second category; they are possibly the larger group, but not necessarily the most influential. There is a core of people who are meaningfully different. The linked article seems to be from someone in between but closer to the second group.
Genuine question - more than happy to be proven wrong.
We do that because:
A it helps us understand them better
B it teaches us how to think, the way Feynman said "Know how to solve every problem that has been solved". Granted, it seems pointless to work through what is easily accessible through machine BUT it teaches how to solve new problems. I wouldn't consider using NumPy or Matlab as the first step towards solving a new math problem.
It's like using Assembly vs using a higher level programming language.
edit-This is of course completely anecdotal experience.
They are both more and less, in my experience, than statisticians (more flexible and solution-oriented, less rigorous and classical), than analysts (they can do more, in general, but a great analyst will be better at analysing and visualizing), than developers (they know more stats, less software engineering, and have great patience for wrestling data into submission). I like to think of data scientists as people who combine the skills of all the above to solve hard problems which exceed the domain of any of specialty (analyst, statistician, developer). It doesn't mean we're amazing at everything, just that we are effective, flexible problem solvers.
And for the record, machine learning, statistical modeling, and data mining are just a small portion of the pie. Being good at modeling and machine learning will not remotely guarantee success as a data scientist.
I could of course be wrong and have a bit too narrow of a view from my particular subfield.
Why would you waste your time re-inventing a wheel.
A good data scientist isn't good because he/she can ace shitty trivia, he/she is good because they know the right question to ask.
In those situations math isn't "shitty trivia," but instead a tool to be leveraged against those hard questions.
You can consider the derivation of SVD to be shitty trivia while throwing np.linalg.svd around while engineering features. That's fine! Good luck visualizing that data in a meaningful way, or dealing with non-linear data, if you're ignoring that "shitty trivia."
What is non-linear data?
That is to say problems that can't be expressed by linear functions.
I.e. Y= mx + B is a linear function.
Y= ax^2 + bx + C is a polynomial (non linear) function.
Linear Programming (LP) involves solving a series of linear equations (something like Excel's Solver can do this).
When you are dealing with non linear functions you need to use a method such as Sequential Quadratic Programming (SQP).
— Stanislaw Ulam