Can't we just be honest and say that most of these are applied statistics jobs with a specialty in large volumes of data? Or is "statistics" just not fashionable enough nowadays?
Can't we just be honest and say that most of these are applied statistics jobs with a specialty in large volumes of data? Or is "statistics" just not fashionable enough nowadays?
IMO, the "data science" label is too broad to properly differentiate statistical engineers. A fine definition for a data scientist is someone who runs experiments on user/company data and can assess the results. It's important work, but you don't need a PhD in stats or ML to do basic hypothesis testing.
You could simply call them "machine learning" experts, but that could be a bit too academic. People who are focused narrowly on theory or niche areas may be experts in ML, but they may also never do anything outside of running matlab simulations. It's unlikely that those people will make very good statistical engineers since they may never have had to think about the challenges involved in scaling algorithms.
My preferred term is "predictive analytics," which I feel kind of straddles statistics and machine learning, and also serves as a nod to a common difference -- "statistical" methods often yield understanding, while "machine learning" methods are often opaque to human insight but yield predictions.
But humans are still the ones responsible for important high level decisions, so it still makes sense to maximize information transparency to enable good decisions in those contexts.
A neural network that given a prediction 'X is most likely' and could answer the question "Why?" with 'Because Y' would be amazing.
It would be a lot more fair to classify machine learning as a subfield of convex optimization. Yet even that classification does not quite fit, so it makes most sense to just accept that it's a separate field which uses techniques from statistics, convex optimization, computer science, and more.
The most value I got from this article was in the realization that, every few years or so, the academic globes align well enough (some paper de joure becomes well-read I suppose) that .. for a brief instant .. terms are defined well enough, and gain enough agreement, that progress is made .. which progress attracts more eyeballs, who tend to want to break off a chunk for themselves, and the terms begin to differ again and we have a whole new 'sub-sub-sub-' variety of the subject.
So its all about globes aligning, basically. I will now go off and implement an AI technique based entirely on the description of globes, alignment, and little chunks breaking off every now and then .. see you at the top of the AI heap in a year or ten.
I think many would agree that "machine learning" and/or "deep learning" are at least cornerstones for "artificial intelligence". After all, nobody singularly defines intelligence.
Statistics is not (just) opinion polling, there's a lot more to it than estimating observable properties of a population.
If you're trying to make decisions, predictions or estimates which involve any uncertainty at all (and in my experience big data almost always is), then it's definitely within the purview of statistics even if you have data for the whole population.
Sources of uncertainty include trying to say anything at all about the future (do you have data on the future population? no didn't think so...), trying to make predictions which generalise to new data in general, trying to uncover underlying trends or patterns behind the data you see which aren't directly or fully observed.
Often people expect big data to be able to answer big numbers of questions, estimate big numbers of quantities, or fit big, powerful predictive models with lots of parameters. In these cases statistics can be particularly important to avoid reporting false positives and to make sure you can quantify how certain you are about your results and your predictions. (Amongst other reasons).
Once you hit millions of rows, it's not humanly possible to survey the data. All you can do is make assertions about the data's structure / buckets it will fall into. You then try to disprove that assertion, or establish an error bounds on it. You will never see all the data, only the results of assumptions you've made about it.
The new machine learning is about building layers of components on top of each other, very much like circuits seen in EE. The "circuit" components being used are no longer well defined mathematical pieces built from the bottom up using ideal assumptions, but less well understood, somewhat black-box newer components that were built from the top down. Far more like a type of engineering than a type of statistics.
If you haven't been seeing all the latest Arvix papers, you're really missing out. It's evolved to look sharply different than statistics now.
In general, AI borrows many more techniques from mathematics than it does from statistics. However, the field of AI has been quite established since the 1960's, and many techniques have been developed within that field as AI techniques, it's more about being accurate than about being fashionable as AI simply isn't 'just' statistics.
There might also be a demand for applied statisticians, but that doesn't make AI experts statisticians. I understand the confusion, as the term AI is often misused, but when you see the names mentioned in the article it's clear they're talking about actual AI researchers.
On the other side the big tech companies are investing heavily in Deep Learning for things like NLP, Speech, Vision, Siri, and wherever else these neural net approaches may work etc...
But isn't this the approach Nature herself is taking?
The knowledge engine you carry in your head spends years just "learning" the world - which means, it absorbs huge amounts of input, sorting the good stuff from the bad. It "knows" what works simply because that stuff happens more often; it "knows" what doesn't work because that stuff doesn't happen very often.
And sure there are higher layers of integration there, but the whole process is strongly supported by a statistical approach.