The reason it carries value is the skills are difficult to acquire. I think the recent decline in interest reflects the rise of new data science candidates that are taking the path of least resistance to a career in data science. Rather than pursuing problem solving, people are pursuing "data science" which is a nebulous term in and of itself.
Machine learning is an area where you need to be able to produce results. Fake it ‘til you make it isn’t going to cut it for long. Either these people produce something that works, or they don’t.
Having to produce results is one thing. Mindlessly throwing tensorflow/pytorch at problems is an entirely different problem.
It's like those front-end devs who mindlessly insist that they need to use heavy javascript frameworks with convoluted build processes such as React/Angular to churn out a static web page with a couple of paragraphs and images.
Well, no. ML solves the classification problem, not the prediction problem.
E.g.: The "is this a cat picture" problem is effectively solved, but we _still_ can't reliably predict something as primitive as a simple binary proportion.
You can certainly produce some plots and numbers, and possibly even plots and numbers that look good to your boss/clients/investors. The (multi)million dollar question is whether those numbers are actually meaningful. I think this is where a lot of ‘data science’, both in industry and academia, falls down.
Some state-of-the-art models don’t even generalize to test sets drawn from the same database, let alone similar data sources or the actual business problem. Unless you run a pet shop, telling breeds of dog apart, a la ImageNet, is probably not your goal.
Can you quote a source for this?
What makes me cringe the most is to see flashy presentations with claims akin to 'Data Science will change your world'.For sure, it can and has been proven to automate decisions (think, credit scores), assist in decision-making (think, sales trends) and anomaly detection (think, security systems). I find so many data scientists that I interview are so hung up about the esoteric techniques they employ, often failing to even explain why was it useful or how it helped their businesses.
What has been transformational and path-breaking is the breaking of enterprise monopolies in this space (for e.g. SAS/IBM SPSS) and a variety of open-source frameworks have made it easy and convenient, apart from opening it up to software developers to build these skills. Important, though, to not lose of the sight that data science is at the sweet spot of expertise in domain, data and technology.
What companies want is to be "in" on the data science hype, while they have no clue what they are doing and the most advanced "data science" they need are simple graphs, boxplots, and linear regressions.
Yeah it seems to have calmed, but I don't think data science was just hype because it comes from (and somewhat is) probability and statistics, and the rate of data/information that's being produced by and extracted from people seems to be ever increasing. But it absolutely was prone to a hype cycle as with almost anything else in tech. IMO this is a phenomenon exacerbated by venture capital.
I think once the hype calmed down, people started to realize that it was largely the same old shit in a much cheaper package — evolutionary rather than revolutionary. Ultimately I think the hype cycle was driven by Moore’s Law more than anything; the fact you could run this type of analysis in a manageable amount of time without needing a huge IBM mainframe was the real innovation.
https://github.com/nemild/hack-the-media/blob/master/softwar...
Not just VCs. It's a whole mafia gang consisting of tech reporters and founders also. They all have their vested interests - reporters want new stories and founders want funding and growth.
Slack, VR, AR - they all went through this cycle. Sometimes, it's a bit annoying.
Several of the Data Warehousing projects I've dealt with could be better described as Data Landfills. One can't just dump data into a hole for years, let it rot, and expect goodness when you go back to look at it.
The realization is that any random pile of data likely doesn't have anything in it that is worth paying for: Here's our analysis! We already do/knew that.
Some people have really good, valuable data sets. Most people don't.
The map is not the territory.
https://www.amazon.com/Raw-Data-Oxymoron-Infrastructures/dp/...
Data Scientist is a buzz word for Statistician. Business Analyst is buzz word for Industrial Engineer. For example 10 years ago if you studied at my university you would witness that some Statistics students were doing second major mostly at Industrial Engineering and vice versa. They are already related for many years but average Joe has no idea.
I'm still not sure what their actual formal responsibilities are.
They ask questions to identify areas of improvement; translate those into functional (and sometimes technical) requirements for other areas (not just IT) to fulfill; and then coordinate the efforts to implement those requirements, potentially as PMs, product owners, Scrum masters, UAT leads, or just a SME.
The best BAs (paraphrasing the data science JD) know more tech than the business and more business than the tech.
It's really not. The skill set we need in terms of some software, system design, and a rich knowledge of modern data science libraries and trends is not something you should expect a statistician to have.
Similarly, I certainly cannot prove asymptotic theorems like a statistician.
I'm not trying to slight the Dept. of Stats. I love statistics. I've found studying PhD stats textbooks more valuable for my data science career than learning the latest deep net framework. I'm just noting there is a need for other tools.
If you think of Data Science as AI sure, but if you frame it as applied statistics + good software engineering practices + cloud scale I think things are in a good place.
I once saw some code written by a "data scientist". The overall code was non-complex, but the Java/Scala code was the worst of my nightmare. Additionally, I think other engineers have also matured enough to understand that underneath the veneer of data science, the fundamentals do not change much.
At some point there is no need to hire a "data scientist", as any python programmer is already expected to know how to use numpy, pandas, sklearn and keras, just like before it was already expected for them to do any kind of data manipulation with SQL without requiring a dedicated database expert.
Of course, if a company wants to truly innovate in the area it will need PhDs or people with great dedicated knowledge in ML/Statistics/Particular Domain, if it needs to scale it will need good data engineers to create the data pipeline together with DBAs and experts in each tools (like Spark/Flink), but for most companies the basic above is already a great improvement to what they had before.
When someone says they want a "Data Scientist" what they really mean is "I want a Data Scientist who is also a Data Engineer".
I have seen so many companies spend a really decent chunk of money on a data scientist and then are shocked to find that this data scientist doesn't know how to deploy models, set up spark clusters or know how many and what type of GPU they need to use to get the job done.
After all - that is not their purpose.
We were in a similar situation, but what we needed was a Data Engineer - we had a rough idea of where we wanted to go and what we wanted to achieve, he was doing a Masters in Data Science so he had that background as context.
We will look at adding a Data Scientist to our ranks in the future - but they will be working side by side with a Data Engineer who can action their requirements!
I think the term "data science" is often misused. It seems to make management feel like they are on the cutting edge. They were talking about AI and a R&D department the other day. They aren't even making use of simple heuristics yet! I guess that talk helps with fundraising though.
Data collection will become more prominent IMO because:
1. Data driven business preference, competitive advantage and FOMO. Already dominates sales and marketing. Starting to dominate in product and dev. Already dominates production.
2. IoT, and more data marketplaces resulting from it.
3. Extensions of the global SaaS value chains (usually connected by data).
Hence data science will thrive in the future.
Enterprises seems to hire less data scientists actually, but they are trying to raise their employees' data skills.
I think that's the cause of the growth of self-analytics tools. Below are examples of them.
1. Metatron Discovery : https://metatron.app 2. Metabase : https://metabase.com/
I used to work for a small start-up and the CTO was very strict on data access, making my life as feature developer and "data scientist wanna be" almost impossible.
He, on the other hand, had not only access to all data but also used the product as a consumer (which didn't make sense for ICs so we ended just playing with sales demo accounts). I ended leaving the company because of that.
What boils down to is that people who have any extra data access privilege will have the lead.
Most of the insights will come from aggregate data, so I think companies could work around privacy concerns but I am no GDPR expert.
Back in my days in academia, there was a saying “if you have the trace, you have the paper”.
On the other hand, your standard tech data scientist may find themselves out of their element if needing to design a very rigorous randomized trial for testing a new drug, and making careful inference (I mean I'm sure plenty could, but I'm not going to trust a 25 year old with two years work experience to do that).
The reality is that most insight from 'big data' are optimizations. They're not going to move the needle on the business as perhaps we would have hoped.
Data Science focused on ad targeting - now that might move the needled.
And of course, maybe some Data Science working along side AI engineers make a breakthrough which could move the needle.
But from a high level, CEO's view, all of these things have trendy undercurrents, the trick is to figure out how much of it really matters to the business.
The 'wins' for consumers will be slow: maybe better product search, better ads. Maybe they figure out how to send flights around ugly weather or how to slot landing times for an x% decrease in flight delays. Or slot road fixing/lights for an x% decrease in traffic delays.
I'm not sure what it would mean for days science to be "just hype." I see DS work on this website alone all the time.
> Since academia is typically a lagging indicator in adoption to new trends in the work place, it’s been long enough that it’s truly worrying for junior data scientists, all of who are hoping to find data science positions. It can be very hard for someone with a new degree in data science to find a data science position, given how many new people they’re competing with in the market.
http://veekaybee.github.io/2019/02/13/data-science-is-differ...
That being said I've observed that the data science techniques and tools that have developed over the past few years have been absorbed and adopted by a lot of people that aren't "data scientist". So while companies are hiring a lot less "data scientists", a lot of "data science" is now done by domain experts and analysts as part of their work.
I’m not from the US and in my country, “(civil) Engineer” is a legally protected title https://en.m.wikipedia.org/wiki/Civil_engineer#Belgium
So the field keeps growing, but in more specialized forums. Increases in Python adoption, for instance, mostly come from re-skilling initiatives in BI teams inside my customers.
Can I ask what city you work in? I'm going to school for CS and I see a lot of students trying to get into DS with little or no avail.