Data Science Is America’s Hottest Job
bloomberg.com
bloomberg.com
The gatekeeping in this field surprised me because in my study of data science and machine learning I did not think the practical use of these techniques was that hard. The math isn’t even that hard if you had to implement these algorithms from scratch. It’s just linear algerbra and calculus, which anyone with an engineering degree is going to at least have exposure to. I couldn’t get the time of day from anyone. Not even a call back to prove that I knew or could learn what was needed to be effective. It was incredibly frustrating and disappointing.
Data science / machine learning is not that hard, and you are turning away good candidates for bullshit reasons. Stop it. At least bring them in and talk to them. Jesus.
Even when a customer's given me a "clean" dataset, I've had to write 400+ lines of code to do the feature engineering on a relatively straightforward logistic regression. Then there's all the other times when a customer asks me to deploy one type of algorithm, and their business problem is actually solved by an entirely different class of algorithm entirely.
Zayd over at Stanford has a nice blog post [1] describing why machine learning is several more dimensions of complexity compared to traditional software development. There _is_ a specific set of Data-first skills that is complemented by dev and CS experience, but a fundamental reason why ML projects fail is due to lack of appreciation for the many different skillsets needed to succeed.
[1] http://ai.stanford.edu/~zayd/why-is-machine-learning-hard.ht...
I've found encapsulating the data preprocessing steps in pure functions helps ensure that the data cleaning can easily/quickly be debugged. When it comes to the actual model, there is no substitute for thoroughly understanding the characteristics of your dataset. Finally when it comes to model selection, a good scoring metric is absolutely necessary; this is entirely dependent on what you're actually trying to accomplish with the model. So there is little universal advice.
When it comes to the long iteration cycle, the only bandaid I've been able to find is a solid test suite, and thorough code review. This makes it less likely you'll introduce unintended problems. Basically you have to move slow and deliberately, instead of "move fast and break things."
Far from impossible, but certainly more difficult than "traditional" software engineering. It's like transitioning from debugging interpreter tracebacks generated by a simple toy script, to debugging a >40k LOC application written in a dynamic language, which happens to be intermittently segfaulting in production.
That's the actual job! The theoretical investigation. Not the programming. The programming is really easy. which is part of why having 15 years in software doesn't have so much weight.
The job in which you are magically given a pristine dataset and can just tweak ML models all day doesn’t exist.
It's a bit of both really. If the data science job is in the field of predictive maintenance for example, then a (relatively) simple model may be sufficient to add business value straight away, and the hard part is a deep understanding of the potential failure modes of the machinery you're predicting on and the kinds of sensors used to gather the raw data. There's no "one size fits all".
This is one of my favourite papers on the subject, written by one of the lecturers on the DS course I did: https://users.cs.duke.edu/~cynthia/docs/RudinETAL2010.pdf
Whoever is running data scientist units needs to realise you need IT/DBA/programming people in the mix. Statisticians in my experience cannot design databases, concepts of code reuse and normalisation do not exist in their vocabulary. A great deal of time is wasted doing repetitive tasks that someone with a programming background would have assisted.
My suggestion is to get a job somewhere with a data science team, do a data science side project with company data, and demo it to the data scientists. That is probably your best bet to get a foot in the door.
Sounds like you may already be too good at data science to ever get any actual experience of it.
In fact, having a PhD does not tell you a whole lot about these skills.
More advanced modern techniques such as deep neural networks, reinforcement learning, etc.. are all extremely proficient at certain niche problems but these do not come up nearly as often in a business context.
This is why I don't advertise myself as a machine learning engineer. Rather, I am a business consultant that knows when, and when not to utilize machine learning methods.
Unless your DS group is big enough to warrant hiring data engineers / cleaners, as a data scientist it'll be your job not only to eventually choose the algorithm, but foremost, to confirm that the data is sufficient to serve the intended purpose of mining it, ideally before you waste a lot of time curating it or paying for a raw data dump you can't use.
So basically, people like you are a huge risk for a company. You really can't prove you have anything more than a practical ability to use the algorithms. Most algorithms are hard to implement from scratch (robustly, let alone correctly), and you would be doing so by reading papers you have a scant understanding of..
I'd bet a large sum of money 2/3 of the candidates applying for a machine learning role without a PhD could not even provide any derivation of OLS or something like that if it came up in an interview.
Incoming downvote train, choo choo
Graduates get mad because they realize they've wasted a bunch of money on something they can learn for free, so they reject the notion and make others go through the same abusive system.
Hn readers are very literal on the whole.
As a dev, you can pick up some popular ML frameworks and learn the basics relatively quickly. The difficulty in this field comes from the amount of theoretical knowledge you need to interpret your results.
All these data science bootcamps/learn quick schemes are like teaching a blind person to drive a racecar. He can work the pedals and the steering wheel - but has no idea where he's going.
I think this is a problem for candidates who don't meet the set requirements, but this can also be an opportunity to differentiate from rest of the pack.
You want a better employee than your competitors have. You'll take all the expertise you can get, plus the ability to invent new techniques as needed. Everybody wants the very best 'cause that last 1% of skill can still be worth a lot of money. Everybody wants someone who can hit a lot of long balls for their team, as it were. There might be a floor on the skill level you could make some use of, but there's no ceiling.
My experience has been that this doesn't matter in a lot of cases. A product can be both, highly useful and subtly broken.
One of the biggest boons to my productivity was realizing when something is good enough and to move on. You can waste so much time tuning for precision and recall.
I think this is an essential, foundational skill for modern data science, not unlike some knowledge of R or Python, but an order of magnitude or two harder to learn. A lot of the difficulty is in lack of good educational material. I also think this is significantly more important than a deep understanding of various modeling approaches; you can learn that as you go, so long as you know if what you have now is "good enough" or if you need a fancier approach.
Unfortunately you can't just throw k-fold CV around, average some fit metrics, and call it a day.
Tech worker's arrogance reminds me of "made in the USA" factory workers who thought their jobs were safe. If people don't start pushing back wages will be pushed down to global equilibrium. Based on how things are going I'm predicting neo-feudalism thanks to the joys of open-boarders globalism putting all power into the hands of corporations.
Yes if it's real. No if it's fake.
This still ignores the question of if entry limits for jobs are set reasonably or not though, you can create a real shortage by limiting yourself needlessly. At which point those who think that it is needless will call it "fake", so which word you use also depends on opinion and interpretation of data :-)
Also, you create a fake shortage by driving wages down: "we have so many positions that we cannot fill (at the price that we are willing to pay)"
I can do everything you can do.
I don’t have an advanced degree, just a bachelors of math from 1995, and have been breadboarding (and more), and coding since the 80s
I can follow along with ML and have implemented toys with the ML algorithms in a couple days.
It’s bourgeois intellectualism. Like a law firm only hiring from Harvard
ML is automated schema design. And the current methodology has known limits of applicability
This is “Mongo DB”, “devops” like hype all over again.
The problem is that there are hundreds of applicants in your situation WITHOUT experience. There are usually a couple of PHD or MS applications WITH experience for every job. Who do you think the company would give preference to?
Companies are using a PhD as a proxy for having research experience because it's the the only qualification like it out there. It's a poor proxy because not all PhDs are created equal.
This is missing my pet step: doing the literature review.
I’m pretty ambidextrous when it comes to Python and R, so I’m not typically a combatant in the data science language flamewars.
But... for as much as the Python community likes to assert their superior coding chops, I’ve observed that the R community does a much better job of reading about prior art.
One of the earliest, most important, and most useful lessons I learned from a senior grad student: "a day in the library can be worth a week at the bench."
This is a grossly defensive overreaction to the parent reply.
All this hype makes it so everyone wants to be a data scientist. You get people who change careers to go into this new hot career. You also have a pool of people who have been working with data well before the hype with experience. The people trying to break into data science will have a very hard time competing with the people with experience over the pool of jobs out there.
So this sucks if you don't fit the model well in a way that has you often end up as a false negative - but that doesn't' mean the model is broken.
You are claiming there is a generalization problem that causes extra error in practice. Another perfectly viable hypothesis is that the classifier is working fine, it's just tuned for true positive rate and accepts a higher false negative rate to get it. Specificity vs. sensitivity is a fundamental trade off, not a training issue (though that can make both worse)
If you're serious about getting a PhD-level job without a PhD, getting someone to recommend and vouch for you is even more important than usual. Since you're up to date with papers, why not email researchers you admire with questions that demonstrate you deeply understand their work? Many will be too busy, but some will probably be impressed by your determination. Once you have a relationship, see if you can assist with their research, even if initially it's just grunt work. It will take time, but integrating yourself into the academic "web of trust" and maybe getting your name on some papers is the only plausible way you can expect a company that doesn't know you to take you seriously.
There's always a domain specificity that sometimes comes from grad school, but data is data and industries are filled with SMEs who understand the domain.
Source: I hire data scientists for Fortune 500 companies.
>I can do everything you can do.
To demonstrate that to an employer wanting PhD workers, go get a PhD like the other PhDs did. Claiming you can do what they do when you haven't done what they have done is not going to cut it.
>you won’t even call me to talk to me that means you are missing out, and you are gatekeeping
Gatekeeping = not spending unnecessary money and wasting unnecessary time.
An employer saves significant money and time by not having to interview everyone claiming they can do what PhD can do but didn't bother to get one. Your skilled workers don't have to stop producing and do interviews, your HR people don't need to spend time and money booking flights, hotels, and such for candidates. You don't have to work through 500 resumes with 30 PhDs in the pool - you sift through 30 resumes.
When these jobs are hot and candidates are plentiful, using signals to narrow the field to a group you can more rigorously interview is typically a more effective use of time than buying into everyone's self-belief. Candidly, I find most individuals from a programming background vastly underestimate the skillset required in this space. I know that is not an uncommon perception and you are likely being penalized for it, fairly or unfairly. I'll say this: anyone who refers to data science as "just linear algebra and calculus" would be immediately removed from any candidate pool I was managing.
Others have evidence of capability, you do not. Programming experience is not evidence enough to elevate you above candidates with more reliable and relevant credentials. A shelf full of books is not evidence either. You either need to find a version of this job created by people that don't really know what it is they want (hint: if the Data Science JD says "Excel" that's an indicator, it's not too uncommon) to create a work history, find a way create a portfolio that you can use as evidence (e.g., Kaggle competitions, hobbyist projects with available datasets), or network with others in the industry and academia such that they will vouch for you.
Sorry but how would you even know?
I understand your frustrations well since I don't have a BS.
Still, as much as you seem to want to talk about how capable you are, you can't seem to understand the perspective of employers. Given what you've said about your history I suspect this is not due to a lack of intelligence but empathy.
Employers have to go through many candidates, each of which has some true capability but of which the employer can only see some signals. Signals have varying degrees of quality, and interviewing candidates costs time and money. That being the case, it is only natural that they try to use the strongest signals they have.
Nobody believes that there are not capable people who do not have an MS or PhD as you seem to be suggesting. The reality is that the proportion of people who have a BS and can do the job is much less than the number with MS or PhDs, and so it's one of the more effective filters they have at their limited disposal.
I'm sure companies are not happy about skipping great candidates like yourself, but they have not figured out a way to do so that is scalable and cost efficient. It's a difficult problem but maybe you can figure it out.
> You are not special because you have a PhD and I don’t. The only difference between you and me is that I had the ability to learn for free what you paid for.
How much do you even know about PhD programs? It seems like not much because PhD candidates, at least in the US, get paid. It's not much but they certainly are not paying for their education.
I apologize for this unsolicited advice but your lack of humility is frankly very off putting. You sound like a very hard working person. PhD programs are extremely difficult to both be admitted into and to finish. I'd expect that you would respect others like yourself who are very hard working.
The days of a researcher producing a model to be re-implemented for production by a programmer are over, or very nearly so. A working data scientist now is expected to produce something that can run in production. That’s something a PhD doesn’t teach and that many PhDs find an uphill struggle.
Right now doing this is somewhere between “state of the art” and “new normal” depending on where you sit.
Is Data Science really fundamentally different? Or is this PHD who barely programs going to either do tasks like cleanup terrabytes of data or risk that a coder with no idea will introduce a bias in the data during that process?
I find the whole emergence of the field fascinating, but I kind of feel like it is just techs recreating actuaries with what is actually a less specific education.
+ (And a worse academic style career path of going through an education you may never get to use instead of going up from apprentice to master)
You have to eliminate a lot of possibilities in a universe with this much detail
The real issue, again IMO, is more of an “expectation crisis”. We expected to repeat this and failed to. Because we’re still a ridiculously ignorant species lacking conscious awareness of many aspects of reality
If you don't reward (By issuing grants for) negative results, or null results, or replicating prior studies, why do you expect that scientists will aim for any of those outcomes?
We have developed tools to deliberately abstract away the complexity of the underlying math, and they work well. They work exceedingly well. I once took a semester-long class in data science where untrained, mathphobe business students were running various kinds of regression models on cleaned up data in WEKA (poor choice of software, yes) by the end of it.
Most data scientists can and should treat the algorithms themselves as black boxes the same way software engineers treat R-B BSTs as black boxes, without necessarily having to know how they work off the top of their heads.
I once worked with a Stanford "data scientist" at a top tech company who couldn't immediately recall Bayes' Rule. He wasn't stupid. His knowledge was just structured in a way that reflected the reality of his day-to-day; this is how it works for all of us.
The primary skills needed are: munging, featurization, analysis (basic stats and then a few other things like ROC, etc), and perhaps most importantly (and the thing I see PhDs in particular chronically fail at) operationalization. You do not need to know heavy math to run a model over data, which is the maximum level of sophistication required for most applications of data science that generate real business value.
I see similar foolishness in data science as I do blockchain, to be quite honest: people hype up and gravitate towards the cutting edge, while forgetting, ignoring, or blatantly obscuring the power of simple math on big data. I guess it's important that people have an inflated view of the complexity of what most data scientists are doing because having the role at least somewhat cloaked in mystique boosts salaries in the long run.
Any argument that a reasonably intelligent person can't be trained to be a legitimately effective data scientist is a counterproductive lie.
The field is absolutely saturated with people who want to be a data scientist but have no experience. This is where some of that gate keeping comes from.
The people who have it easy are the ones with a MS or PHD and years of experience doing data science work at companies under their belt. There are very few of these people right now.
There is this idea that data science is needed everywhere and there is a HUGE supply of jobs. As an example if you search for data scientist jobs at glassdoor.com in San Francisco there are ~2000 jobs. If you search for software engineer in San Francisco there are ~9000. Similar ratios can be found in any major tech city. Data science does not scale like software engineering in companies but the narrative out there is that this is the job to be in and there is this huge unmet need. It is all hype.
What are the data science "gotchas?" A lot of people can pick up basic programming in a weekend, but they wouldn't necessarily know what they don't know and might well get deeply mired in problems with concurrency or algorithmic complexity.
It's such "gotchas" which justify gatekeeping. Otherwise, gatekeeping is just unproductive manipulation of the market.
There are a lot of people out there who have memorized how to implement k-means and PCA but would absolutely struggle to understand what they're actually doing or interpret the results in a meaningful way. A HUGE part of being a data scientist is presenting information in a useful way. That's why PhDs are favored because with their experience having to write grants to get funding for their research they're exactly the type of people that can take a naive problem, work it to a result, and then sit in front of a board room of non-technical people and explain why their result was worth the money that was given to them.
My wife has a PhD in comparative lit, but she now works in banking. Her PhD gave her a superpower: Reading. She can quickly absorb large amounts of text with abstruse, complex, and subtle distinctions and tell you very nitpicky things about it. She thinks financial/banking regulations are light reading in comparison to the stuff she waded through to get her PhD. (It's also quite surprising: the number of C-level people in banking who have little patience for reading. She's won a number of boardroom battles because she has actually read things.)
what exactly do you mean by literary analysis here? i have (in my opinion) extremely good reading comprehension skills, in that i can read and understand the literal meaning of almost any text (provided i understand the context), and i got an 800 on the critical reading section of the SAT. on the other hand, i can't for the life of me read a book and pick out any of the major themes without having them spoon-fed to me. i was always terrified when i was expected to have my own opinion about a text to use as the topic for a paper.
1) Data engineering. I suck at this, don't ask me.
2) Inference. One big gotcha is often of the form of not accounting for all the sources of variation in your estimator and thinking you have something when you don't (often coming from unaccounted sources of correlation in time or space or repeated measures). Another is that correlation isn't causation. This pops up in surprising ways. Or things not being as independent as you thought.
3) Prediction/classification. Gotchas take as many forms as the things you look at, but the birds eye view is that you apply a method and it works ok, but either not well enough, or you then try it in the real world and it doesn't generalize as well as it did on your test set. The ways models break down depend heavily on the model and the data, so the way to diagnose and fix the issue depends on both understanding your toolkit really well and understanding the context of the data (business logic, etc). Another gotcha is in understanding uncertainties of your predictions. If I predict that this word is a noun, how sure am I of that? Many beginners skip those kinds of questions, but don't realize it.
I'm a data scientist with (barely) a bachelor's in physics working with mostly PhD's and, while the academic degree based gatekeeping is bad and frustrates the shit out of me, I get why it's there. The learning investment to learn the basics is dwarfed by the learning investment to be able to flexibly apply the right things at the right times and tweak/fix them as appropriate.
Of course we can abstract the root argument; for a given job, among those qualified to fill that job, there exists at least one person who has auto-learned the skills required to perform the job. This is probably true.
For data science the big one is over fitting, which everyone talks about, but can happen in really insidious ways in production. You have to be very disciplined and careful with the data to prevent over fitting.
Another big one is productionizing data science, which in my opinion most data scientists don't have a ton of experience with.
The actual training of the models part of data science isn't that hard, its actually making it work with the crappy data that exists in the real world and putting it into production that are the really hard parts.
If I'm reading between the lines of your comment correctly, I don't think that's what you mean, I think you mean that this was a problematic experience. If that's right, I don't really understand that reaction: this is a nascent field, the ready-made experienced work-force is expected to be much smaller than the demand, with most positions filled by new entrants gaining experience rather than folks who already have it. This is the fundamental challenge and simultaneously the great opportunity of running a business in a nascent field!
Are there enough of those people to meet the demand, even recognizing that it is smaller than the overall software engineering demand? If not, it seems like the parent's point still holds.
Edit: But thanks for the perspective on the size of the market and level of hype, that's a useful data point.
Then you go off on a jag about how simple the math is and how easy to implement the algorithms are. Totally true, but this isn't what data scientists are paid for. Data science/machine learning positions are more about understanding the limitations and pathologies of the algorithms, the data, and their interactions. This isn't necessarily hard, but it can be -- and the pay tends to scale in proportion to the difficulty. Since theory doesn't provide much guidance -- you can learn everything that anyone knows in a year -- employers will necessarily prefer someone who can demonstrate practical experience with data. Selecting for PhDs is one way of filtering for that.
If you've got experience with data, emphasize that and you should get callbacks.
My partner has a successful career in data science, but she has multiple degrees in mathematics - significantly more than your typical engineer is exposed to.
My training heavily stressed bias and confounding, study design, problem specification, validation, and interpretation of results. I've seen a lot of software engineers dabbling in machine learning jump to training a deep learning model for a problem where regex would suffice. I've also seen multiple people build models that reflect the data collection instead of the biologic/medical process, and present it without even realizing how wrong their results are. The problems in these cases is often their "objective" measures of performance (e.g., precision, recall, accuracy, AUC, whatever) look pretty good, but they don't see that it's because the model AND data are both heavily influenced by this larger problem. For example, is an increase in complexity of a particular disease due to people actually being sicker or because some payer changed a reimbursement program so now billing departments are using higher acuity diagnosis codes for their patients?
That said, the best engineer I've ever met didn't have a college education. I know a bunch of awesome data scientists who have taken pretty circuitous journeys to their current career. So, "PhD required" seems like it'll lead to a lot of false negatives. So, acknowledging that, my main point is that doing data science in a meaningful and ethical way—particularly when it involves human subjects—requires a lot more thought than just being able to implement some machine learning algorithm.
This is great. It's something I notice I have to mentor my junior data scientists and ML practitioners on with some regularity. Data science and machine learning aren't necessarily useful without some additional domain knowledge and discipline to recognize the need to view the data from many different perspectives and with many different relationships highlighted. Too often they see high P/R numbers and call the job done, or something similar.
I think a nice team can have some balance, with enough overlap that we can find common ground. I'd be uncomfortable without knowing I can rely on some of my peers on some of the heavier algorithmic or SDE stuff.
It's not particular to ML, though ML is worse than others. CS degrees are luckily becoming less needed for basic programming jobs.
Why should someone bring you in? Do you have a kaggle profile with your notebooks that people can look at? Have you replicated research papers ? Do you have any independent results that you can share, publish, or talk about?
I actually think that requiring an advanced degree doesn't help toward that end.
In the finance world we call these people quants. And actually in my experience having a phd does very little for someone; the critical skill required is software engineering.
Thinking about investment strategies is about testing hypotheses, and you can't test things properly if you don't know a few things about how to organize code. This is actually an insidious problem, because there's nobody telling you how to actually build an alpha generating strategy. And if you can't investigate properly, you fall into all the traps (you make excuses to do the following): too many features, choosing too small a sample that happens to do well, filtering the data in so may ways one of them is bound to "work", and so on.
Imagine that you're a chef, but you can't chop stuff effectively. You would then work around that limitation, maybe work on dishes where it's not needed, or perhaps get a junior guy to do the chopping. You might think this solves the problem, but actually it just swaps one problem for another, because now you need to communicate and coordinate with this other person. Or you don't explore that whole area of food with chopped stuff in it.
CI pipelines, version control (branching, diffs, etc), database maintenance, a bit of OS basics. All things that tended to differentiate the productive quants from those who merely thought they were useful. I've seen things done in a few weeks that others had spent years not achieving.
As for the ML skills themselves, you do need a bit of math to do it, and the math is relatively easy to learn. There's loads of materials to help you as well. What's not explained so much is certain philosophical issues around what is being examined. A course in economics has examples of these things: endogeneity, Lucas critique (which is Hume rehashed), experiment design (do the observations mean what you think?).
1. A little bit of the selfish "oh no, the secret's out, at what point is my salary going to drop when the demand is met by the dedicated Master's degrees and bootcamps?"
and
2. These articles seem so incredibly corny, it's almost embarrassing. The "hottest job"? Ahhhh, stop it. But these things go in an out of phase, similar to back in the day when "anesthesiologist assistants" (CRNAs, AAs) were the hottest thing for Bloomberg to talk about. It will not last forever.
The irony is that I probably only knew "data science" (always in quotes) existed because I read one of these cheesy articles. I mean, we all know that statistics have been around forever, but that there were dedicated positions where you could run stats, build models, and then deploy them all in a single role was foreign to me.
So it's a combination of a potentially irrational fear of self-preservation, and laughing at the state of affairs where some basic stats work will pull in that kind of money.
I tend to have fears about the future, always wanting to hedge myself so I don't become outdated. In the data science sense, I see the field becoming super super broad and eventually saturated with new supply, so I debate on whether I should pivot into management of analytics in general or not. I.E. getting my hands off the keyboard. Ultimate goal would be to help define, strategically, how statistics/data mining/machine learning/yada/yada/yada are used at a company.
A boot camp can easily teach someone to, say, estimate a linear model or run k-means. I dread the future when the industry decides the right way to put up barriers is by creating ever less-realistic interview loops that are even more coin-flippier, dice-rollier, card-shufflier.
Imagine you're hiring someone to build a house for you. Would you feel comfortable with someone who's just been drilled on how to use individual tools? I would want someone who had been taught a step by step process for how to put together a house.
Very true. On the other hand, it's a pleasant rarity when I see positions that appear to index more heavily on, "how well is this person able to conceptualize the problem and choose an appropriate method?" than "can this person do X?"
Lots of folks can do X; fewer can conceptualize a research question and choose the appropriate X; even fewer can carry out the X and communicate robustly what it means.
The latter two start to get into squishy territory, but also are where the value is. They also seem to get the least focus in advertising / recruiting / interviewing data scientists.
It reminds me of studying evaluation methods in planning. One that people are really familiar with (at least anecdotally) is cost-benefit analysis. Conceptually, it's very simple. The problem is that the costs and benefits that are hardest to measure are very often NOT measured. And they're very often the sorts of things that people find the most important. So, you end up with an answer that encodes a ratio of easily measured things rather than important things.
So too with data science. Easier to check whether someone can remember basic probability rules and carry out a linear regression than it is to diagnose whether someone can reason carefully about an amorphous business problem.
Big data is like teenage sex: everyone talks about it, nobody really knows how to do it, everyone thinks everyone else is doing it, so everyone claims they are doing it...
Tim Hopper https://twitter.com/tdhopper/status/916383020835368960
I think a lot of companies are hiring "data engineers" but don't know it. The guy who started making Superset (https://github.com/apache/incubator-superset) wrote this and I think it's apt.
https://medium.freecodecamp.org/the-rise-of-the-data-enginee...
It's the HR equivalent of deciding to rebuild your solidly performing website with some hot new v0.1 frontend javascript framework purely because it's the new hot stuff. Perpetuated and reinforced by C-Suite level desires to appear trendy and cutting edge.
When I ran an Analytics/BI team, I was well aware of what I needed and wanted. But constantly had to fight to not have HR label open positions as Data Science roles.
Once I got them to start posting the position with Data Analyst and Data Strategist titles instead, the team satisfaction scores and attrition rate markedly improved. The (very competitive) pay rate and actual work was the same as it was under Data Scientist titles. But it better scoped the candidate pool and aligned expectations up front, rather than hiring a bright-eyed and overqualified Masters/PhD graduate and being the company that disillusions them to the reality of most "data science" roles.
Although I can understand the desire to ride the hype train for those executives. Being pragmatic improved my team and benefited my company, but cost me the ability to add Data Science as a buzzword to my resume, and "Created, managed, and scaled a Data Management and Operational Analytics/BI organization" doesn't have as much market value.
In my case, I'm a reasonably solid R / Python programmer (who occasionally dabbles in racket, clojure, and others). I've got the sort of applied statistical training of someone who took a quant-heavy course load in a PhD program. I've even (by title) been a data scientist a couple times now.
Having recently decided to reenter the job market, I'm reminded that finding the RIGHT data science role is a major challenge. When someone wants a data scientist, they may be looking for someone with a lot of specialist depth in operations research, financial forecasting, machine learning, data-focused software engineering, or some other not-at-all-universal area of expertise.
In some ways, I almost think that someone with a bootcamp level understanding of stats may be at an advantage. Whereas I'm very inclined to be, "Oh! You want this other kind of person. Would you like me to put you in touch with one?" I think someone more junior is inclined to be "How hard can [X] be?"
You don't sound like the Dunning-Kruger sort, so I'd chase the fun sounding problems in organizations where you can use more senior people that you respect as sounding boards / mentors.
Good luck (to us all)!
Beyond that, there are labels that refer to specific technologies that get you even closer: Rails, Elixir, Django, React, QT, C#/Winforms, Verilog.
There really aren’t the equivalent shorthand filters in data science.
There are, though, right? The catch-all "Data Scientist" title is being split into Data Engineer, Machine Learning Engineer, Data/BI Analyst, etc. Each of those, in my experience, have clear definitions.
targeting health-care customers for hospitals
people who can turn social-media clicks and user-posted
photos into monetizable binary code is among the biggest
challenges facing U.S. industry
“sentiment analysis,” or finding a way to quantify how
many tweets are trashing your company or praising it
determine how customers prioritized paying bills
“recommendation engines” - those programs that predict
what you may want to buy next
advertising
My background is traditional business intelligence, finding actionable data for high level leaders.A common response on this thread is that little data science is actually occurring in the business world. It would be incredibly useful to me (and I suspect other readers of HN), to hear from other participants on the thread what data science and methods they are using.
I'll kick it off with data analysis examples from my workplace:
1 - Analysis of patient accruals to various clinical trials.
Mainly tracking against goal numbers. No statistics or ML.
2 - Analysis of tissue collection opportunities to answer the question:
Are we obtaining samples useful to future research opportunities? No statistics or ML.
3 - Creating models to accurately predict patient accrual rates for individual studies across various
different variables (race, ethnicity, gender, age).
Simple statistics, probably just a linear regression model. This is a new effort.
4 - a long, on-going, and currently unsuccessful attempt to extract useful data from pathology
reports (free text descriptions from pathologists examining various collected tissues for
both medical treatment and research purposes)
In addition, I know of a few NLP motivated efforts to train classifiers - say, given a set of 500 manually labeled papers, can a classifier be built that would be effective in bucketing an additional 8,000 papers?Am I right?
Despite data science being a hot job, the sheer, growing number of MOOCs available will cause the high amount gatekeeping from many prospective employers to get even worse.
I am very happy as a data scientist now, and yes, it's more complicated than doing Excel VLOOKUPs! Although I did get rejected from those VLOOKUP positions many times in my job search...
I did a MOOC, nothing wrong with them. A lot of it turned out to be just revising things I had originally learned in my BEng or MSc, but I learned some things too. And a bit of revision never hurt. I don’t think I had done much calculus “in anger” since graduating for example, so that was rusty.
In 5 years or 10 years there will be no data scientists as a full time job - the skills will just be folded into regular programming jobs or accounting jobs or whatever. If you’re in that job now, make hay while the sun shines
Oh and it will all run on JS.
-everyone in the 90s
If you want data scientists, pay a good salary and good equity.
Kind of disgusted though, that Equifax "is shortening the hiring process to keep anyone from slipping away." They could fix their security practices before hiring anyone with Scientist on their resume.
"A data scientist is someone who knows more about statistics than your average software developer, but more about software development than your average statistician".
"Is America's current fad job, like tons of similar jobs before it" (have lived through 4-5 of those "hot jobs").
In the 90s/early 00s there was the biology/bio-engineering/bionformatics trend -- everybody was thinking of going into biology when I went to university -- that dried up soon as well [1]
[1] https://iubmb.onlinelibrary.wiley.com/doi/pdf/10.1002/bmb.20...
Here's a good list written by another HN member:
I have been flooded with candidates for the data science position. There seems to be a glut of well qualified [1] data scientist applicants. I have turned down many candidates who I think would also have done well in the position. The CRM developer position has been extremely difficult to fill. This has had the effect of bringing the salary of those two positions in our company to near equivalence. My initial expectations for the salaries were that the data scientist would be paid substantially more than the CRM developer. I know that Sugar/SuiteCRM aren't super popular but it also could be that I have been effected by all the attention that ML and AI have gotten over recent years.
[1] well qualified for us has generally meant masters level or above with real work experience. maybe this definition is the source of the 'glut' of applicants?
People compete in machine learning for fun (e.g. kaggle.com). I doubt anyone works on CRM systems for fun.
I hope the tide changes here and I think it will, but this has been my experience so far.
Those kinds of jobs are advertised as “data science” jobs. In fact any job that involves data manipulation at all is getting re-branded as “data science”. Partly this is employers doing a bait-and-switch to attract better candidates and partly its just jumping on the bandwagon.
But candidates aren’t entirely innocent either. How many programmers - an honest title for honest work - call themselves Senior Certified Enterprise Solution Architect Team Thought Leaders or some such nonsense? We’ve all met a few. LinkedIn is crawling with them.
I see this is as the greatest threat to the demand of the "in house" data scientist.
If this turns out to be the case, I see the greatest demand for those who can write production grade code (i.e. software engineers) and those who are effectively trained data scientists. We see this job often called research scientist or research engineer.
FTFY.
I'd even argue that - for data science - it's more important than any ammout of years doing engineering.
It appears to me that about two or three people were hired that way over more than a year.. at the same time, the company was being sold to a larger company in Seattle.. the founders made at least seven figures in the sale, I would guess... draw your own conclusions..
Let me tell ya'll all a little secret. Execs have been mismanaging infrastructure... and it's all so close to crumbling at the first puf-o-wind.