For example, if you have a bachelors degree from a top engineering school (MIT, Cal Tech, Stanford, Berkeley, etc.) you have proven that you are intelligent and can work hard.
People without a masters degree, but more business experience, bring a different perspective, and are often more business results focused, and potentially work more collaboratively than an individual who just graduated from a masters program.
Source: I am a Data Science hiring manager, and have interviewed 100+ candidates at several companies
I think the commenter’s point is (implicitly) about specific, directly relevant Master’s degrees. Obviously a general Master’s wouldn’t provide much of an advantage. The difficulty isn’t demonstrating intelligence and work ethic, it’s demonstrating targeted expertise.
> People without a masters degree, but more business experience, bring a different perspective, and are often more business results focused, and potentially work more collaboratively than an individual who just graduated from a masters program.
To be honest with you, this sounds to me like complete speculation. I’m not saying it’s wrong; rather it seems like it’s at best unempirical, and at worst unfalsifiable. The qualifiers you’re using (like “potentially”, or “often”) don’t seem like strong heuristics.
I think it would be helpful to discuss straightforward job descriptions. For most real data science roles, I would not weight any of what you’ve listed (except collaboration) as being remotely as useful as demonstrable expertise in computer science and statistics. For candidates without a Master’s degree, I wouldn’t take business experience or lack thereof as a signal whatsoever - I’d look for a relevant heuristic to replace it.
I've had questions ranging from reversing strings on a whiteboard to checking for valid email addresses. I had another question about flipping biased coins and calculating probabilities. It's all nonsense and totally unrelated to the skills I developed during my PhD which primarily consisted of performing massive amounts of machine learning on high performance computing systems over large sets of data to extract important insights.
But — if solving these algorithm puzzles quickly and without errors is the key to a $300k+ job, so be it. I'll just practice this nonsense until I've optimized for the skill of "interviewing", and then maybe I can contribute in some kind of meaningful way to the company with actual data science.
Is there a consensus about what kind of Master's would be most useful data-sciency stuff? Computer science? Stats?
I think the actual term is Data Engineer.
It’s why technical interviews can be so brutal, unfortunately. There are a lot of frauds out there. Money attracts frauds.
What’s the fizzbuzz test for data scientists anyway?
My phone screen "fizzbuzz" is having them calculate a standard deviation from an array of data w/out with only basic operators (no numpy.std). Then explain why they choose population/sample and explain the difference.
I studied math in undergrad so one of my requirements is "knows more math than me".
What kind of questions are you asking to ensure that they’re correct when they’re speaking about math you don’t know?
1) People who are great at the mathematics behind the statistical tooling
2) People who are great at conceptualizing a relevant question, operationalizing it, and then using a computer to apply appropriate models.
I think in most cases, for businesses needing to solve business problems, the latter kind is probably more useful. There are applications where the former is required, but you probably know if you need this kind of data scientist.
I should also add that these traits aren't mutually exclusive, but that individual data scientists typically are stronger or weaker along approximately those axes.
In general, I still dislike the term "data science" because it obfuscates meaningful distinctions between math nerds, computer science nerds, and research nerds who happen to do some applied stats.
I do, however, think anyone with some lick of statistics background should know the formula for a standard deviation. Considering how fundamental the idea of variance is in statistics.
Also, for our role, we're specifically hiring someone with extensive stats background since a large part of the role is learning domain-specific statistics of the industry we're targeting and figuring out how we can adopt those models with our data.
It’s a filter that theoretically allows false positives (which is why you continue with other questions), but it really shouldn’t have any false negatives.
This is just my n=1 opinion, but this is a terrible test for data science skills. I've had to calculate standard deviation by hand many times in my life, but my short term memory is such that despite doing that dozens of times over the past two decades, I still can't recall the formula off the top of my head. And then there's the whole n vs (n-1) thing in the denominator which has something to do with degrees of freedom, but I would just Google that as soon as I needed to know (depending on exactly what I was trying to do with the data).
So I don't understand how your question in any way tests someone's skills at analyzing data to extract valuable business insights. At best, it tests someone's ability to memorize formulas and minutiae (although I'll grant you that understanding the difference between a sample and the population is important).
Personally, I think take-home interviews with real data sets are the best way to gauge a candidate's skills. You're actually testing them with a work sample, and they are not under artificial time or memorization constraints.
Read through, and do all the exercises in, one textbook each for:
1. Calculus
2. Linear Algebra
3. Abstract Algebra
4. Analysis
5. Topology
6. Probability Theory
7. Number Theory
...more or less in that order. Make sure your calculus book covers single variable and multivariable calculus. Supplement with applied mathematical statistics. Do that, and you have the equivalent of a mathematics undergrad (as far as relevant courses are concerned).
You could even do this with something like UIllinois’ NetMath program, or some courses on Coursera. You can swap out Number Theory for Complex Analysis or deeper Probability Theory and it’d be more relevant.
Skip abstract algebra, topology and analysis. If you find yourself in the same room as a number theory book, walk away slowly without making eye contact lest it cast a spell on you.
Sure, skip number theory. Like I said, you could swap that out.
One can learn the necessary topology, analysis (etc.) in the relevant places (and the relevant depths) that they come up.
1. The context is knowing more math than someone who has an undergraduate degree in it,
2. Abstract algebra is part of such a degree, and contributes significantly to overall mathematical maturity, and
3. You can avoid some subjects in the short term, but in the long term you can’t progress further without a reasonable mastery of algebra and analysis.
Probability theory and linear algebra are heavily used in data science. You won’t be as competitive a candidate for a job if you don’t have a firm grasp of both subjects. At a certain point, linear algebra ceases to be distinct from abstract algebra, and those exercises you were doing become applicable to real world results.
But, what you're describing I would consider "data engineering" (at least how I have been hired to do it). Working through the business problems and pragmatically facilitating data, pipelines, databases, and models to solve those problems. It's less established and less "hot" but, IMO, it's a much more valuable job to most businesses.
Much like how the early hires at Twitter were not deeply experienced in high availability work -- segregating the architecture of a predominantly RoR code base to be resilient at scale, which lead to countless "fail whale" outages, before they eventually landed someone who helped them re-think their architecture to use RoR for what it's good at while introducing the JVM and other languages to handle other aspects of their workload.
I'm having a little trouble trying to parse that sentence. Could you explain it better?
Based on what I think is being asked, the question is essentially: What is a STD? I think this is a very straightforward and fair question.
For less Stat-y HNers: For normally distributed data, the STD is the root of the Variance. The Variance is just the average of the square of the difference between the data points to the mean. Essentially: Take a point, find the distance to the mean, square that, average over all points you've done that to. That's the variance. Root the variance, that's the STD.