Mathematicians becoming data scientists
quomodocumque.wordpress.com
quomodocumque.wordpress.com
Industry can cheerfully, usefully absorb thousands of solidly "average" (in this particular sense) mathematicians (or similar) in a way that academia just has no plan for. Even if they do not work on anything quite as technically interesting as they had previously in other ways the job may be more rewarding; And I don't mean simply financially.
That being said, most academics are not a good fit initially.
That's very different.
No one can keep up with the flood of JS frameworks, either.
Information fire hoses are very different from extreme differences in ability.
I'd suggest figuring out which niche of data science most aligns with your personal interests and start there. Hiring managers don't discriminate based on "too many years in academia", it's more what's the reason they should hire YOU over candidate y?
Diehard animal lover? I bet the Sierra Club would love a statistician who could help them quantify how many elephants are being poached in Africa per year.
Interested in sports? Strava gets millions of data points a day, I'm sure they're looking for help finding signals in the noise.
Can't get enough of financial markets? Finance jobs all over the place for quants.
If your profile is too generic they'll see no good reason to hire you, but if you're passionate about X it can really work in your favor.
The job search is depressing but never give up hope.
But that brings me back to my original point, you have to find a place where your personality and interests are a match, not just RandomJob.com.
If you have a higher degree in math, there is always another job out there waiting for you.
Typical interview for higher jobs aren't focused on tests or quizzes. If that's what your getting your not applying to the right places - and there might not be many places.
It could be true in some fields (biotech and life sciences, pharma/chemistry, "real" engineering, ...) but in math/CS?
Most companies hire PhD grads at one job grade/level above entry-level, i.e. it's typically not worth much more than a ~2 years head start salary-wise. If a candidate has spent 4/5 years gaining experience in the specific skill set you're hiring for, that's a bargain.
- why: http://p.migdal.pl/2015/12/14/sci-to-data-sci.html (on academia vs industry)
- how: http://p.migdal.pl/2016/03/15/data-science-intro-for-math-ph... (also got reprinted at KDnuggets)
Now, I work with a lot of people way smarter than I am, who are mostly useless because they can hardly prototype their stuff in Python or run an SQL query. And they'd expect to only work on the best, cleaned and formatted dataset and only do high-end maths on those. Reality hits hard, we're losing money paying them and they're wasting their time not doing what they like. Add to that the frustration / jealousy that this creates.
In that regard, I like that famous definition for Data Scientist: "A programmer that know more about statistics than most programmers, or a statistician that knows more about programming than most statisticians".
And don't get me started on the general repulsion for understanding the basics of how a business runs from academia. Data Science is all about application.
His original Code was cli program who's sole prompt was ?
You had to enter integers , 1, 2 3 etc to select the next option - the possibilities for errors where immense - by this time we had scaled to 1:1 tests where the chemicals for a run could cots over 10K£ per run
So you want employees who can understand complex mathematics and science but are also good software engineers.
Those people exist, but you have to pay to get them.
(As an aside, lots of Ph.D.'s -- especially in CS -- build systems that are as good or better than a lot of industry code. All of the best and worst code I've read has come from Academia.)
If they ever want to transition into DS they're going to have to skill up and do the equivalent of undergrad CS math curriculum (discrete, calculus, linear algebra, math stats, etc.), which let's be honest, you cannot pick this up in any meaningful way in a "few months" of after hours/weekend study like you can when learning to program or learning some new "framework"; either that or just remain another dime-a-dozen code-cutter. It's sad because if you just did the 4 years of computer science and all the math and the stats that go with it you'd be so close to pivoting right now if you wanted to.
You can't pick up coding like this either. See Peter Norvig's famous "Teach yourself programming in 10 years" article. The delta in the wisdom you obtain, between a few side projects over months and battle hardened experience with real products and code bases over years, is immense.
I've known many CS undergrads who are better mathematicians than people with graduate Mathematics training.
I'm from an engineering background, going in the direction of data scientist. Sometimes I find that my math skills could be stronger, and I try to read up on things when I encounter them, but still it sometimes feels like there is an infinite amount to learn. Maybe I could use some more systematic approach to it. Anyone else who has walked this path, and could come with some useful advice/resources?
I can give recommendations on books for the former but not the latter.
Just take statistic, especially Multivariate for big data (big data for statisticians is huge predictors not petabyte of observations).
Math people can do so much in data science. I think statistic is better for a non math person and it's much better suited for data. Since statistic is all about data.
On the linear algebra front, you get some understanding from your first course, but the more you internalize it by meeting the ideas in different contexts, the more useful it will be. A decent amount of higher math is turning things into almost-linear problems, and then trying to sort out the parts which don't quite fit...
Check out the syllabuses and self-study at your own pace
I feel Kaggle is sufficient for the exploratory part of data science. But Kaggle's relevance to testing and honing "production quality" code-writing skills is sometimes minimal.
In the larger context of programming and software engineering, scientific programming is fairly easy to code-up (though they are harder to conceptualize and understand mathematically). The coding part of non-CS/non-CE/non-some-parts-of-EE academia is pretty much mostly scientific programming. Yet, production quality code is seldom just scientific programming.
Do you enjoy kicking the ball towards the gate and scoring? Then you might be good for Football.
This said, I don't have an MBA, and most of the MBA people I know are not technical at all or try to stay away from the "technicalities". There's for sure a huge need of technically and business savvy people.
IMO, an MBA proves most powerful when backed by real-world experience.
The reason for the MBA was so she had a sense of the business value of the analyses she was performing, and, was more 'promotable' to executive positions.
MBA courses for data science will bore her to tears. If she adds the R and Python and takes a couple graduate level math stats courses before she graduates, she'll know more than almost all the graduates from those programs
I would not be able to do my job at all if I didn't know Python, R, and JavaScript well (and know my way around various Linux flavors). The modeling is fun, but it comes at the end of a long pipeline requiring a lot of skills that are more engineering oriented.
Rarely do my tasks sound like "model this weekly and give me the result."
Often, they sound like, "I need you to pull together data from these 5 sources, model it, and produce a weekly report showing these derived KPIs. And I need to be able to access it in a web browser so that I can send links to colleagues. And it needs to be secure. And generating a report across an arbitrary date range needs to take less than a minute."
By the way, that is a request that I've gotten at three different jobs. To give you a sense of what this looks like: most recently, I wrote Python scripts to harvest and ETL the data into a Postgres database (running on Google Cloud) and a BigQuery table, then wrote a Flask app to accept the arbitrary report requests and query the database, run models, store computed results in a local SQLite database for fast future retrieval... finally producing dynamic Reveal.js slide decks available through the Flask app.
That's a long-winded way of saying that I strongly suggest that your daughter get some practical experience with the data engineering side of data science, preferably using Python for fetching, manipulating, storing, cleaning, and preparing data. It's the most flexible tool for the job and it easily the most important tool in my data science toolkit.
Register as a Twitter dev (free) and get the tokens, etc. needed to access the public API. Pick a topic of interest -- hockey, for example, or whatever floats your boat -- and write scripts to harvest all tweets coming off of the public API related to the topic. Design a relational schema for the tweets and push them into a SQLite database. I suggest SQLite because it's ubiquitous and has a low barrier to entry. Something like MongoDB also works well for dumping everything straight off of the API. You can then have another script pull out of MongoDB and push into SQLite, for example. Once you're collecting tweets, storing them, etc... try some unsupervised clustering; k-means, for example, or a decision tree. If you feel up to it, go through 1000 or so of the tweets manually and label them according to some target variable of interest to the project. Then, use that labeled data set to run supervised models. Maybe start with a binary target variable and run a simple logistic regression. Then, visualize the data. There are a lot of ways to go about this, but I suggest trying to use something JavaScript based, such as D3 or p5.js, since it allows you to create interactive web-based visuals. Create a public GitHub repo and push work to it as you progress through the project. When done, use GitHub pages to put a summary of the project online.
Twitter data is great because there are tons of variables. It's horrifying because it's like reading the refuse of language, littered with abbreviations, emoji, and other weird characters. However, it's the terrible part of it that makes it great for learning.
There are other similar public APIs that would accommodate similar projects. Having a few self-initiated projects similar to this under your belt will really help you when it comes time to apply to graduate school or to jobs. If nothing else, it will give you something to talk about in interviews.
As mentioned separately, I suggested a data science boot camp after getting the Stats BS, work a couple of years as a data scientist, then go back for the MBA. My thinking RE the MBA was to give a sense of the business value of the analyses she's performing, and, to make it easier to promote her to executive positions.
Maybe that's old school thinking, I know that the MBA in general gets a mixed reception these days, but we can look more closely at it after she's out of school.
What many data scientists (myself included) find is that they often are excluded from meetings and conversations that provide the context for the analysis they are doing. I always tell my supervisors that it is very helpful for me to sit in on as many business strategy meetings as possible, just to listen, because it builds context around the work that I do and helps keep me properly focused.
On the flip side, many involved in business strategy do not understand the nuances of analysis performed to support their business questions. There can be many reasons for that, from not being directly involved in the analysis to being excluded from data science meetings. Effective business analysis and data science requires trust between the players and that trust is built through showing an ability to deliver focused, relevant, and accurate results that support decision making.
As somebody who likes to write code and dislikes sitting in tons of meetings, I tend to avoid climbing the career ladder to management positions. I'm happy in a Senior Data Scientist position. At my last job, I was being groomed for management and I never got to do any actual data science work. It was boring as hell. I made a lateral move to a different company so that I could be more hands on and work in an industry that is more fun. Now, I get to write code and build and run models every day. I also get to present the results and have a trusting relationship with the managers.
As your daughter finishes school and gets some job experience, don't be too quick to suggest routes leading to business strategy and management. While it's the "top of the career ladder" at many companies, so to speak, it's not for everybody.
MS's and PhD's in stats, applied math, physics, computational bio/chem were far more common.
She turned out she didn't like managing people and politics so yeah...
If you want data science stat and comp sci is the way to go imo. I'm bias cause I'm doing stat for master now.
There are some stat classes in MBA and they do prediction and stuff but your daughter may not like it. I certainly don't, they do power point and excel and visualization. Their regression classes ignore checking if their model's assumptions are valid are not (QQplot, residual, etc...) it's very dumb down.
There's too much temptation to hack something in Python or whatever, google as you go, etc. What I did was sit down and practically memorize entire programming manuals. Many academics would refuse to do something so plebeian.
After years of trying I have been unable to make the transition from Mathematics into Data Science and have since shifted my attention to other things.
Theory is nice. But Experience trumps all. Esp when you think it mostly boils down to PCA/DBSCAN, regression and scatter matrices.
A month of Python/R/D3 on some initial data like pubmed/twitter or any toy dataset would go a looong way.
This is so snobbish. What here is even think about ?
The startups and Googles of this world are the exception when it comes to employer tech-savviness.
Second, programmers at companies almost never work in isolation from other programmers, and in most open office environments are close enough they could touch another one from their desk. On the other hand, a startup with five to ten engineers may be hiring you as their first data scientist, and bigger companies may be putting you on an embedded team, with the nearest data scientist a hallway or floor away. And this isn't a big deal in terms of teamwork, but most data scientists don't have a PhD in mathematics, so if you find higher level mathematical ideas and notation to be the most efficient way for you to think about a problem, your colleagues may not. That's also not a math specific thing though -- an economist, statistician, mathematician, computer scientist, physicist, etc. are all going to have slightly different ways they think about things internally.
His audience is academics - who up until this point have been mostly in and around other mathematicians. Its important to ask yourself, "are you OK with having to sell/explain your work to people who aren't versed in even the basics of what you do?"
Obviously that isn't all data science jobs, but in my experience most of my colleagues have to evangelize their work within their organization to some level.
When you interface with non-specialists, you can't rely on the shared specialized vocabulary. This means you have to do a lot of work to distill which insights are key and figure out how to communicate them (along with supporting ideas) to an audience whose highest insight resolution is going to be varying and who may each bring their own language to the table.
So on top of building your problem domain model, you're going to have to build a model of how the people you're working with can understand it. And you will have to do this over and over.
You can be snobby about that, and say that it's soooo hard being so much smarter in your specialized field than non-specialists are, but you don't have to be snobby to realize that this is a dimension of the work that you might not enjoy or may even not be cut out for.
And of course, you don't have to be a mathematician to have experienced this. It's certainly sufficient, but not necessary. If you've been a developer in a company that has non-dev coworkers and you've never hit this turbulent boundary, you've either been remarkably well insulated, or you're so remarkably natural at that kind of job that you definitely should be doing it. :)
Translation: Can you be a good assembly line worker, devoid of pride or curiosity and just do the damned job? If so, great, we can use you! If not, please return to academia.
Lovely stuff.
I mean, by all means open up a little boutique data science consultancy that charges three orders of magnitude more for a solution that can't be updated by anyone without a ph.d and only beats an out of the box svm by .1%, where you put your pride and curiosity (your vanity) first. You might even find some suckers to patronize you, but man... I must be misunderstanding you because that is the most clueless and entitled thing I've read in days.
98% of the time data science customers actually just want some basic statistical insights into their data to make better informed decisions. If you aren't willing to help with that, and also can't find a research group to take you in (to work on their projects) and also can't bootstrap your own startup, then yes, stay in academia.
At best you're talking about a career change, at worst you're talking about a watered-down version of your intended career.
The money is good though, that's true.
Not so much. "Wants" as "Thinks it needs", but even then, the amount of student debt suggests that neither of those are entirely true either.
What matters is not how elegant your solution is, what matters is solving the problem. And doing it under cost and time constraints.
An engineer is first and foremost, a person focussed on the end goal, above all else.
You've never met a number theorist.
And you've never met a cryptographer ;-)
That is - things that rather solve a concrete problem (given all constraints, including: dirty data, finite time, understanding the needs of clients) rather that have a pure and beautiful, but utterly useless for practical purposes, theorem to prove.
But sure, tastes do vary.