The real prerequisite for machine learning isn’t math, it’s data analysis (2016)
r-bloggers.com
r-bloggers.com
Even if you're using very little of what you think of as "math" to interact with your machine learning framework, to dismiss math as not being necessary is ignoring all of the underlying mathematics that went into creating it. You may not need calculus, or the like, but the math is there -- even if you don't recognize it as such.
This seems like a straw man to me. The article never claimed that math is not necessary for the development of machine learning frameworks. In fact, it explicitly states the opposite:
"First of all, math is particularly important if you’re doing machine learning research in an academic setting."
"In particular, there are people at companies like Google and Facebook who are pushing the boundaries of machine learning – people working on bleeding edge tools. These people almost certainly employ calculus, linear algebra, and more advanced math routinely in their work."
IMO the "pragmatic rules-of-thumb" approach to teaching/understanding statistics is probably the genesis point of the reproducability crisis. This mindset toward stats has probably done more damage to Psychology than anything else in the history of the field.
So, proceed with caution.
Sure, but then we have another problem: if you need to know measure theory to run an experiment, nobody will run experiments. And lots of mathematical statistics just goes ahead and uses the measure theory.
Having done both machine learning and data analysis more generally, these points are correct. The article is explicitly aimed at the question of what is needed to do an entry level job in this kind of setting.
It's like the difference between calculus and analysis.
i think the tone of the article is good, however, because it's definitely possible to start playing with pylearn, make some SVMs, and have a lot of fun without worrying too much about L^p spaces are in the first place. later on, it could be possible to interpolate more mathematical ideas for a smoother understanding of how to measure how much preprocessing is flattening out their data sets.
(in other words, sssh this is another opportunity to trick people into doing math without immediately recognizing it as such! :-)
Like programming, machine learning has its 10xers, and I've worked with several. There's one thing they all have in common (beyond experience): they can think like the model. They've read a bunch of ML papers, get the intuition behind them, and envision from start to finish how data can be captured by various models. Beyond a small amount of hyperparameter optimization they often build a great model on the first or second take, and can put out production models in days (which would normally take novices weeks or even months). This requires math knowledge, because ML models are built on top of math. Even in higher level ML packages that do a lot of the work for you, novices get stuck in dead ends where their models don't work and they don't know why. The 10xers drive right through it because they know why their model isn't working.
When I interview for my team, I usually have non-junior candidates describe a ML algorithm they would use for a specific problem, then ask why certain things could go wrong - e.g. if they say they would use a logistic regression model to predict an action, I ask why the model may be returning the same score for every test case. A good candidate would understand that this means the coefficients are getting pushed to zero, and would list off reasons why this could be occurring (too much regularization, underpowered data, constant target variable, etc). You learn these things from experience, but you understand them because you know the math behind the models.
Similar attitudes in the past have lead into misconceptions like if you are only applying statistics in your research, you don't really have to understand what p-value is, just run the magical tests and report if it's "significant" or not.
In some organizations, "data scientist" is the modern name for a business analyst and the responsibilities largely involve basic data sanitization and analysis, streamlining processes that were manual and tedious in previous eras. In other organizations, people with advanced mathematics skills find ways to apply mathematics and machine learning to gain a marginal advantage over competitors or to deliver something novel to the market. This latter case does indeed require a deep understanding of mathematics.
Using off-the-shelf implementations of ML algos might seem like it obviates the math skills, but problem definition, algorithm selection and feature engineering are difficult for laymen to do well, and can insidiously poison the validity or efficacy of the results if done poorly.
Knowing how to manipulate tables is relational algebra, even if you never learn the formalisms.
Representing the parts of a problem as symbols is just algebra.
The hardest may be thinking in terms of higher dimension tensors (matrices that have 3 or more dimensions) but that is not even mathematics, just an intellectual effort.
Which supports the headline. The math will naturally follow as you gain experience with ML, but not having good data will make even just starting your journey difficult.
That is where everyone starts though. With time, they will become the professionals that you want to hire.
Otherwise algorithms will learn nonsense and random correlations made from datasets of random noise.
Look at finance to see how it "works".
But also : can work in a team and leverage relationships for insight.
No matter where the inspiration for a model comes from, the final formulation is always mathematical, and I think that without an appreciation for the mathematics, it's hard to get a true feeling of a model's effectiveness and limitations.
In short, to _apply_ machine learning tools, you need a good understanding of _applied_ machine learning, which is not a small subject area, but is very approachable for most.
Data analysis needs understanding casualities and correlations, and ability to use some handy tools, like chi-squared. The more tools you can use, the better analyst you are.
I know some psycology scientists who have no math background at all, but who can setup an experiment and do statistical analisys on data gathered. They have a pretty good understanding on what they do while its not strictly math understanding. Math is a tool and as with any other tool you need not to understand how a tool works, you need to know how to use that tool.
In recent times you even need not know nothing about calculations behind different statistical tests, because there are computers and software that are happy to calculate anything for you.
If we add something like ability to use R or SPSS, then we get scientist skilled enough to setup good experiment.
The replication crisis is not error of individuals, it is system error. Psychology is much more complex than, for example, quantum physics, there are much more causal links in psychology and no one know even how to speak about mind, for example: is it possible to differentiate perception from memory or from thinking? Perception cannot work without memory, and there are no way to separate them as phenomena. Psychology is much more complex than psysics, and at the same time for physicist is is normal to have p<0.001 or sample size of 10k data points, while psychology is bound to p<0.05 (it is probability to get false positive) and sample size of 30. This is itself explains while physics have no replication crisis while psychology have one.
Model knowledge certainly helps, but it isn't neccesary.
You can certainly practice statistics in a way that mathematicians would take issue with, but it's intrinsically a branch of math.
Just because you're not attacking problems formalized with the underlying probability theory doesn't mean you're not practicing mathematics when you practice statistics. It's not a separate discipline.
Lots of folks like to slag these articles and talk about how "If it's not real Maths, it's crap" (http://www.dailymotion.com/video/xgzfxs), but if you've worked in larger orgs, you recognize that there is a need for more than just advanced math. There really are myriad needs, and the failure of AI/ML (yes, I'm combining them for simplicity here) in an org is usually the lack of understanding these needs. Here are some I've seen:
1) What is the problem that AI/ML is being applied to? What is it optimizing, deciding, predicting, forecasting, categorizing? How will this decision be used in a process or flow? This requires a business analytic approach, understanding available data, _what it means_, how it's generated, and how the business might evaluate the impact of the ML/AI. These folks need to interact with the business folks as well, so some communication skills are helpful... though that's true for every role these days.
2) How will said decision be implemented both in tech build, test/QA/FUT/UAT/etc, and prod? This technical architecture and approach is data engineering, but some folks in the data science world are amazing at this. BTW, a model that works in dev may not scale in prod. The fact that so many folks keep "re-discovering" this is scary to me. There is a whole class of amazing folks who can re-implement models to scale them, and if you know them, reward them well.
3) How will models/algos/systems be built? This workflow is often pretty sloppy, just a bunch of jupyter notebooks or a tonload scripts (but they're in Git, so it's ok), and so replication and scaling becomes painful... esp. in regulated industries. Again, data engineering approach, but needs a more nuanced understanding of the vagaries of ML/AI. I find that solving business problems may not always fit a traditional software workflow (call it "agile" all you want, it doesn't always fit) but ymmv. So, you may need to create new workflows for your org's needs, and this tooling may not be off the shelf, but like automated testing, this tech debt will need to be paid sooner or later.
4) What ML/AI approach should I use, esp. if I'm using pre-existing approaches? This becomes more of the data science analytic approach, mixing the understanding of how ML/AI works with the data landscape (how was my internal data generated? What does it mean? Is it stable? What's available at score time, and what's it's latency?) and how to build various working predictive models. Note that this is traditionally where data scientists/model builders/analysts spend their time, from data cleansing to preliminary analysis to generating/training various models to crying when none of them predict well to stumbling onto a fascinating and amazing combination of models and approaches at 3am. Yes, math is defn helpful here, but _understanding_ the underlying math is often helpful enough, vs. _mastery_. A good understanding of the levers affecting each model/algo/DL architecture/etc. and how to diagnose them can get you pretty far, though you may violate assumptions or overfit if you aren't careful.
5) Real Algo design: Using all that math that underlie the models to not just diagnosing a pre-designed package but making your own optimizer, or your minimizers, or your own way of computing the Hessian, or a new weighting approaches, or whatever new approach your expertise is in. From writing your own primitives to making your own deep learning architecture, this is often the real wizardry, but it's not needed in EVERY case. But the effort can really pay off, from optimized prediction, improved use of resources, and unique IP which can be a competitive advantage. Remember, custom work is great, but maintenance of the resulting product can be painful, esp. if done in a language few folks in the org understand and if few resources are available outside.
Finding skills to meet all of these needs in one person is certainly possible, but somewhat unicornish. And you may not need all of these either. But the best orgs that have to scale have at least a "Data Prep/Data Engineering" role, a "Data Science" role, and if needed, a "Data Analyst" role who can help translate specific business needs into analytic problems, and also carve off the easy stuff (ad-hoc pulls and simpler analyses).
So, to original post: yeah, sort of. You need to be able to analyze, but you also need to be able to do some of the math, and understand the impact of the rest. If you are awesome at the math, go up to making your own magic. But if you aren't, at least try to understand as much as you can, while trying to also be familiar with the other areas. I do think that the emphasis on visualization is somewhat glossy-ily ignoring real analysis skills, but it's one of many helpful approaches to understanding the problem. (Pet peeve: Just drawing graphs/charts is usually not analysis. Sorry, but true.)
If you are hiring for a data scientist, at least be clear with yourself on what of these needs you are looking for, and if you need them all. Also be clear about where you hope to grow, and if you are hiring for now, near future, or the year 3000.
(Ok, one more: do yourself a favor and learn experimental design. I am sort of shocked at how many data scientists I chat with who haven't really tested champion/challenger models, or compared 2 or more processes or approaches in a randomly assigned test, or gotten as close as they can with quasi-experimental or matched-groups approaches. The simple ones are indeed really Excel simple, but if it's so easy, why haven't you at least tried it? And the causal-modeling stuff is a bayesian's dream come true, so it's worth learning.)
1. who teaches it
2. what gets taught first
3. what forms and notations get used
https://web.stanford.edu/~hastie/pub.htm
http://bayes.cs.ucla.edu/jp_home.html
http://andrewgelman.com/books/
I guess it all depends on what is a mathematical background and what is advanced!
What I found is that many "data science" books cover how to use R or Pandas at a very introductory level. Books like ESL focus on core theories (which is great) but do not focus on how to tackle a tough real-world data.
I suppose much of data insight come from experience, but I was wondering whether there are sources to help me jump start.
Sometime it suddenly becomes apparent and we suddenly see how to do it and everyone feels pretty sheepish!