How to Get a Job in Deep Learning
blog.deepgram.com
blog.deepgram.com
I'd like to clean up a bit the air from the hype fog:
DL is giving amazing results only when you have big sets of labelled data. Hence it will be much cheaper for companies to buy Google/Microsoft Vision/Audio REST APIs rather than paying the costs of: cloud + find data + deep learning experts. So, I don't think we will see a massive growth of DL gigs.
e.g. Google Vision API: https://cloud.google.com/vision/
Except those areas where your own CNN implementation is needed (automotive, industrial automation), Deep Learning will be another "library" in the ever increasing Software Engineering mess of gluing many open source libraries and REST apis to get something useful done. You need 1 guy training a Neural Network for every 100 software monkeys maintaining the infrastructure complexity. There are now many Software Engineering jobs because it's hard to glue and maintain publicly-available code to solve some specific business problem.
I think the the same applies for many Data Scientist jobs, which are these days more about fetching/cleaning/visualizing data than making machine learning on it.
The advantage is not the Deep Learning algorithm.
Some theoretical progress has been made on Neural Networks lately, but largely it's the same stuff from the 90s, with much more GPUs and data.
The competitive advantage is the cloud, and the software mess that keeps it alive.
I think Deep Learning experts will be like Linux Kernel experts. You need 1000 kernel experts in the world, but you need 10 million javascript monkeys that code what dialog message appears when the user does something stupid in some app.
Does anyone call a doctor a "diagnosis monkey" because she isn't in the business of inventing antibiotics? Is an architect a "wall-placement monkey" because he doesn't create new kinds of construction materials himself?
In the software industry, professionals who are just as important as doctors and architects routinely think of themselves as mere monkeys or data plumbers at best -- and the reason is this never-ending cycle of nerd hazing.
This phenomenon is also there. But I think, more prominent is the derogatory usage by others, especially when coining terms like <xyz> monkey. For that matter, I am Okay with critiques of tech/practices e.g. Javascript. But definitely don't like derogatory usages like calling the practitioners JS monkeys. Which was the case here. Difference between attacking techs/practices and people.
We mainly do fraud detection and security related work. We have also seen operations workloads (forecasting when machines are going to break or preventative maintenance) Most of our core business isn't even CNNs.
One thing that's missing from the narrative is that researchers who vision and speech because it's "pretty" and you can demonstrate results on a large feature vector. It's also very relatable for normal people.
Most of corporate america (not silicon valley) has more traditional things like time series data. Not images.
I would say here that deep learning use cases aren't explored by most people and that there are other areas besides what the marketing with self driving cars is perpetuating.
The other thing I would add here is when companies get bigger they typically need to take core competencies like vision and speech in house. It can be hard to justify outsourcing your core business to google as you grow.
This trend however might change over time. I would love to hear contradictory statements here.
Disclosure: I am a deep learning founder in the same batch as the authors of this blog post.
I work with .net and java shops where running on windows is a requirement and hippaa is considered "lite" and they dont know what a gpu is.
Despite the marketing noise that is still most of the world.
Look no further than another comment that showed java job postings vs machine learning.
My biased comment out of the way: if you can convince ops on why you need this deep learning thing then be my guest. Their first question to you is probably going to be: Does dell/hp sell this as a reference box?
It mostly comes down to roi and what your existing stack is.
My job is hard enough as it is. If you can convince IT to allocate an r&d budget be my guest. That is more or less our specialty. Email is in profile if you want specifics.
"do research using vision and speech"
Too late to edit.
It's true that DL will probably become just-another-library, but that will happen only once computing becomes extremely cheap on the petaflop scale (it isn't cheap yet). Even after that happens, the people that spend time doing DL now will be trained in a way of thinking that will be in demand for a long time.
Just like how there were countless excel spreadsheet business processes that birthed CRUD applications.
Well, creating ML algorithms is a small part of data science. And getting data, and making sense of it, is a highly non-trivial part (though, somehow underappreciated). Very often it requires a lot of knowledge, skill and experience in used methods, dataset and statistics in general.
Anyone can plug&play a scikit-learn algorithm within 2 lines of Python, yet it does not make everyone a data scientist. Anyone can copy&paste&run a deep learning network architecture, but without proper care it is likely to terribly overfit to underprepared dataset.
But I just don't see it - machine/statistical/deep learning gigs just seem really rare.
I know this isn't a great metric, but searches on Indeed.com:
"deep learning" - 873
"machine learning" - 9,762
"statistical learning" - 65
java - 72,802
javascript - 43,785
Same searches on LinkedIn: "deep learning" - 646
"machine learning" - 6,952
"statistical learning" - 34
java - 43,845
javascript - 30,818
Even the "machine learning" search on Indeed, with 9K+ results has 1300+ from Amazon, followed by a much smaller number (in low hundreds each) from Microsoft, Google, others (including some that look like staffing companies).Even on HN's who's hiring Sept 2016 thread phrase counts: 14 "deep learning" 79 "machine learning"
I completely agree with the idea that being able to use some deep/machine/statistical learning is going to be a toolset that data hackers need to have. I even think that there is a bit of the "build it and they will come" magic waiting out there.
But I think the best way forward is to be working in data and figure out how to generate value with deep learning - this will be much more productive than trying to seek out a deep learning gig in terms of promoting deep learning in the workplace. Heck, that's a suggestion I would be wise to take myself . . .
It could be that demand for deep learning jobs is growing faster than ML jobs, or vice versa.
Here is a graph of something similar: http://www.indeed.com/jobtrends/q-%22Data-Scientist%22-q-%22...
There is definitely more demand for data science than deep learning, and much more demand for software engineering or development than data science. Of course the supply also matters, there are many more software engineers than data scientists. But still, deep learning is a niche skill.
As an example, one of Fei-Fei Lee's recent grad students got multiple job offers upon graduating, one for more than a million dollars a year: http://www.nytimes.com/2016/03/26/technology/the-race-is-on-...
And more so recently with their being quite a significant demand for them in enterprises looking to establish big data analytics programs.
Do people who implement (albeit real, useful) deep learning systems, but who have no formal machine learning background, who don't really know much or care about implementing derivatives or softmax functions because the frameworks abstract all that away - are these people getting offered jobs?
No, I don't think so yet, but even if it did, would it even matter? There would still be world of different between the teams that can build a Twitter/LinkedIn/Github in Rails to someone who knows vaguely how to string something together because they learned it on codecademy
The barrier to entry for AI stuff is quite a high. And IMO will be limited to certain kinds of people who like working on the hard science /math stuff, and have a strong enough early education to learn advanced linear algebra and statistics (for example) .
In many ways flying itself is easier than driving because you have less chance of a collision.
The really hard parts are takeoff and landing, not sure if one could learn to takeoff and land a Cessna badly in a week long bootcamp.
Yes, you can. Taking off is extremely easy, you just need to accelerate with the flaps down and pull up at around 90km/h (if I recall correctly, it has been many, many years since I last flew an airplane).
Landing is much more complex, but you can do decent landings[0] in normal weather conditions after a few classes. The course I did (as I said, many, many years ago) had 35 flight hours and I did it in a summer (2-3 classes per week, usually 1 hour of theory and 45 minutes of flying) and you were perfectly able to do the exam (and pass) after it.
[0] Of course, in the eyes of many pilots, a decent landing is one you can walk away from. A good landing is one where you can use the plane again.
Maybe someday all of that will be abstracted enough that you don't need to know math to do ML (like how now you generally don't need to know machine architecture to program), but how soon will that be?
For example, for a while there's been work on doing sentiment analysis using machine learning and they typically train them on a data set of movie reviews. It turns out that as soon as you apply that trained system to anything other than movie reviews the actual results are quite poor, but you might not catch it.
I started college in '97 and specialized in robotics and AI. Dropped out during my fifth year (2003) for several reasons (some having to do with the university, others were personal/family related). My working experience begins in a startup in '99 doing web stuff.
Now, in my country and in the early 2000's there were no jobs in robotics/AI, so I continued doing web stuff and, through the years I've moved up and down the ladder (I was development director of a multinational, left for a senior developer position because I was bored), lived in 4 countries and moved into different sectors and technologies. AI/Machine learning/Robotics were a hobby and I tried to keep up to date[0], but I had very few opportunities to apply that knowledge in my day to day.
In terms of education, I'm a drop out: I went through five years of university and have more than enough university credits for a bachelor degree, but have never felt the need to go back and officially get a degree[1]. I started a master's degree in Ireland and left it after the second module[2].
Now, finally, the point I wanted to make (sorry for the long introduction, I felt it was important to give a bit of background):
About five-six years ago, I got my first official position working full time with AI and machine learning[3] and suddenly found that my maths knowledge was definitely lacking. Even today, although I'm always improving, I find that when talking about deep learning, the maths are definitely my weak point even though I'm able to implement models, systems and tweak them enough to get good results.
So, to answer your question: Yes, I keep being offered jobs. Maybe not as well paid as someone with a similar level of experience and a PhD and definitely not research-oriented. If the company does have a data science team, those jobs are usually working with them to take their models and integrate them with the rest of the product/s.
[0] I went through Andrew Ng, Sebastian Thrun and Peter Norvig courses originally through Stanford Online and I'm constantly following one or two coursera or MITx courses in the area. I also have an account in Kaggle but I don't usually push my results, I'm more interested in the datasets.
[1] I'd need to go back to my country and fight the bureaucracy there while trying to get documents from a university that no longer teaches computer science.
[2] I felt a bit cheated, to be honest: The syllabus looked great, the content was mediocre at best and the papers we had to do had very little to do with the stuff we learnt and even with IT in general. 4K euros per module (for an online master) and the time spent wasn't worth it just to get an official piece of paper.
[3] "Data science", I don't specially like that term and, if anything, I prefer to call myself a data engineer. I don't do science, I don't do research. I can take a research paper, implement it and add all the stuff that makes it useful and production-ready. But I don't call that science.
Do not label yourself as a data scientist or machine learning expert. Go for the domain, i.e. become comfortable with the actual data and the methods used there:
- predict land use in aerial imagery - become comfortable with photogrammetry, geography, etc.
- predict biological tissue(s) - become comfortable with specific branches of biology or medicine
- predict $something_relevant
I actually stole this advice from the epilogue of some text about programming, and it really stuck with me. Otherwise your expertise is just too generic and you compete with a big pool of people who call themselves machine learning experts, because they can write a for loop in Bash.
Curious to know if anyone has had success learning/re-learning these as a mid-20s or older adult who works fulltime, and if you could potentially provide a list of books/courses to go through. I personally never learned anything past geometry (in high school). The most advanced math class I took in college was College Algebra. That means I never learned trig or anything past it (so no calc, linear algebra, or probability), and I'm sure most people on HN surpassed me math-wise sometime in high school :)
I've been able to skate by with my embarrassing lack of math knowledge/skills as a developer, but I feel like it's only a matter of time until the mathematical steamroller becomes a serious threat career-wise and I get crushed.
for probability and stats: https://www.amazon.com/Introduction-Probability-2nd-Dimitri-...
for linear algebra: https://www.amazon.com/Coding-Matrix-Algebra-Applications-Co...
this was my college calculus textbook: https://www.amazon.com/Calculus-7th-James-Stewart/dp/0538497.... I can't comment if it was good or not because by college, I had taken calculus twice so it was all a refresher
best of luck! You sound educated enough (yes, I'm judging from the couple sentences you wrote) that I think you won't have any problems acquiring math knowledge with persistence.
You say didn't take trig? Check this out as one example: https://www.khanacademy.org/math/trigonometry/trigonometry-r...
I graduated from CompSci in 2011 and worked as a proud code monkey since then.
Last year I took up a Master's course in AI (well, "Intelligent Systems") which was chock-full of machine learning modules, as expected [1]. I finished just a few weeks ago.
I struggled a lot, particularly because I did the course part-time and I was only offered an (optional) maths course in the second year after the machine learning and image processing modules- and concurrently with an NLP module.
The maths module helped a great deal and it cleared up a hell of a lot to do with differentiation and linear algebra. I still struggled with the maths module itself and to be honest I got a lot of help on the homework from my roommate who is a maths wiz, but in the end I managed to feel comfortable with the material and to get a good mark (75%) in the maths module.
I think I impressed my maths lecturer even -but that was my coding skillz, haha. Teach was amazed I managed to implement an LSTM RNN in Python [2] (we went over optimisation and had to code a hill-climb/gradient ascent algo, plus a RNN). In truth once you get the maths down pat the coding is nothing, but I guess mathematicians are crap as coders :P
Anyway, what I mean to say is: yes, totally. You can totally get at the very least comfortable with the material required for deep (or standard machine) learning. And if you've built a good level of coding skill from work you'll find it gives you a big boost, and can even help you clarify some of the maths. Frex, I feel I got a good intuition about optimisation in general from implementing hill climb/ gradient ascent, especially because I spent hours watching it trying to get over a ridge in a datascape and getting stuck every. single. time. the dumb thing XD
I think you'll find linear algebra in particular the easiest to learn because it's almost just array manipulation and you've done that in spades. Differentiation is a bit harder, but at some point it clicks ("it's like stepping on a break pedal" or some such analogy) and it goes swimmingly from there. Probabilities are not hard either, it's just a form of logic. Other stuff- ymmv, but it's all doable.
As to books- I don't have any recommendations. I think any popular textbook is going to be good enough to get you started :)
[1] How I ended up doing that course- I got into AI via logic programming, that I learned during my degree. Many Prolog textbooks are also AI and particularly NLP textbooks so I thought it would be a shame to let all that kewl stuff I learned by reading them go to waste. At the time I had no idea how hot machine learning is in the industry.
[2] Plug: https://github.com/stassa/lstm_rnn
The most frustrating thing will be going through basics that don't seem connected to your eventual goal of ML; this is potentially a long phase.
I'd recommend Khan Academy for the math basics. Schaum's Outlines series for lots of worked problems. I'd recommend Strang's MIT OCW videos and books for Linear Algebra. I first learned calculus from an economics book, probably best not repeated ;). Bishop's PRML book is a good source on probability, Bayesian stats and machine learning.
I did a part time MSc in applied stats while working full time. Even in the MOOC era, there's nothing like real exams to focus your efforts.
Although I was pretty confident in my maths up through basic calculus, I began fresh earlier this year starting from pre-algebra on their videos and problem sets.
I took the approach of looking at the mind as a muscle and if I had taken a long break from lifting, I wouldn't jump right back in lifting the same weights. The analogy isn't perfect, but I feel like it helped to reinforce the old basic neural pathways in order to prime my mind for more difficult topics.
EDIT: Also the achievement points aspect helped as a learning tool as well.
I also have one on linear algebra: https://gum.co/noBSLA
Piece of advice: don't skip the exercises. It's great to learn and understand math, but you don't really learn the material until you have to solve problems and make use of the math. Since you're a developer, you'll probably also enjoy this short tutorial on basic math using SymPy: https://minireference.com/static/tutorials/sympy_tutorial.pd...
It is highly unlikely that you will get a job in which you exclusively use deep learning alone, and not any other ML/AI technique.
Once you learn DL, then, "congratulations... here are 100 other topics you might need to know about before getting a job". http://scikit-learn.org/stable/tutorial/machine_learning_map...
In the past you need to know that to recognize lines: hough transform, recognize polygons: line simplification, recognize face: cascades, etc etc.
Now? You can almost just feed it arbitrary labeled training data and do well without any sort of feature engineering. Just another api to glue.
Would someone in their right mind create a self driving car product with the help of someone who learned deep learning from a blog or youtube? probably not.
Don't read too much hacker news, it kind of becomes stressful and I would try enjoy your weekends, don't worry too much about the market, a lot of good people just burn out and have breakdowns by trying to understand all that's going on and become useless anyway. Just know the basics well and learn what you need to in work hours, make time.
What I'm finding is that in the end, most of the good / important stuff ends up condensed into a nice O'Reilly (or similar) volume that you can read at you leisure' later on when the hype has evaporated. If you invest yourself too much in the latest tech constantly, you run the risk of it being redundant / replaced anyway.
If you are okay with median compensation and median project importance (internally and externally), then sure, wait a couple years when the interest has died down.
I imagine a product like this could actually charge a fair bit of money helping companies and people improve the 'virality' of their tweets.
"How to get a job in deep learning" would include:
- What specific topics will be asked during interviews
- What the interview question format is like
- How to prepare for the interviews
- How to get interviews without a PhD. What do you need to show competence in your self learned skills?
They want to see if you did cool stuff before you applied for the job. If you didn't then you won't get an interview, but if you did then you have a chance no matter what your background is. Of course, the question of "what is cool stuff?" comes up. If it is building small projects with a a little bit of success, that probably won't do it (it might work for larger companies, or companies that need light ML/DL performed). But if it is "built twitter analysis DNN from scratch using Theano and that can predict the number of retweets a tweet will get: here's te accuracy, here's a link to a my write up on it and here's a link to github for the code.".
Edit: added words similar to this at the end of the blog post.
Statements like this contradict what you are saying here - to really build a model that predicts number of retweets based on the content of the message (not something like the average number of retweets this user has) is very non-trivial. If your threshold of a side project is publishable [1], it is an unrealistic expectation.
- have good problem solving and coding skills - spend a month or two learning how to build networks using good libraries
Then you will be able to get a good result with the Twitter task I pointed out. It takes being able to work input data correctly, think about what matters and what doesn't, then synthesize using readymade DL tools, usually in python. None of that is complicated esoteric neural net stuff, it is just motivated problem solving.
I'll add that as a big point — you should be able to code, problem solve, and be motivated.
Sometimes software engineers need to accept they aren't the smartest ones anymore. This is why they can't get the "smart" and "cool cutting edge" ML/DL jobs.
If you are that type of person then you can kill it with just a few weeks of hard study on your own.
False. Being a great software engineer is not enough to get a job in deep learning. For the same price, or a little more, you can hire an "expert" scientist or engineer with a PhD in CS/stats/ML.
Would love to hear if I missed something!
Another point is that this article is really about how to learn deep learning, not how to get a job. I would really like to see some evidence that: "The good news is that basically everyone is hiring people that understand deep learning." Most data scientist jobs I have seen don't require or use deep learning.
Had it been, i would have clicked the back button.
We are trying to address this soon!
I posted about this earlier today but ML really should be demystified. You can write a lot of commonly used algorithms in 100 lines or fewer. The math is not complex. If you can get past the notation and buzzwords like "deep learning" (it's an artificial neural network, itself a grandiose term) you'll see it's not as daunting as most think.
The reality is most "data scientists" will be working on implementation rather than creation. They'll be working on data sets and error analysis, not creating the next buzzword-laden algorithm.
Convex optimization is not high school level math.