Get some textbook suggestions and make a minimum of reading 5-10 pages per day. In about a month or two, you're done with a 300 page book. Repeat that for a few years and you're an expert. Once you have the foundations, read papers too, but don't skip straight trying to using AlphaZero to solve a curve fitting problem.
This. Highly recommend Russel & Norvig [1] for high-level intuition and motivation. Then Bishop's "Pattern Recognition and Machine Learning" [2] and Koller's PGM book [3] for the fundamentals.
Avoid MOOCs, but there are useful lecture videos, e.g. Hugo Larochelle on belief propagation [4].
FWIW this is coming from a mechanical engineer by training, but self-taught programmer and AI researcher. I've been working in industry as an AI research engineer for ~6 years.
[1] https://www.amazon.com/Artificial-Intelligence-Modern-Approa...
[2] https://www.amazon.com/Pattern-Recognition-Learning-Informat...
[3] https://www.amazon.com/Probabilistic-Graphical-Models-Princi...
Statistical Rethinking https://www.amazon.com/Statistical-Rethinking-Bayesian-Examp...
An Introduction to Statistical Learning http://www-bcf.usc.edu/~gareth/ISL/
Russel Norvig should be treated as a subtle intro to AI.
The start Bishop to understand concepts.
I'd be ecstatic if I never again see a comment about how folks suddenly and completely understand a class they failed years ago after watching a 3blue1brown video.
So do exercises, spend time digesting and trying to explain things to others. If you feel it's hard, well you are correct. Get comfortable feeling that way. Hopefully theres light at the end of the tunnel. Dont buy into the hype. Know the basics
Edit: so didnt see op said exactly this. My bad, new year and all.
Finally I started studying serious books in a disciplined manner. I wish I should have done this earlier.
I recommend Kevin Murphy's ML a probabilistic approach and Ian Goodfellow's Deep Learning.
Those are the books used in most of the ML courses I took in grad school.
There is also Chris Bishop's Pattern Recognition and Machine Learning, but I think it is less popular now, than it was before.
For purely professional purposes, it's probably a better idea to take a non-academic approach.
A minimum of 5-10 pages a day for 6-12 months seems pretty reasonable for a career change into a competitive field.
Patience. Spending that extra time (it helps if you really enjoy it or can program yourself to really enjoy it).
https://www.youtube.com/playlist?list=PLZHQObOWTQDNU6R1_6700... for inspiration.
To answer the question though: I'm not sure what you produce other than maybe blog or publish some analysis using your data skills. And maybe: https://github.com/MaximAbramchuck/awesome-interview-questio...
(source: non-CS engineer at amazon who watched the amazon videos internally before they were made public. I'm not a data scientist yet but sometimes, esp. when people talk about the challenges of AGI, I think about transitioning.)
I am not an expert, but from what I've heard/seen, being really solid on the fundamentals of regression and feature modeling (and not being afraid to read and apply ArXiv papers) are all key. And eagerness and statistics go a long way and are valuable to companies.
MOOC: http://course.fast.ai
I just found a resource a few months ago that I'd love to recommend, but haven't started yet. It's mentorship you pay for, but not up front. You sign a contract to pay a certain percentage after you're hired. I plan on going through this program if my current job leads don't pan out.
Mentorship: https://sharpestminds.com
I'm interested in comments about either program in general. Speaking as someone who also has an EE degree, went through a web development bootcamp, and was disappointed by both at the help in getting hired that was offered after the curriculum was finished, I am also interested in your findings and results.
EDIT: Also, I strongly concur with the fast.ai recommendation for deep learning, especially if you're starting from a background in software.
Some stats about our mentors:
- There are about 60 of them now
- Geographic distribution is ~1/3 in the Bay Area, ~1/3 in the Toronto region, the rest across the USA and Canada
- About 50% are deep learning engineers, the other half are a combination of ML devops, data eng, traditional ML (clustering, boosted trees, etc.)
- About half work in (or are alums of) the AI labs of major companies such as the ones whose logos are on the website
Why we haven't listed some of them on our website yet: no good reason. We'll probably do this soon. It's a good idea.
I think the advice about getting in as a hardware engineer is solid. At my workplace, there's a ton of need for people working on specialized hardware for DL, and for people working on the software that works with it (optimizing compilers, etc).
If you are looking to break into the software side of DL, the first two thirds of the Deep Learning book [1] contains all the math you need to know to pass the interviews. Then, it's just a matter of getting interviews; I found that I needed professional experience deploying DL/ML to do that. I got that by doing side projects at work. For instance, we had a long standing operations research problem, and I spent some free time at work implementing a RL algorithm to solve it. I didn't get too far, but I was able to talk coherently about the papers involved and about how I planned to conduct the project, which went a long way.
Are you implying that, once prepared well enough, the contents of the interviews are simpler than getting actually noticed in the pile of applicants ?
Much easier to quiz the applicant how they would solve a problem, or to discuss a previous project or paper they've published (or are interested in). Some people will find that much easier than whiteboard coding, others will hate it.
It really depends where you apply and if you want an applied or research role. Some places won't touch you unless you've got a publication in somewhere like CVPR. Others will go _hard_ on the stats questions. Other places want to see a strong Kaggle rank or some personal projects. It's really useful to have a portfolio here.
Does this mean you'll be good at the job? No. Is this very wasteful? Yes.
Getting interviews, on the other hand, requires you to read the recruiter's mind, and can vary depending on what the recruiter had for breakfast, or if they fought with their significant other that morning. It's much less formulaic.
I'm an iOS engineer without a STEM background, and I've been contacted by Amazon recruiters for entry-level ML/AI positions. I thought it was weird, but they said they've hired a few people with iOS backgrounds and no prior ML/AI experience who are now excellent ML engineers. I backed out because I knew I would fail the interview process at this point, but it's something for me to think about for the future.
There are tons of youtube videos and books (Cracking, Dynamic Programming for Interviews, etc). Definitely do research into questions that will be asked.
Linkedin, Cloudera, Redhat, Quora, Robinhood, Asana, Salesforce, Dropbox
Lots of good advice about acquiring skills, I don't have much to add beyond that. I'll just mention that before you jump into the advanced stuff, please understand the terminology and basics very strongly. I've interviewed over twenty people for roles in ML the last year and many (despite having ML on their resume or even some experience in it) could not even explain the difference between training/inference, the meaning of validation, etc. The field is so hot right now that many unqualified folks are trying to get in, often by faking more experience than they really have. In response, I've created a simple 'fizzbuzz' test just so I can quickly screen people.
You're mostly on track with your plan to build something. You do need to demonstrate that you have the skill set, but building one giant thing isn't the answer. There's so much that goes into building a giant thing, that I can't accurately access your ML skills.
Ideally I like to see a lot of small things over a reasonable amount of time. Someone with a solid GitHub showing 6-12 months of paper implementations, weekend geez-wiz hacks and various other projects would go right to the top of my call back list.
Hope that helps, good luck with the job search.
Non-FAANGs may be less picky but the competition in the field is too great at the moment (due to MOOCs/Bootcamps increasing supply), and even with an excellent portfolio it may be impossible to stand out. (in my case, despite my data science "fame" most recruiters tossed my resume out immediately during my job hunt a year ago; the only interviews I got were by going above the recruiters. And that was for data science, not even ML/AI)
Even after working as a Data Scientist for over a year, I've received practically no recruiter spam.
I am a Director of Data Science and Software Engineering for a mid sized firm (~1000 employees and $150-200MM revenue). I started with a Finance degree then shifted into an analysis position at a FAANG (lots of excel, SQL, learning how to query big data). This eventually led to learning more about tech (python, AWS cloud stack, messaging queues) and after 8 years in the industry giving me enough experience to manage teams of data scientists, software engineers and data analysts.
Although it is so important to know all the software engineering stack, many companies will benefit from simple business intelligence and data analyst roles. I guess my recommendation is to also keep an open mind in looking for these types of roles in the market (data analyst, business intelligence engineer), because given your desire to learn and existing background, its clear you can make a big impact in those companies as well. And it will be much less competitive than traditional CS crowd.
Some food for thought
A lot of companies that say “data scientist” when they really mean “spreadsheet analyst” are places to avoid if you have career aspirations in ML. In the worst cases it can be a bait and switch (very common) to get overqualified people to babysit rudimentary analytics. Especially avoid places that might do this to pad their staff for any type of acqui-hire or investor signalling reasons, because your career goals will not be acknowledged.
In the best cases, it can be some befuddled IT manager who vaguely thinks they need “AI” but really they don’t have projects that would actually benefit from it. They might be sympathetic to your dissatisfaction in the reality of the job, but will have little power to do anything about it.
Somewhere inbetween is another very frustrating case: situations where the business or product clearly can materially benefit from “real” machine learning, and from the perspective of making customers happy & making money it’s a no brainer to invest time to research implementations, but risk averse management, often with no ability to gain an understanding of the benefits of investing in machine learning, or who want to act as credit / politics gate-keepers for an existing system, puts the brakes on it and retasks you on things that just waste your talent.
I'd suggest you look for an opportunity to apply ML/CV/AI in your industry (deep learning for PCB inspection maybe?). Show the possibilities, get some research funding, do a pilot program or similar. Lead and drag your company (kicking and screaming if need be) into the 21st century.
Then you will have ML/AI on your resume, and recruiters will come looking for you.
Also, you may not realize but we’ve been using ML techniques for decades in communications. Gradient descent is used to optimize equalizers; maximum likelihood estimation(and equalizer optimization) used for phase estimation in high order QAM. Plenty of other examples. So you probably are already familiar with much of the basic tool kit. I had a wannabe startup founder in ML tell me that there’s no way I could possibly understand the stuff if I didn’t have PhD in that area in CS (I am physics). I just smiled and nodded.
The usual answer here is look for suitable business / r&d cases within your own EE industrial domain and use ML/AI as any other tool instead of as a black box or a magic wand. Good luck.
I cant seem to get past HR. My resume has that I'm a Chem Engineer BS, Industrial MS, 7 years in engineering, 2 years of Electrical Engineering.
The first page of my resume is my 10 years of non-career programming experience. Built a Dishwasher(embedded C++), full stack app(RN JS, Mysql PHP laravel), and smaller projects.
I cannot get past HR.
Every real life programmer I show my work to, knows I'm capable. Heck even some got me in touch with HR. Nothing came of it.
This must be part or most of the problem. Cut your resume down to 1 page, if possible. Include a meaningful cover letter catered to the opportunity and specific company youre applying to. Shove the last ten years stuff into the very end, and start that first page with your software knowledge and related projects. Ping me someday here if this ends up getting your foot in the door.
ML is unfortunately better done in a big company due to data but also a b*Ch due to tremendous friction within org to get things done.
Another key strategy is to commit yourself to build an end 2end ML application, structure your learning around it. I found this a tremendous technique to turbo charge my learning.
BTW, I think the you may have a bit of an advantage because of the math background you presumably have with a BSEE (linear algerbra & differential equations).
This is a very tough track but uf you have software development experience it should be easier to get a role in ML or Data Science.
I know ML/AI is all the rage. I just feel that targeting it so heavily is a bit shortsighted.
Maybe that is too simplistic but I can't help after my 1 semester ML course think that most of the ML problems people are solving aren't really suited at all. Like SWE see this cool hammer and now everything is a nail. Maybe I should read up on startups using it successfully for anything but I haven't seen many of those on the frontpage.
I can’t reccomend Kaggle enough for those who are looking to prove their abilities in the field.
Or put another way: are there plenty of problems where ML/AI are valid tools or are they largely cool tech looking for problems to fit into?
Although I went right after my undergrad, there are several in my cohort who were in industry for as long (or in one case much much longer) than you have before starting their PhD.
There are certainly ways to merge your hardware experience with learning. Either applying AI to hardware design, or applying hardware design to speed up or otherwise improve learning, lots of research going on in both areas.
ML-SWE: SWE with ML focus - building architecture around models, feature engineering, distributed training, etc. Relatively limited ML knowledge needed (IMO). The math won't be helpful for this role. Much more important to have SWE background. If you want this, keep building your programming knowledge (Python) and read books. Would focus on understanding the popular frameworks PyTorch and TensorFlow b/c your work will likely interface with those.
Research engineer: Mostly for MS/PHD background. Farther away from the product and closer to actual research. This doesn't sound like what you want to do.
Data Scientist: ML is a subset of the knowledge needed. Applied statistics as important, if not more so. Doesn't sound like you want this.
A path forward:
(1) Program a lot. On what? Anything at all, b/c you need programming skill to work as a SWE.
(2) If you want to do ML-SWE, program with an eye towards ML applications. Maybe do a simple cloud project that leverages ML - Google Cloud makes this particularly easy for classification tasks. Focus on breadth here, not depth. No sane person outside of academia can keep up with state-of-the-art and truly understand it. Far too much material, so focus on fundamentals.
(3) Work towards your strengths. You aren't some hotshot kid out of college proclaiming to be an AI guru. That would be silly and no competent recruiter would believe it. You know hardware - and AI (neural networks) leverages a lot of hardware. Why not focus on the hardware side of AI? Demonstrate your knowledge of how/why TensorFlow is so effective across distributed hardware, or how CUDA accelerates NN computation, or why TPU claims vs. Nvidia may be up to interpretation, etc. This should be a natural transition given your background.
TLDR; Know what you really want to do. Your background is valuable. Play to your strengths. Don't ring the bell.
From your question, it sounds like you want to be a software engineer rather than an ML/AI engineer -- is that a fair assessment?
I have a PhD in EE, working in semiconductors. I have done a couple of MOOC specializations on Coursera, and am trying get some data science projects on my resume. Also trying do some Kernels / Scripts on Kaggle to build up a basic portfolio.
on FAANGs, the teams are usually huge, 100+ people doing what a nimbler company does with 3 or less, to the point the employees don't even see that, because the product is now broken into several pieces to give the illusion of complexity. middle managers then break it down further that engineers start being know as the "person that writes the java files in that one directory" and nothing else. This creates constant fear of becoming irrelevant. All while you see 2~10% pay raises while hearing about undergrads making the same you make now with "new tech du jour". This creates even more pressure.
And because this cycle (stagnate, fear, learn, relief) repeats often, engineers start to associate learning a new tech with happiness, just because it offsets the psychological fear for awhile.
In Amazon, it’s easy to move around, but not between job families. I think it’s a bad idea to join as a SWE and try to transfer because they want people that have done it before, and you’re unlikely to do that sort of work as an SWE. I think it’s better to get experience in the role you want at a less prestigious company. You’ll learn a ton. Pick the best company that will have you.
My personal turning point was when I did free work for a local startup in exchange for them letting me take the Data Scientist title. To a recruiter, it’s totally obvious to hire a data scientist for a data scientist role, and isn’t clear at all what physics has to do with it. Recruiters are the first step when you’re starting fresh, so make it easy for them.
I also somewhat disagree with many of the comments here that textbooks are better than tutorials. If you buy these 1000 page graduate level texts when the idea that you need to read them cover to cover, you’re likely going to give up and fail. Instead, buy the books and put them on the shelf, and then work through tutorials and examples. Then reference particular sections of the book that are relevant to your work to add depth.
Finally, I recommend against starting with deep learning. There’s a whole helluva lot to learn with basic techniques. Very few companies are actually using deep learning in production systems. Start with linear and tree based methods to learn all the stuff about how to frame the problem and build robust systems. Then you’ll have a deeper appreciation for DL.
A reasonable person could disagree and say that there’s so much domain specific stuff around the art of DL that it really behooves you to start there ASAP. I would counter that you’re unlikely to be considered for positions using DL unless you’re pursuing your PhD in it, or have proven yourself in industry. Since that isn’t your situation, I’d wait until you got your foot in the door somewhere and then pursue DL on the side. That’s what I did, and then you look like a hero to your boss. This strategy led to my first publication in the field and I’m now working on DL almost exclusively.
Edit: one more thing. Think carefully about the type of work you want to do. My advice is assuming you’d like to be a person that trains/deploys ML models to solve problems in industry. This is much different than an ML Engineer, who’s implementing algorithms in low level languages and squeezing out efficiency. Obviously that would require a much deeper understanding of SWE. And a totally different person is an academic researcher that’s developing theory or technique. It’ll be hard to do that without a PhD.
>My advice is assuming you’d like to be a person that trains/deploys ML models to solve problems in industry. This is much different than an ML Engineer, who’s implementing algorithms in low level languages and squeezing out efficiency. Obviously that would require a much deeper understanding of SWE. And a totally different person is an academic researcher that’s developing theory or technique. It’ll be hard to do that without a PhD
Can one not only train/deploy ML models, but in addition to that be able to implement the algorithms in low level languages and also be able to develop theory?
I’d imagine these are all skill sets that someone in PhD program could pick up.
If they could do all three, what kind of job should they be looking for?
In the context of a big company, I think it makes sense to have a specialized workforce. Why look for the one in a billion person that can publish top quality theoretic papers and then implement them on distributed gpus in an optimal way while also building simple Random Forest models for your business? I’d rather that person do more of the most valuable thing, and then hire someone else to do the rest.
I suppose my question is more along the lines of, if someone is specializing in deep learning in a PhD program then shouldn’t they at the very least be able to implement models and also know optimization tricks?
In other words shouldn’t they be able to develop enough skills to go deep in one area but also know enough to be dangerous in the other three domains?
I am deep into semiconductors, and am facing the dilemma of giving up my expertise so far, to join a startup as an entry level engineer.
I have done a couple of MOOC specializations and am trying to find projects within my industry to gain some credibility. Also trying to stay active on Kaggle to build some basic data analysis portfolio.
The reason I did this is because there just wasn't enough hours in the day, and my job was taking ~10 hours a day with commuting...etc. It was a risk, but the idea was that I would be able to transition much more quickly if I worked full-time towards it. I also had the financial savings to support myself for 6-9 months and was willing to get a part time job if necessary. Once it became clear that my job's only purpose was to pay the bills in the context of my goals, and I had enough to pay the bills for the near future, it was clear that quitting was the easiest way to free up a lot of time.
This turned out to be the best decision of my career, but YMMV. I doubled my salary in less than 2 years. It's also nice to be part of an industry that isn't so cost-sensitive. I also have a skill set that's in much higher demand, so you can live almost anywhere and there's a ton of companies that want/need it. With semiconductors, you're much more limited.
It's true that you're giving up some expertise and will start in a less senior position in a different field than if you stayed in semi. Sticking around because you have experience is classic Sunk Cost Fallacy. Think 5 years down the line. If you leave now, you'll have 5 years experience in ML. You'll definitely be giving something up if you leave, but there's huge opportunity cost if you don't leave.