It should be possible for a competent software engineer to get up to speed in AI in less than 6 months and much of that time can be on the job itself.
It should be possible for a competent software engineer to get up to speed in AI in less than 6 months and much of that time can be on the job itself.
AI is ill-defined so the premise of your comment makes it difficult to answer. For small well-known tasks (image classification, object detection, sentiment detection) that is train-once on a single dataset and deploy-once what you are saying is true, but for more complex products there is a lot of arcane knowledge that can go in training/deploying/maintaining a model.
On the training side, you need to be able to define the correct metrics, identify bottlenecks in your dataloader, scale to multiple nodes (which is itself a sub-field because distributing a model is not simple) and run evaluation. Throughout the whole thing you have to implement proper dataset versioning (otherwise your evaluation results won't be comparable) and store it in a way that has enough throughput to not bottleneck your training without bankrupting the company (images and videos are not small).
Finally you have a trained model that needs to be deployed, GPU time is expensive so you need to know about compilation techniques/operator fusing, quantization and you need to be able to scale. The requirements to do that are complex because the input data is not always just text.
So yes all the above (and a lot more) require specific expertise.
I know cause I did it.
And I knew the math beforehand. I was a Physics major in college with a CS minor.
I've been in industry and now I do research at a top university. I hand pick the best people from all over the world to be part of my group. They need years under expert guidance, with a lot of reading that's largely unproductive, while being surrounded by others doing the same, in order to become competent.
Writing code is easy. You can learn to use any API in a weekend. That's not what is hard.
What's hard is, what do you do when things don't work. Fine, you tried the top 5 models. They're all ok, but your business requirements need much higher reliability. What do you do now?
This isn't research. But you need a huge amount of experience to understand what you can and cannot do, how to define a problem in a way that is tractable, what problems to avoid and how to avoid them, what approaches cannot possibly work, how to tweak and endless list of parameters, how to know if your model could work if you spent another 100k of compute on it or 100k of data collection, etc.
This is like saying you can learn to give people medical advice in 6 months. Sure, when things are going well, you could handle many patients with a Google search. But the problem is what happens when things go badly.
But becoming an X-Scientist (Data/Applied/Applied Research) is a whole different skill set. Now, this kind of role only exists in a proper ML company. But, just acquiring the Statistics & Linear Algebra 201 level intuition is about 6 months of fulltime study in its own right. You also need to have deep skills in one of the Tabular/Vision/NLP/Robotics areas and get hired into a role accordingly. Usually 1 year intensive masters level is good enough to get your foot in the door, with the more prestigious roles needing about 2 years of intensive work with some track record of State-of-the-art results on 1 occasion.
Then you have proper researchers, and that might be the most impossible to get in field right now. I know kids who have only done hardcore ML since high school, who are entering the industry after their masters or PhD. I would not want to be an entry level researcher right now. You need to have undergrad math-CS dual major level skills just to get started. They're expected to have delivered state-of-the-art results a few times just to be called for an interview. I'd say you need at least 3 years of fulltime effort if you want to pivot into this field from SWE.
I don't think AI is hard to learn. The fundamentals are extremely simple and a competent software engineer can learn all the required concepts in a few months. It's easier if you already have a background in mathematics but not required. If you can write software then you can learn how to write differentiable tensor programs with any of the AI frameworks.
Edit: You asked what it is about these jobs that requires expertise. I answered: it requires expertise to create competitive models. So companies that need competitive models requires expertise.
Edit: Why do you ask? I don't see why it is relevant for the discussion.
Several years ago on HN there was a blog post which (attempted to) answer this question in detail, and I have been unsuccessfully trying to find it for a long time. The extra facts I can remember about it are:
* It was by a fairly well known academic or industry researcher
* It had reddish graphics showing slices of the problem domain stacking up like slices of bread
* It was on HN, either as a submission or in the comments, between 2016 and 2018.
If anybody knows the URL to this post, I would be stoked!
--
I decided to have a crack myself and see what come back.
Here's a few names / blogs that might be useful:
Chris Olah: He's written extensively about deep learning and AI. His blog, colah.github.io, has a unique graphical style that helps explain complex topics.
Distill.pub: This online journal publishes clear and visually engaging articles on machine learning topics. Some of the articles have been discussed on HN.
Andrej Karpathy: Director of AI at Tesla and previously a researcher at OpenAI and Stanford. He's known for his blog, karpathy.github.io, where he delves into various AI topics.
Ian Goodfellow: Known for inventing Generative Adversarial Networks (GANs) and for his deep learning textbook. He might have some writings that match your description.
Ben Recht: A professor at Berkeley who writes about the challenges and misunderstandings in machine learning on his blog, www.argmin.net.
Sebastian Ruder: He has written many articles about NLP and machine learning at ruder.io.
You can also try searching via https://hn.algolia.com/
Knowing the library is the least hard part about ML work just like knowing the web framework is the least hard part about webdev (both imo). It's much more important to understand the actual problem domain and data and get a smooth data pipeline up and running.
Scaling, optimizing inference, squeezing out better performance and annoying labeling. There's a pretty solid gap from applying some framework to a preexisting and never changing dataset vs. curating said dataset in a changing environment. And if we're talking about RL and not just supervised/unsupervised then building a suitable training environment etc. also become quite interesting.
If someone asked me "what's so hard about webdev" my answer would be similar btw...it's fairly easy to set up a reasonably complicated "hello world" project in any given framework but it gets a lot harder when real world issues like different auth worklflows, security, scaling and handling database migrations etc. enter the picture.
If something is already done, i.e. a model is available for your exact use case (which is never), then for using and deploying that can be done by a good SWE and any ML/AI specialist is not needed at all.
To solve any real problem that is novel, you need to know a lot of things. You need to be on top the progress made by reading papers and be a good enough engineer to implement the ideas that you are going to have iff you are creative/a good problem solver.
And to read those papers you need to have solid college level Calculus and Stats.
If this is so easy, then why don't you do it, and get a job at OpenAI/Tesla/etc?