Practical Deep Learning for Coders 2022
fast.ai
fast.ai
It wasn't until the last few years I saw people start freezing the model and just fine tuning the last layers. I've watched presenters from flamingo and imagen talk about their similar approaches. I heard it here first, at fastai.
fast.ai is a fantastic educational resource and a great way to approach solving problems. But the library itself is lacking, and if you are an experienced programmer, when building real-life projects, you will be frustrated with fast.ai library.
The goal, IMO, should be learn from Jeremy Howard, s great instructor, communicator; learn his attitude, and then move to PyTorch (keeping the attitude, the knowledge, and the lessons with you.)
It's also the only library I know of that consistently bakes in best practices like super convergence techniques or making things like test time augmentation very seamless. Many libraries lag behind fastai 1-2 years in this regards, and frankly it can be frustrating to use other frameworks sometimes.
There is a slight learning curve, for example to learn the DataBlocks API or the callback system, but once you really understand what is happening you will understand how nice the API is and how well engineered it is.
Side note: Regarding being an experienced software engineer, I highly recommend digging into how the python language was extended for this project (fastcore) and the development workflow used (nbdev), which I think could be interesting for those software engineers you mention as well as heighten your understanding of the ecosystem of tools.
What is your take on the current state of autonomous driving? Do you think we can achieve "full autonomy" with the technology we have currently?
Any new advances in DL that you are excited about?
The new advances in DL I'm excited about are things I show in the class: the accessibility of modern NLP thanks to the Hugging Face ecosystem; the power of ConvNeXt for even better computer vision models; the way Gradio and HF Spaces makes it trivially easy to get a working prototype application using DL online.
I'm also excited about hosted models and applications like GPT-3, DALL-E, and Codex. All the illustrations on our course website are from DALL-E, for instance!
In some ways Jax is almost "non-python deep learning" since it's treating Python more like a DSL for the XLA backend. Normal Python code doesn't work in Jax. It's a pretty reasonable compromise since you still get all the benefits of the Python ecosystem.
Julia seems like it has the best foundations for deep learning, since everything can be written directly in Julia. But it doesn't have a great ecosystem as a general programming tool.
F# might turn out to be a good option.
I have a search problem of my own and I have had a hard time applying what I have learnt (including the coursera DL specialization). The chief characteristics are: (a) It is a fuzzy search of a corpus that is in a non-English language. (b) The search should be able to run on a mobile phone _offline_.
Is this possible? Can training be done elsewhere and transferred to TinyML or some such? What would be a good forum to go seeking answers?
In my case, I don't want the model to be general. I can afford for it to be like a database index, tailored to that data.
a) BM25 after some preprocessing (lemmatization etc.)
b) fastText / GloVe (possibly weighted by BM25)
The results can be surprisingly good. Often no need to bother with big language models or GPUs.
For the same reason, GloVe is of not much use to me.
How far away is the fast.ai from working on a Mac? PyTorch recently gained support (https://pytorch.org/blog/introducing-accelerated-pytorch-tra...) but that's only the start. Is this something that is being worked on?
Mac support for all the libs used in the course will probably continue to improve in the coming months and there should be no reason you won't be able to run the stuff for the course locally on a Mac at that time. Having said that, even the M2 trains deep learning models much slower than even the free NVIDIA GPUs provided by Kaggle. So you'd only want to use local development for the smallest and simplest models. (The course shows how to train models that are fairly cutting edge and some take a while to train even on modern GPUs, so they wouldn't be a good fit for a Mac.)
There will always be smart individuals and talented small teams that can successfully integrate AI into their products, but it's not thanks to the courses like this.
If you're going to aim at coders, there has to be clear path demostrated from the beginning to the end. From starting up your first notebook on your local dev machine and running training on your local training machine to setting up inference in the final app (.net app or whatever)
How easy is it to do the course on my own hardware rather than cloud notebooks? Would that make it closer to practical deployment?
Re the course, I just skimmed it and I think you can do most things on your own hardware but if you will actually use this for something practical (not just for you or a side project), being familiar with cloud tools is a big thing especially once you scale.
I disagree with needing none and just going along as needed. That’s how you have machine learning models that look like they work but you don’t understand why they work so there might actually be problems.
Most practical deployment is done to cloud environments rather than local notebooks. The deployment exercise we do in the course is designed to show the key components you'll need for deploying simple models in practice.
I have one question, and one only. Please answer:
Second part, when?
.b EnThe amazing thing about these courses is how simple Jeremy (and team) are able to make machine learning. I didn't need to understand python dependency management in order to learn how to train an really good image classifier. Their approach helped me have lots of little wins, gave me confidence, and helped build the motivation to slog through the harder stuff when I needed to.
From the bottom of my heart, thank you @jph00. You changed my life immeasurably for the better. I learned that I AM good enough, I AM smart enough, and I CAN do hard things... I just had to find the right way to learn them. Your courses completely changed my perspective on what was possible for me and opened the door to some of my life's greatest passions.
I was studying a masters in statistics and computer science that had 1 neural networks lecture and nobody knew anything about deep learning. Fast.ai and Jeremy’s teaching style helped me start playing with deep learning models really quickly and I changed my thesis topic to computer vision.
I ended up consulting on the topic and doing various startups leading to the startup I’m working on now which just finished YC (AiSupervision W22).
I doubt be here without fast.ai. I highly recommend and appreciate all the work that Jeremy and the rest of fast.ai do!
Even DL used for audio processing (classification, separation etc) seems to convert audio to spectral graphs and apply DL to that.
Changing a problem to be expressed as image inputs will be an advantage when using DL as a solution. Would you agree ?
I think a major reason for this is because of transfer learning. For computer vision, there are many good pretrained models that were trained on huge datasets (like ImageNet) that can be fine-tuned for custom tasks. Other fields often do not have such pretrained models and huge datasets to work on, so it turns out transforming a dataset into an image dataset and fine-tuning a pretrained model works better than training from scratch.
Take convolutional models, for example. Very effective for working with images because they're (a) parameter efficient, (b) learn local/spatial correlations in input features, and (c) exploit translational invariance. As an oversimplification, we can train models to visually identify "things" in images by their edges.
If you think about what's going on with an audio spectrogram, you can see the same concepts at work. There's local/spatial correlation - certain sounds tend to have similar power coefficients in similar frequency buckets. These are also correlated in time (because the pitch envelope of the word "yes" tends to have the same shape), and convolutional models can also exploit time-invariance (in the sense that convolutional models can learn the word "yes" from samples where the word appears with varying amounts of silence to the left and right).
That being said, the addition of the time domain makes audio quite hard to work with, and (usually) not as simple as just running a spectrogram through a vanilla image classification model. But it's definitely enlightening to think about how these models are "learning".
Your comment about time domain making audio difficult - before doing some research I thought it would make it impossible. But looks like people have had some success with using spectrograms of short audio samples. What techniques should I try to learn to deal with the time component of audio?
One idea is to chop up the audio into short samples and treat the resulting images as a video. Then look at DL algorithms that deal with video. Am I on the right track?
Extremely grateful to have found it. Changed the course of my life.
I can vouch for it's quality.
Jeremy is an excellent instructor. So much clarity in his teaching!
I love that this is a hands-on course, and there are ZERO hand-wavings. I also really like the top-down approach of teaching. Now, whenever I try to communicate something or teach someone, I try to do it top-down. And I have Jeremy to thank for that.
Currently, I am attending his APL study group and having a blast!
Only question for @jph00 is: second part, when?
Now moving on from here, do you have any resource recommendation where I can dive deeper into machine learning and deep learning theory? And also any resources to become a much much better programmer?
I am currently working in as an assistant in a research lab. My coding skills are not that great.
Or you could do those bits at 2x speed in case there's some concepts there you haven't seen before.
The NLP lesson could possibly work reasonably well standalone if you already know some DL basics, since it uses a different framework (Hugging Face) to the earlier lessons.
The CNN lesson would probably largely make sense if you already understand multi-layer perceptrons, since it mainly shows how a convolution is just a special case of sparse matrix multiplication.
Most ML advancements will use DL at the core in interesting ways.
Windows not supported on AMD cards Navi series cards not supported in general. Heavily biased towards CUDA, despite AMD cards drivers being open sourced far more than Nvidia cards.
Remote machines and kaggle notebooks go some way to improving these limits for the course. I'm complaining a bit more in general here, I think.
One should invest too much time just for the sake of learning the library's weird API, and then using it.
Doing something custom is too difficult, in contrast to Jax, PyTorch, and even (poor library) TensorFlow.
The coding practices are whimsical. The codebase wouldn’t pass code review in any respectable company.
Variable namings are weird and super-problematic.
I fully stick to what I said. Learn techniques, best practices, and, most importantly, Howard's attitude. Then take them with you and move onto something like PyTorch.
Howard is great with one problem: he kinda hates math. It might also seem that he ends up promoting anti-intellectualism.
Can you explain this last sentence (which I understand to be insulting and without basis): > Howard is great with one problem: he kinda hates math. It might also seem that he ends up promoting anti-intellectualism.
This is not insulting. That man is my hero, and I deeply respect him.
But his 2019/20 course was riddled with such statements. He repeatedly said that one doesn't need math, and showed tools like drawing math symbols on a website to learn their names and ride on that. No further math needed.
It's like you can wing it in Deep Learning without learning Math. His behavior throughout the course reinforced this attitude. It is harmful for new learners.
But I am fortunate that I didn't learn from that, but learned from some successful alumni example that Howard gave. One woman who was also a musician ('19/'20), she made it big, but Howard mentioned that she did the Ng course, and also read the Goodfellow book.
So, I took the cue, and did DL the proper way. Anybody I know in DL made it because they know the Math.
There are some influencer types in fastai community who has 10ks of followers and shills stuff and do media stuff. Other than that 1-2 people, everyone who made it in DL, did it because they knew the math.
So, I think that people might get the wrong idea hearing from Howard that "you don't need math".
This is one fault I find. It's not like I dislike him. I like the rest of him. I love his attitude on almost all other things. I love Jeremy Howard, and he is my hero.
You don't need the math in the beginning to train a model and get first results. Later, you will need the math and Jeremy clearly knows the math.
He gives a great example: In sports, you don't start with learning about physiology and train individual muscles etc. (I paraphrase), you start playing basketball or baseball or soccer, and understand the overall game. And if you like it, you can then become better and better and get deeper and deeper.
It's not helpful to start with linear algebra if - what motivated you - was the application of ML. We lose people who could have otherwise become experts later.
Yeah, I know.
But you need a lot of math to do Deep Learning.
But I do not think Howard tries to communicate that.
You can't show me people who knows high school math only and gets to work in FAANG, or PhD in DL/related, or CTO of an AI start-up, or anyhow "made it" in DL.
In the courses he has always been clear you don't need a ton of math to begin. He's also always been clear that as you progress you will encounter math that you need to learn to continue. He's always clear that that is ok if you don't know it before you start and it's ok to learn it when you need it.
But also use ISLR, Goodfellow, Bishop, etc.
Start with Andrew Ng's ML, then do the first part of Aurelien Geron book, then do Ng's DL specialization, then do fast.ai. Then learn PyTorch. A great book would be Sebastian Raschka's book. Also d2l.ai. A fast-paced, but really good course would be the Neuromatch DL tutorials.
Then move forward based on your interests.
Yann LeCun has THE best MOOC on DL on YouTube.
For the Math, I majored in Physics, so stuff came naturally. I suggest Imperial London's MOOC on Mathematics for ML specialization, Robert Ghrist's Calculus course, VMLS book for Linear Algebra. For stats, haven’t found a good one yet.
What you read, how much- these all depend on what you want to do. Where do you want to see yourself, and so on.
If you just want to brag about DL and put it on your resume so that you can get a job writing SQL queries and make PowerBI presentation as a "Data Scientist", then the bars are low.
If you want to do some DL, then that is another league altogether.
You need to be able to quickly read papers, understand ideas, use those for your own projects or papers.
Makes sense?
Everyone that I can see that fits that profile work at real companies doing real deep learning work, or are building infrastructure and tools that we all use. Nvidia, Huggingface, Etc. I don't see pure media stuff at all, most people are developing libraries or doing other applied work, and talk about their work publicly. Frankly, your comments come across like you are salty. Being a unpleasant person in online forums that enjoys insulting people seems correlated, which likely doesn't bode well for your professional aspirations, regardless of how much math or python you do/don't know.
> You don't need math
He's saying you don't need a PhD in math, not that you should ignore math all together. I have graduate level math and CS background and I don't thing either of those helped much, other than overcoming gatekeeping. The thing thats far more important for applied ML is to practice DL on lots of different problems to be effective. PhD level math might be useful for research, but that isn't necessary in practice for most people.
I'm sorry what?
I run a math study group 4x per week.
Right now the book I'm reading during my rest time is a calculus book.
I've co-authored a lengthy paper on matrix calculus foundations for deep learning.
I wrote a lot of the math materials in our numerical linear programming course.
It really seems like you have very very little understanding of me or the software library I've created, but yet are nonetheless comfortable publicly pronouncing your opinions about both.
We detached this subthread from https://news.ycombinator.com/item?id=32189308.