From Deep Learning Foundations to Stable Diffusion
fast.ai
fast.ai
Anyone aware of an in-depth intro-level text-based explanation of Stable Diffusion that covers the whole pipeline, including training on an extremely limited dataset?
Here’s an example, but open to suggestions too:
Of course you can start with a more traditional course and then learn something like stable diffusion afterwards, but as a newbie it’s quite hard to figure out where to even start. A full-fledged course that takes you exactly where you want to go is a lot easier and I think it can help learners to stay motivated because they have a clear goal in mind. If I want to learn how to create cool images, I want to spend as little time as possible predicting housing prices in the Bay Area.
I think that's somewhat of a dangerous mindset to have. If you want to create cool images you can use pre-trained models and high-level APIs without needing to understand any of the internals.
But if you want to truly understand how these models work, you need to make effort to study the basics. Maybe not predicting housing prices, but learn the foundational math and primitives behind all of the components from the ground up (and the Diffusion models are a complex beast made up of many components). And getting an intuitive understanding of how models behave when you tune certain knobs takes much longer. Many researchers in the field have spent years developing their intuition of what works and what doesn't.
Both of these are fine, but I think I think we should stop encouraging people to be in the middle. Have courses that that promise "Learn Deep Learning / Transformers / Diffusion models in 7 days!" but then go on and teach you how to call blackbox APIs, giving you an illusion of knowledge and understanding where there is none. I don't know if this applies to this specific course, but there are a bunch of those out there, and highly recommend staying away from those. I know it's a hard sell in this modern instant gratification age, but if you actually want to understand something you need to put in some possibly hard work.
fast.ai do stuff pretty well. FWIW, I did one of their earlier free courses and, as a maths grad, got my fill of maths learning as well as my fill of practical 'doing stuff with ML' stuff. If I didn't have my plate full I'd probably pay the 500 quid or whatever to do this course now rather than wait for the free version.
I was very confused by this in the beginning of my journey. I was trying to learn everything involved with ML/DL, but in the end everything is already implemented with APIs, and your boss doesnt care if you know how to implement a MLP from scratch or if you use Tensorflow.
My (poor) analogy is: you don't need to know how a car works (or how to build one) in every detail to drive it. When I understood it, it was liberating.
What's not so great is the huge number of people believing to understand something when they don't, i.e. the illusion of knowledge they're getting from some of these marketing-driven courses and MOOCs. I see that in job applications. Every resume has "Deep Learning, PyTorch, Tensorflow" on it now, but if you ask them why something works (or why a variation may not work) these candidates have no idea. And for some jobs that's totally fine, but for other jobs it's not. And the problem is when you can't tell the difference.
It's kind of like putting "compilers" on your resume because you've managed to run gcc.
Interesting. Do you have an example? Is it common for people to practice ML problem-solving à la LeetCode nowadays?
I'm just making these up because the questions we previously asked were domain-specific to our applications, e.g. "why is this specific learning objective hard" or "what would you modify to help generalization in case X"
These questions are very easy to talk about for someone with a strong ML background. They may not always know the answer and often there is no right answer, but they can make reasonable guesses and have a thought process around it. Someone who just took a MOOC likely has no idea how to even approach the question.
> fast.ai
> Do that and your life will change
Sounds like Emad Mostaque of Stability AI / stable diffusion thinks this course probably won't fall into "do this, no understanding needed" trap (I'm not contradicting anything you said here).
Originally took a break from being a hedge fund manager to build AI lit review systems to investigate ASD etiology for my son along with neurotransmitter pathway analysis to repurpose medication (with medical oversight) to help ameliorate his more severe symptoms.
Had 100% satisfaction from programmers with some math knowledge trying fast.ai and members of team active there, really nice take off point into a massive sector.
It digs nicely into the principles and is not a surface level course. The stable diffusion one will need some good work to get through.
But yeah my job now is to get billions of dollars into open source AI to make the world happier, happy to do my best and let the smart and diligent folk buidl.
Do you know other efforts in that direction?
Will be aggressively investing in this area and making the output available openly next year after our education launch.
If there's a way to contact you (Sharing similar challenge) I'd be happy.
I would also highly recommend FastAI’s Deep Learning for Coders (and their new course that came out this year). You’ll start immediately with some cool applications (basic image recognition and NLP) and then drill down from there to learn how they work in detail.
It’s set up such that you can learn as much as you want (basics with no depth: first chapter; basic understanding of how a neural network is trained with SGD: first four chapters; understanding of decision trees, LSTMs, and CNNs: first half; detailed understanding of how to build everything from scratch: whole book).
I haven't taken the course, but that sounds like a horrible place to start a course on understanding deep learning. GPU matrix operations are literally an implementation detail.
I think the proper way to teach deep learning "from scratch" would be:
1.) show simple example of regression using high level library
2.) implement same regression by writing a simple neutral network from scratch (explain going through, multiplying weights, adding biases, applying activation function, calculating loss, back propagation).
3. use NN on more complicated problem with more parameters and a larger training set, so that user sees they've hit a wall in performance, and now implementation needs to be optimized
4. at this point, say okay, our implementation of looping and multiplying can be done much faster with matrix multiplication on GPU, and even faster parallelized across GPUs on a network. If you're interested in that, here is an optional fork in the course that gets into specifics. Anything after this point will assume that implementation of NN calls will be using these techniques under the hood
5. move onto classification, q learning, GANs, transformers.
95% should have skipped step 4 and only revisited if they become interested in this for a specific reason. To start with it is crazy. It's like starting a course about flying by explaining how certain composites allowed us to transition from propellers to jets, and let's dive into how those composites are made.
Sounds like you support the course author's decision to make this part 2 of the course series to be taken after part 1 is completed!
> Three years ago we pioneered Deep Learning from the Foundations, an in depth course that started right from the foundations—implementing and GPU-optimising matrix multiplications and initialisations
They're talking about how Part 1 starts
You should consider taking a look at part 1 of the course if you want to understand the starting curriculum: https://course.fast.ai/
I keep turning to easy distractions like twitter and Pac-Man.
That helps them have something they can use right away, which helps with a few things:
- at each milestone, they have a distinct goal that’s reachable
- they can understand concrete use cases immediately
- they get a little dopamine hit if satisfaction of having completed something.
Whenever you’re doing a course that doesn’t structure itself that way, it’s good to try and break down the components and set yourself little tasks as you go.
You can only switch to something that it doesn't find boring, even if the final result is the same.