Deep Learning Foundations to Stable Diffusion
course.fast.ai
course.fast.ai
We study and implement a lot of papers, including many that came out during the course, which is a great way to get practice and get comfortable with reading deep learning literature.
If you need to get up to speed on the foundations first, you should start with part 1 of the course, which is here: https://course.fast.ai
If you've got any questions about the course, generative modeling, or deep learning in general, feel free to ask!
I came across your APL study session videos while exploring the other material. I used APL professionally for about a decade back in the early 80's. I am always pleasantly surprised when I see interesting work being done in APL. Interestingly enough I always thought APL would eventually evolve into a central language for AI. There were attempts at designing hardware-based APL machines way back when. Of course, much as the language, they were ahead of the technology of the times.
I don't do much with APL these days. I do keep it around for quick calculations and exploration, far more so than doing any projects. In many ways, a thinking tool of sorts.
It was obvious to me on first encounter that APL would never become widespread. Its character set was too abstruse for most programmers and at the time required a special monitor. Plus it seemed hard to maintain if you weren’t either a mathematician or a full-time APL developer.
Your thoughts?
What caused me to hang up my glyphs were the inconsistencies and head scratching behavior of the language in so many corner cases. Working with very nice mentors I found that one could get around them, but because of the legacy of the language they have to be kept around to support existing code bases.
I loved the notion of having a language that gave a first class experience with matrices, though, and after looking around the space, I finally came to Julia and have been very happy.
Thanks for this material.
On the whole, we cover the needed math at the point where we use it in the course. But there's also a bonus lesson that focuses on just the mathematical side, if you're interested: https://youtu.be/mYpjmM7O-30
As in, with the pace of improvements from other AI startups and general availability of their APIs (e.g. GPT-4), is there a specific advantage (aside from maybe cost) to learning to build my own models? Or is the course more suitable for people wanting to become ML engineers (or similar) and to find a job as such? Thanks in advance
The more time you spend on marketing, the better.
We found that as our AI got worse, our product got better.
> We found that as our AI got worse, our product got better.
Interesting. Could you elaborate on this too?
So, specifically, our startup started with the idea that high-brow tech would be a key differentiator and give the best user experience. This is the common story trotted out by survivorship bias stories that make good tech news articles.
Whereas making cuts to our R&D time and focusing on UI/UX and working with simpler science ultimately led to better product.
From consulting, sales, and corporate work, one learns that the dirty secret of big-iron large tech companies is that all stuff sold as ML is just nicely packaged logistic regression. Or it was was five years ago. Nowadays I guess it would MAYBE be transformers, but the point being that off-the-shelf ML with solid data engineering work is what drives 99% of good products. Rarely is it truly innovative tech. I think Pete Skomoroch was the one that joked: "People say I'm a data scientist. I'm actually a data plumber."
I guess one could take this lesson from science. What you learn from a good PhD advisor is: a) read the latest work b) note the simple baseline approach constantly trashed as scoring 2% worse than the sophisticated intricate new things proposed and c) implement the simple baseline. Achieve impact not by hillclimbing on a standard metric but define a new problem or arbitrage insights from adjacent fields, etc.
I’m not smart enough to decode this. I can imagine multiple conflicting answers.
Is this correct? I skimmed the notes for both parts 1 & 2 and they stated requiring PyTorch?
So that means that nothing is mysterious, since we know how it's all made, but we also see where and how to use existing libraries as appropriate.
In practice, during the course we end up using PyTorch for stuff like gradients, matrix multiplication, and convolutions, and our own implementations for a lot of the stuff that's at a higher level than that (e.g. we use our own ResNet, U-net, etc.)
While I appreciate your top-down approach and buy your arguments for doing it this way in previous courses, I also like this new bottom-up where one really learns what lies beneath.
I recommend it. I feel I can now read an arbitrary paper, frown a lot, and eventually understand what it's talking about - to the point where I can implement my own buggy version. And hey, I built my own stable diffusion!!
I found the previous version of this course[1] to be a good complement: it's older (predates SD) but I feel it explains core concepts slightly better. Very understandable given how the close to the bleeding edge this new version is...
Perhaps an even better complement was Karpathy's famous course[2] - similar material but builds towards GPT instead of SD. The fastai coding style is somewhat esoteric (to me) so it was helpful to contrast with Karpathy's more familiar style. I recommend doing both courses. Also I believe the fastai folks are planning a part 3 which covers LLMs; looking forward to that.
Concepts from part 2 helped my hobby project, a 7-day forecast of renewable electricity and power price[3].
Feels pretty great to have built my own Stable Diffusion and GPT! I am grateful to Jeremy and Andrej.
[1] https://course19.fast.ai/part2
One will get much better grasp of DL concepts doing those versions rather than this.
E.g. 2019/20 version IIRC.
I enjoyed how Jeremy and his teaching assistants step through every detail so that you can understand how these fantastically complex systems actually work, from the ground up. Nothing is glossed over or taken for granted.
Much of the course material is building an AI programming framework from scratch. It’s well worth the hassle; I now at least have an intuition about how these systems work, rather than a glossed over impression with many holes. The other thing I greatly appreciated were the insights from other classmates in the forums. Some people way smarter than me took this course and Jeremy would incorporate their flashes of insight into the lectures.
I give this course a 10/10 and I hope my life gets a little easier so that I can get back to the lessons and finish the homework one day.
Is that for part 2? Or also for part 1? How long did it take you, doing it 10 hours per week?
Link for more info: https://www.fast.ai/posts/part2-2023.html
Quite easily I managed to use a pretrained resnet to do some quite astonishing things on completely unrelated pictures we have. And it only took like 10 epochs at ~4 seconds each, so less than a minute of training and it was already quite good.
What I did was that we have taken pictures of things during a process. At that point, I know what the volume should be. So used ~thousand of those images to train a regression model to estimate volume of the thing in the image.
Then I run the model on other images and compare what the model thinks is the volume compared to what it should be. When it's off, the algorithm is mostly correct and we have an error on our side.
Quite astonished about how quick and easy this was to do with the high level api of fastai. Took me 2-3 days.
Or we race ourselves to fulfill the last jobs on earth? Is this progress for humanity or enslavement and the end of our species? I am serious. This question is honest. Maybe my IQ is too low to understand the logic of this innovation.
Maybe I have seen too much tech implementation doing bad stuff to humans, who knows? Please, enlighten me.
The question is why to cut your wrists with a blunt object and stream the event for validation is suddenly a rational behavior?
To be fair I know I'm light years behind the cutting-edge so I feel I can sleep assured that I won't take too much of the blame for having created our new AI overlords, or taking the jobs, etc.
Also, generative art and LLMs are undoubtedly impressive, but not completely my cup of tea. But you can always take your ML knowledge and find a domain/problem you care about to apply it to. It's not like literally everything will be solved in the future.
The folks at OpenAI like Sam Altman are certainly very aware of the potential hazards of LLMs and advanced AI generally, and Elon Musk famously is. The fact is though, these technologies exist and we can't bury our heads in the sand. There is also an enormous upside if we, as a society, marshal the technology and regulate it properly.
I think a degree of nervousness is an appropriate reaction, I'm a little nervous. I'm also excited about the potential. I also see it as a somewhat inevitable evolution of computers. From at least the early 1950s this was "the plan" in a sense. Alan Turing certainly thought so.
But yes, nervousness is warranted. We must tread carefully.
It's not like I could tell Chat GPT "I want a fully featured browser" or "I want a machine learning framework" and it'll just come up with a working Chrome or PyTorch. But those are way, way, way outside your abilities they might as well be magic as far as you're concerned.
Not cool.
When I go to learn something for my career I find myself asking how much time I should actually dedicate to it when a LLM could do so much better.
I want to learn deeper and understand things deeply - but if my time could be spent instead on surface level things and using LLMs to fill in the gaps - it makes it difficult to motivate myself to find time.
I still am trying to learn things deeply, I've had this course bookmarked for a while now and have gone through it a bit.
I guess for me the core fear I have is that the programming field is fast paced. I could learn something deeply, but someone with a surface level understanding and using an LLM can maybe surpass me. Will an employer pick whoever understands something deeply, or whoever has the solution first?
Maybe then the core question here is this: As LLMs improve, is it worth it to learn things deeply for your career, or is it better to gain a surface level understanding and defer to an LLM?
There is also the argument that understanding something deeply will allow you to use an LLM much better than someone with a surface level understanding... but I still have that fear that that may not matter anymore.
I'll still learn things for fun though, so at least I have that!
Assuming that I don't learn is wrong. I have a local installation of SD with a lot of models to test. The only useful thing in this gizmo is the Control Net module or maybe the Photoshop plugin for outpainting. You can upload your linear representation (sketch) and generate. But in essence, this is not progress.
I don't extend my drawing ability or produce any original work by synthesizing all the available artworks. The long term negative effects are obvious. After 3 to 5 years, kids will not bother to draw at all. In my view, this is not progress for humanity, there are a ton of scientific data about the importance of drawing for development of the mind.
The same applies to every human form of creation. You cannot beat the ultimate "calculator" with a model trained on all human intellectual production. Don't get me to start on corporations and greed and how they will view the necessity of human labor.
You enter a room where people are discussing a technical subject, trying to find ways to collaborate and have a better understanding and then you start giving your speech about the dark future of humanity, education, poverty and politics.
Picture yourself in the physical world doing that and imagine the reactions. Why do you think it's acceptable online?
Second answer: I am quite capable, technically speaking, to assess the technology. I don't see the benefits for humanity in A.I. "art" generators. And implementing A.I. outside the narrow use cases, which must be regulated in a form close to how we regulate nuclear energy, is not a good thing. Exponential growth is a reality with this one.
Outside the hype cycle, a lot of specialists are ringing the alarm bell already.
Dismissing their expertise because it represents an obstacle to startup and corporate ROI is not a form of rational thinking.
On the other hand, the signs are clear and finally the lack of empathy and ethics in tech industry will come to fruition.
The original question is a bit "why bother with anything?" so a little hard to answer I think the way it was phrased.
2. Because "all the machines will learn how to create and control themselves and powerful elites will offload the burden of governments to control us?" is a ridiculous sci-fi scenario, and the chatbots aren't nearly as good as so many people pretend they are.
I find the gitlab employees (developers) being excited over copilot to be almost the definition of insane
short term gain, but in the long term it destroys demand for their own product
developers that don't need to be employed don't need a subscription to github
automating your own company/business out of existence
If you like writing algorithms and enjoy the mental problem-solving aspect of it, then you might not like it.
If your main motivation is to protect your job/livelihood and ensure your existing skillset is in demand, then you will probably be worried.
But in any scenario, the cat is out of the bag - you can't un-invent it so might as well get excited and be on the train, rather than be the person that gets left behind.
there are plenty of other ways to deal with it
politically push for AI output to be banned, made un-exploitable or highly taxed
or the luddite approach
time will tell how the several billion people about to be made destitute will react
Society will need to adapt - doesn't necessarily make sense having several billion people doing something that a computer can do quicker, more accurately and more easily, just to keep people in employment (doing something that a lot of people hate).
this argument is extremely similar to that which was used against proposals to ban slavery
Automating jobs does not have the same moral implications as slavery.
That might be a valid argument for keeping slavery if it wasn’t entirely immoral.
But I suspect AI isn’t like the discovery of slavery - it’s more like the discovery of electricity which put candlemakers out of business.
"if we don't have child labour/slavery/no workplace rights/lax environment law/planning restrictions/data privacy law/... then we'll be outcompeted by those that don't"
unregulated AI fits in there perfectly
these persisted for so long because of economics over ethics... maybe we'll make the correct decision this time
(separately: if we do get to AGI then the slavery comparison suddenly gains an extra dimension... as that's literally what its operators will be doing)
But in terms of the first point - the ethics of forcing someone into labour are just not the same as automating peoples work.
Although some people do genuinely enjoy their work, most people work because they are forced to (pay for shelter, food and a good life for their kids). Not forced in the same way as slavery, but there is still mostly a requirement to work in modern society.
It’s one thing asking people to do tasks that are valuable to society, but asking people to do jobs they hate which could be easily automated in the future just for the sake of them “being in a job” seems pointless as a society level to me. Why would we have people doing things they hate that don’t really deliver ‘real’ value because we could automate it and have them do something else? (Preferably something they hate less).
I can read many times faster than listen, and more easily skim back if I realise I don't understand something.
So many tantalizing possibilities for next-gen A/V blending text (toc, transcript, summary, elaboration, explanation, interaction) with massaged video (speed, excerpts, mashups, synthesis) and HID ("hmm, my viewer just frowned and shifted gaze to the text pane, so let's cut to a synthesized reprise of that last bit with the conceptual steps filled in a bit"). So much fun to be had, modulo patents.
Jeremy does an excellent job making the content approachable, but, if you want to go beyond being an excellent practicioner, you cannot skip the theory bits.
Of course there's loads more to do on gaining intuition, and loads more theory to be learned once the course is done! Only so much can fit in a single course :)
It doesn't matter where in the world this is (I like to travel), but does anyone have a recommendation for in person bootcamps on AI?
I feel that while this course is surely superb, I don't want to invest time in a new framwework. What's your opinion?
We've got lightning, fast.ai, and another one that I forgot about... why?
But to answer your question -- the reason IMO to learn the fastai framework (or indeed any software which incorporates a significant amount of novel research) would be to learn about the new ideas that it introduces, in order to become a more effective and informed practitioner. And if you find you like it, you may even decide to use it.
I loved the previous FastAI courses, and the top-down approach taken in them is really fun to go along with.
I guess now it's time for an LLM-themed course?
How much additional effort would it require to incorporate these topics into your course, potentially as supplementary content? This could make the course more accessible to students from countries with less comprehensive educational systems.