The other thing is this seems to be very CNN heavy. Four lectures on the topic seems like a lot.
Also, I don’t see embeddings explicitly mentioned as a topic. They’re a huge component of industrial research, and creating good embeddings and retrieving them quickly is a topic I feel students should also be exposed to. (Yes, they mention “representation” with autoencoders but quite frankly the code bit is generally not useful for similarity metrics.)
Finally, it would be nice to expose students to multimodal learning. Something like CLIP would be pretty neat to expose students to. It’s a great insight when you realize that you can train projections of multiple modalities into a shared high dimensional space. If they’re going to cover diffusion models certainly complexity isn’t a concern.