Cutting Edge Deep Learning for Coders, Part 2
course.fast.ai
course.fast.ai
Let me know if you have any questions about the material or approach. There's also a discussion for the course here: http://forums.fast.ai/c/part2-v2
(BTW, I used to think I didn't like videos for learning, but actually I now think some material works best in this format. E.g. in this case we're walking through interactive notebooks, and you can follow along too. Some of the material is animated, and often I'm building up drawings as I talk. Maybe it won't be for you, but it might be worth giving it a go to see...)
But it's much easier to handle smaller chunks. This is probably my only critique of/gripe about fast.ai---the videos pack a lot of topics in, and they could be broken at topic boundaries to make them easier to use.
Personally, I really hate the Coursera approach of lots of separate short videos - I totally get that some people like it, and that's fine, but it's not something I want to do myself.
- http://forums.fast.ai/t/how-has-your-journey-been-so-far-lea...
- http://forums.fast.ai/t/how-has-your-journey-been-so-far-lea...
Some folks have taken the Udacity flying car and self-driving car courses as well as fast.ai, and had success building on that combination, e.g. http://forums.fast.ai/t/meet-greet-thread-introduce-yourself...
This approach helps ensure that if I haven't explained something clearly enough that I get another chance to do so, and, more importantly, keeps me fresh and energized throughout each lesson (I get stale and boring without some interaction).
So I think this more effective way of learning/teaching even though there is a loss in efficiency.
I do try to limit the time I wait to take a question, since I don't want to move on with a topic where I've failed to properly explain some foundational piece.
So to answer you question, yes, it would be more acceptable if you were saying that to me in person. In fact, I agree with you about the flow of the lectures, and I was looking for someone to bring that up and I'm glad that it's getting discussed here.
However, Jeremy and Rachel are both reading these comments and we should strive to provide thoughtful, fleshed out feedback. The work they've done on Fast.ai is a tremendous lift, and deserves more than a drive-by comment.
I found it to be disappointing especially since they had hyped the collaboration with Siraj, which was nothing more than linking to certain YouTube videos.
The project feedback was sometimes helpful. I felt like most of the time though, the feedback was "you did this wrong, read this article" instead of something more personal like an elaborate explanation on why you should do things a certain way. I even once explained why I initialized a model a certain way and the reviewer ignored it when critiquing my model, which almost felt like "all students have to do it this way."
It wasn't all bad. My favorite parts were learning about GAN's with videos and a notebook from Goodfellow. And when I was trying to build more intuition about CNN's, the videos with Vincent Vanhoucke were helpful.
But altogether I felt a little disappointed in the actual projects. Maybe it was because I felt like the math was glossed over and it was too many topics with shallow exploration for a single course. I actually wished that Udacity offered a single course for say, CNN's and GAN's, going very deep into the math and processes behind them.
I'm taking another Udacity course taught by Thrun (this time, it's free) and again, he kind of glosses over why certain mathematical operations are done, at which point I spent a lot of time watching lectures by other professors who spent more time explaining it.
I think that's my biggest criticism about MOOC's in general, they can be very hand-wavy about very important concepts that underly a process. I've spent a great deal of time reading papers and course material from other colleges, writing throw away code, and watching videos from other profs in order to shore up an intuition that was simply not strongly built by the MOOC.
I think that's what are they going to do next with their AI school. They announced separate nanodegrees for CV, NLP and RL. However, I hope they aren't going to be such massive disappointments as AI ND term 2, where they basically didn't deliver what they promised, castrated projects and one could finish each of term 2 specializations in one weekend (i.e. removing image captioning project, removing real-world NLP project, etc.). A similar story happened with Robotics term 2, where instead of a real robot they promised you received a standard discount for NVidia TX2, and reinforcement learning was butchered from robot walking to robotic arm movement. So I am doubtful they can really live up to their promises with their current staff that keeps underdelivering. The only ND they made absolutely breathtaking, worth every penny, was Self-driving Car ND (IMO the best [not only online] course I've ever taken, and I took many from top 10 universities).
Just wanted to ask a quick question if you recommend any prerequisites before taking this course? I am a second year Computer Science student looking to get into AI, and I was just wondering if this would be a good starting point for me, or should I take some other courses and learn something else before?
Thank you.
We assume you're familiar with python and somewhat familiar with numpy. If you're not sure if you're ready, you can always just start the first lesson and see how it goes. There are plenty of resources on the forum to help fill in specific skill gaps as they come up.
In part 2 of the course I provide quite a bit of advice about how to approach papers in general, and you'll get plenty of practice in reading and implementing papers - but we don't cover the specific math in this particular papers.
My view is the best approach to the math in papers is to generally learn what you need as you get there. It's nearly impossible to know all the math that covers every paper you'll come across, but if you learn the meta-skill of how to learn it on demand, then you'll be just fine! :)
I finished my degree with economics, went to work in finance, and worked my way to being a developer by first doing advanced Excel, then Access, then database servers, then web development, then native development, and now more math heavy research and development. As the need arises, I keep going further down the stack and gaining more understanding. It's much more natural (for me) to learn that way.
But you do need to know abstract math to build stuff, particularly if what you want to build is deep learning stuff, both models and implementations.
For example, how do you expect to understand how to minimize an utility function if you have no idea of what a gradient is, how you calculate it, and why you want to descend through it.
The course teaches all those things - as the comment you're replying to states, you go deeper and deeper during the course to understand all the details.
There's been a lot of research into teaching strategies that shows that this is often a more effective approach for many people than the bottom up approach widely used in math and CS. It doesn't mean that you learn any less of the foundations - just that it's in a different order.
I seriously doubt that anyone can effectively learn linear algebra, multivariate calculus, optimization and regression models from an onlone tutorial on deep learning. These are subjects whose basics alone take multiple semester-long courses. If a bottom-down approach was remotely effective, no one would bother teaching the basics.
This is a purist approach.
Sometimes, you sacrifice the details to broaden the audience. This has the result of getting more people interested.
Somewhat tangential question, but how do you feel about sticking with PyTorch for future lessons/research work? What are your thoughts on the new Swift integration with TensorFlow?
I'm glad you asked about Swift. It's one of the two directions I'm most excited about for the next stage of DL software (the other being javascript!) Swift is a much better language than Python, and having autograd and tensor types built in will be quite transformative. But it has nothing like the same data science library and tooling, so it'll take time to be a real option for most people.
Javascript has the potential to make DL accessible even for those without a discrete GPU, and allows people to get started even without installing a tool-chain. It also allows more work to be pushed to the client - possibly even on mobile devices.
Note this link is for part 2 of the course which is significantly more advanced than part 1. Start with part 1 if if you are new.
I've got part 2 on my to-do list for this year. Looking at the 2018 part 1 course, quite a lot has changed (from Keras to Pytorch?) from a code perspective in comparison to the 2017 part 1 course. If I want to start part 2, does it make sense to go-ahead with the 2018 version? Are there any 2018 part 1 lessons you recommend to bridge any gaps from the 2017 part 1 before I start 2018 part 2?
Thanks!
One thing I liked is the positive attitudes of Jeremy and Rachel, but there was not much theory and a lot of time was spent on questions and answers for the participants of the lecture that I don't think necessary.
I wanted to see definitions and "why", but the course spends too much time on "how". Sometimes I see "why" in the lecture, but many times it is just an answer for the question from a random participant and the answers seemed not well prepared, so it makes it hard to deeply understand the content.
I think fast.ai (at least Part 1) may be good after taking other deep learning courses and before participating in the Kaggle competitions that when you are already familiar with the theories.
I recently went through the videos, of Andrew Ng's Coursera course to get feel for the theory and intuition,and I'll be taking another stab at fast.ai courses to compare his assignments with Andrew Ng's.
In my case, I stopped doing data MOOCs entirely once I realized I learned 10x faster just by reading documentation and working on a project from the bottom-up, since MOOCs rarely highlight the annoying gotchas present in real-world data applications.
On the other hand, with the bottom-up approach of Andrew Ng in deeplearning.ai, you start with a lot of 'why', and later on get to 'how' (although in less detail and fewer best practices than we show). So it's good for people who want to understand the theory right away, and don't mind waiting a bit to understand how to use it.
A lot of our students did Andrew's course after ours, and many did it in the reverse order. All have reported finding the combination more helpful than either on their own. When we describe 'why' it's mainly with code, whereas with Andrew it's mainly with math - so which you prefer will also depend on which notation and framework you're more comfortable with.
(But I promise - you do get all the 'how' with us, particularly in part 2! Our students have gone on to OpenAI, Google Brain, and senior AI leadership positions at well known startups, as well as writing and implementing new papers. Here's an example of a student who just implemented a paper that was released within the last month: https://sgugger.github.io/deep-painterly-harmonization.html#... )
Much of Deep Learning is still experimental in nature and requires quite a bit of educated guessing. A number of times I have been stuck on a particular deep learning problem and a passing comment from one of the fast.ai videos has given me the perfect insight.
I would really like to know how to solve natural language translation. I think everyone would. Many people have been trying to solve this devilishly hard problem for several decades and failed. So I'm really curious how fast.ai has finally manged to do it.
Quite a few new network architectures in this space have been updates to this model, which uses an RNN encoder and a decoder, along with attention between them and a beam search for better results.
I wouldn't call this model a solution for natural language translation, nor would anyone else. But I think fast.ai meant that they're going to explain and go through this model, and how it's helped bring a new generation of models with good performance in this particular space.
By "solve end-to-end problems" I only mean that we show how to do the whole process from beginning to end - I didn't mean to imply that the final model would be human-equivalent or perfect or anything like that.
Yeah, I understood the intention behind that statement. Great work with the course!
While you're here, what do you think about using temporal convolution for sequence tasks? I've read a few articles, this particular one by my professor comes to mind now [0], which say CNNs could work extremely well for the tasks traditionally done with RNNs. A recent paper by the people at Google Brain [1] mentioned that their CNN with attention network beats traditional RNN approaches. More surprising is that the network is 130+ layers deep, and yet trains faster than RNNs. Do you think we can potentially switch most machine translation tasks to CNNs?
[0]: https://towardsdatascience.com/the-fall-of-rnn-lstm-2d1594c7... [1]: https://twitter.com/lmthang/status/989261575482560513
Then why not write just that? What is the point of using language that implies you can teach people how to solve a very hard problem that nobody knows how to solve yet?
I find it extremely disreputable to claim to be able to accomplish feats that go far beyond the limits of current technology. That is the tactic of charlatans and snake oil salesmen, not of scientists and technologists.
Well, the passage I quote, by fast.ai, does exactly that.
It's pretty damn amazing that they build a library (and also the new fastai.text) that goes one level of abstraction beyond PyTorch with the goal of implementing interesting/helpful papers ASAP and because they wanted to make their teaching even more efficient.
I'm always open for ideas in other ways we can help people get involved and stay up to date. And I'm very interested in supporting students who are trying to build these things in their local community.
Completing a MOOC isn't going to help you land an interview for a job, generally speaking (although if it's any good, it should certainly help you once you do get an interview!) However if you use your study time to, for example, create useful projects that you make available on github, web apps that you host on heroku, deep technical blog posts you post on medium, etc then you should definitely get plenty of interview opportunities.
I've seen many many students go through this process, so I know it works. That includes students that hadn't been getting interviews or job offers - until I convinced them to spend time building a portfolio, and then they got multiple offers from top companies.
(There are some companies - not too many thankfully - that strictly require a PhD for DL applicants. I've noticed that such companies, that focus on credentials over actual work output, seem to overlap a lot with companies that turn out to have toxic cultures. So I'm not sure you should worry too much about them...)