Introducing Pytorch for fast.ai
fast.ai
fast.ai
The difference is that Autodesk relies on a mature, deterministic technology (3D graphics rendering). Deep learning is a stochastic process that depends on the data and the model. The training code, and especially the framework hooks, is the least important part. The example they give is three lines of code to train a cat vs. dog classifier. I've tried this classifier on a different binary image classification task: livers with and without tumors. It didn't work very well. There's lots of reasons: little variability between images, grey-scale images, different resolutions, etc. You can tweak the network, throw in more middle layers, try different kinds of layers, whatever, to get better results. All of that is guesswork if you don't understand what the CNN is doing at each stage. At this point in time you do need a formal education in linear algebra, calculus and statistics to investigate why a model does/does not work. It's not enough to know how to use the libraries.
On the flipside, you also need to know how to manipulate data and parse it into the correct format. This generally requires a year or two of programming practice in a good scripting language like Python. I will echo their thoughts that Ian Goodfellow's Deep Learning Book is remarkably lacking in this area. As a simple example, you cannot even use AlexNet without pre-processing your images to be 227x227 or 224x224 for GoogleNet. That's 10,000 images resized, labeled and loaded into the model before training can take place.
tl;dr IMHO in terms of being a competent user of deep learning: mathematics >= programming >>> knowing how to use a framework
It is a question of motivation. If you're reasonably proficient at programming and want to hit the ground running on a specific application (especially in a domain that has well-established methods) fast.ai is probably what you're looking for.
If you're new to programming and mathematics, or want to work on the state of the art of these methods, you need to first truly understand more fundamental ideas. For those who fall into this camp, I am collecting resources (work in progress) and trying to organise them into a learning pathway: https://github.com/hnarayanan/deep-learning
You or anyone else who's interested is free to offer suggestions on learning material.
But there is actually one more problematic area for jpeg: encoding of graphs, drawings etc. with a limited color palette and straight lines. Here, jpeg artifacts become more visible, and to reduce them, you can either turn up the quality, or use a better approach like svg or png. For this, at least a bit of more technical knowledge is required. How many non tech-savy people do even know about svg or png?
But an even more appropriate comparison with ML would be to ask how to improve jpeg to better deal with straight lines. For this, you clearly need to understand the maths.
It's a top-down vs bottom-up teaching approach. The advantages are better motivation, better context and more immediate usefulness.
> All of that is guesswork if you don't understand what the CNN is doing at each stage.
Pff, it's still mostly guesswork even if you do understand what it is doing.
Well, what the CNN is doing at each stage is very simple to understand. There is a forward pass which is a matrix multiply (and addition of there is a bias) and then the matrix weights are learned in the backward pass, which is just basic differentiation and chain rule application. Now I am not trivializing differentiation ( when you try take the differntial of a vector, you are tearing your hair apart) but it's fundamentally a simple concept to understand.
Even with this understanding, designing deep neural nets and tuning hyper-parameters is mostly guesswork. Yes, the frameworks have little or nothing to do with this.
What I've found is that TensorFlow is difficult to wrap your head around for programmers because it's more like a DSL. You declare a computational graph and then run it multiple times. So when you are declaring a computational graph, you have no way of debugging that graph, unless you run it. Also the conversion from numpy arrays to Tensors and back is an expensive operation. PyTorch simplifies this to a great extent. You just create graphs and run like how you run a loop and declare variables in the loop. This is great for imperative programming. However, think about it - every time your graph is recreated. Now if it's just a variable re-initialization it's not a big deal, but we are dealing with Tensors so you give up efficiency for flexibility.
Again, all of this is immaterial for learning how to build deep neural nets. I would say, just stick to whatever framework you can wrap your head around. I am learning that my ability to tweak numpy arrays, visualize them in pyplot, load data from csvs using Pandas and the like will take me a lot further in learning deep learning.
This is just gatekeeping bullshit. Most of the time the right answer is "collect more/better data".
The Nielsen ebook: http://neuralnetworksanddeeplearning.com
It's also because I believe it's in everyone's best interest to have more than one widely used framework controlled by a single company (TensorFlow).
Also, I think fast.ai's approach to teaching deep learning is the right one for the vast majority of developers: start with practical, immediately useful know-how instead of theoretical underpinnings. People who want to delve deeper, say, so they can develop innovative architectures, can always do so at their own pace after taking fast.ai's course. There are a ton of other online resources for learning subjects like linear algebra, multivariate calculus, statistics, probabilistic graphical models, etc.
[1] Here's why I love PyTorch: https://news.ycombinator.com/item?id=14947076
It's not quite as bad as the Javascript npm/gulp/bower/whatever insanity but it's not too far off. Get it together ML people!
Really welcome new libraries and frameworks to make deep learning more accessible. My only fear is that the field becomes cluttered with a myriad of frameworks like the JS world. It would add a lot of confusion and apprehension for people entering the field IMHO since they would not know what to use and where to start.
He also states that PyTorch is hard [1], which does not seem to be HN's overall opinion. So I guess they found Keras was limited in customization and PyTorch required some boilerplate for loading/processing data and training loops, and this new framework tries to fill the gaps?
I'm looking forward for their folloup posts.
[1]: https://twitter.com/jeremyphoward/status/906653539161694208
Have loved your material, thank you very much!
What does that mean, customize models during training?
Also, how are dynamic-graph architectures performing vs. models where the architecture doesn't change? Are they winning competitions?
Would be nice to see use of cluttered mnist instead of coco as an initial toy example for teaching https://github.com/kevinjliang/tf-Faster-RCNN/blob/master/RE...
May I suggest the name "entelechy" (the realization of potential) as a candidate name for the framework.
Kenneth J Hughes (Engineer and Founder, Entelechy Corporation)
Would love to see some benchmarks for that claim.
1) Compilation speed for a jumbo CNN architecture: Tensorflow took 13+ minutes to start training every time network architecture was modified, while PyTorch started training in just over 1 minute.
2) Memory footprint: I was able to fit 30% larger batch size for PyTorch over Tensorflow on Titan X cards. Exact same jumbo CNN architecture.
Both frameworks had major releases since May, so I am sure these metrics might have changed by now. However I ended up adopting PyT for my project.