Free “Deep Learning” Textbook by Goodfellow and Bengio Now Finished
facebook.com
facebook.com
My main issue is that the book tells you all about the different parameter tweaks, but passes little concrete wisdom to the reader. It doesn't distinguish between modeling assumptions, and it replaces very simple explanations of concepts with complicated paragraphs that I can't make sense of.
I think it boils down to something that I have been feeling and hearing a lot in the past few years: the statistical jargon is so overwhelming that the authors can't explain things clearly. I can point to many examples in this book that I feel are unnecessary stumbling blocks, but the fact is that I'll spend an hour or two discussing parts of this book with a room full of smart machine learning researchers, and at the end we'll all agree we don't understand the material better than we did at the start.
On the other hand, I'll read research papers that don't force the statistical perspective down the reader's throat (e.g. http://arxiv.org/abs/1602.04485v1) and find them very easy to understand by comparison.
It might be a cultural difference, but I've heard this complaint enough from experts who straddle both sides of the computational/statistical machine learning divide that I don't think it's just me.
For those that may not know, j2kun also writes an on-going expository machine learning series on his mathematics blog[1].
[1] http://jeremykun.com/2012/08/04/machine-learning-introductio...
It is good to see critical views but it will be even better if you could give concrete examples for statements like the above. Also, what other books do you recommend?
This has been my frustration. I've been wanting to grok deep learning, but haven't found any source that can explain it in a way that doesn't overcomplicate simple things (which I can only tell it's doing when I already understand the topic). I also don't have the time or incentive to really dig in and do the math myself from scratch, since so much of it is wide open research directions.
I've also had this experience with much much simpler areas of statistics and statistical ML (cf. Markov chain monte carlo), so this is a sort of recurring theme for me. Considering how simple MCMC is now that I do understand it, it's difficult to dispel the nagging feeling that all the statistical ML literature is (likely unintentionally) obfuscated.
I could give concrete examples of excerpts and entire sections of the book that don't make sense to me, but I don't think it's all that productive because a lot of it boils down to organizational disagreements, cultural assumptions behind the math, and differences in priorities. Individually it seems like nitpicking, but they add up quickly to general muddled confusion. This is especially true when, for example, almost always the right answer to the question "Why does this particular technique work?" is, "We have no clue, but here is some anecdotal evidence and half-substantiated oversimplified theories." Instead these answers are passed as well-known fact.
I think both explain statistical ML clearly at just the right level of detail. But I'm not an advanced math person, so it's certainly possible that I've failed to notice any flaws in their approach.
[1] https://web.stanford.edu/~hastie/local.ftp/Springer/OLD/ESLI...
Too many people are willing to accept (and defend!) sub-par explanations and overcome them through titanic mental grind. The scary part of it is that often this leaves you with a broken mental model that continues to require tons of effort in application. Meanwhile, much better explanations exist. They make learning easy and fun, while also giving you intuition that is easy to apply and even expand to other areas.
I am pretty good at spotting bad mental models, but it's really hard to prove that they are bad, until you find one clearly superior. Recently I stumbled upon MIT's linear algebra class at Open Courseware and was blown away by how much easier it was to follow than the stuff I had in college, while covering the same material in much greater detail.
It necessarily shorter on detail in terms of the tricks of implementation that have radically improved performance of these techniques over the past 5 years, and might serve as a good read before diving into this via Goodfellow, Bengio, and Courville.
On the other hand, if we find the cited examples clear due to e.g. different backgrounds, expectations, perspectives, etc, then that might be a good signal that we should look into the book.
Even if the material is already 'covered' elsewhere, simply re-explaining the same concepts would do a world of good for those who think in the same way as you.
If this book is trying to do more to bring statistical or probabilistic insights to bear on deep learning than I think that's a very good thing. It might make it less accessible to those coming from a pure computer science background, but potentially more so to those who like to think about machine learning from a probabilistic modelling perspective.
If they're using stats jargon in a gratuitous way that doesn't actually cast any light on the material then that's another thing, but from a quick skim I didn't see anything particularly bad on this front. Do you have any examples of the kind of jargon you're talking about?
To others reading, I just wanted to emphasise that statistics is really important in machine learning! Deep learning lets you get away with less of it than you might need elsewhere, but that doesn't mean one can treat it as an unnecessary inconvenience. It's a language you need to learn, especially if you want to try and get to the bottom of how and why aspects of deep learning work the way they do. As opposed to just an empirical "using GPU clusters to throw lots of clever shit at the wall and see what sticks" engineering field. Bengio seems very interested in these kinds of questions and I'm glad he's leading research in that direction, even if clear answers and intuition aren't always easy to come by at this point.
From a mathematical perspective, the book does not define hardly anything. Take, for example, the important chapter on Deep Feedforward Networks. They hem and haw saying things like "quintessential deep learning models" but never actually explain _what a deep feedforward network is_. There are no definitions: they give what amounts to preamble of a definition, and then move right on to using the notion, but don't actually give the definition itself. In the deep feedforward network example, the closest thing to a definition is, "A feedforward network defines a mapping y=f(x;theta)". That doesn't really say anything though. If you're asked to define "dog", you can say, "A dog has four legs", but how does the reader know that a cat isn't a dog?
I find that for basic machine learning it's much better to just go through machine learning framework tutorials. For advanced learning, I'm not there yet so I can't speak from experience but I'll be shocked if this book fits well in any role except maybe last-resort reference material. If that.
The fact that beating an accuracy test on some arbitrary data set is seen as more important than understanding why and when your methods works (or do not work) is deeply disturbing.
Or alternatively, reach comparable level of validation accuracy with significantly less compute or memory usage (computational complexity).
Or alternatively, reach comparable level of validation accuracy with only a subset of labeled samples + unlabeled samples (sample complexity).
I found http://neuralnetworksanddeeplearning.com/ far more accessible.
Journal papers are the pinnacle of unclear, convoluted context depend writing I know of.
This is an extremely damning critique.
1. It also covers "classical" artificial neural networks, i.e., things like backprop from before Hinton and others made breakthroughs for deep learning. This means you can start with this book even if you are new to ANNs. The later sections cover "real deep learning".
2. The language is great for beginners and users. You don't have to be an advanced math geek to follow everything. They seem to cover a fair amount of ground too, so its not dumbed down either.
3. I guess it covers most of the underlying theory and practical technicques but is implementation neutral. You should probably pick up a tutorial for your favorite implementation like Theano, TensorFlow, etc.
All in all, I like it a lot.
Another great great free online book on this topic: http://neuralnetworksanddeeplearning.com/
pdftk $(ls -tr *.pdf) cat output DeepLearningBook.pdf
Out of respect for the authors' contract, I won't post the resulting file here, but anyone can reproduce it with about 5 minutes of work.Actually, they were mostly just very hostile and unpleasant. This is among the many reasons why we're probably staying independent and printing the book ourselves - so that the digital version stays unencumbered among other things.
Preview: http://i.imgur.com/5OXv28R.png
Call me evil but I did it:p
Same course if you prefer the classroom lectures http://ocw.mit.edu/courses/electrical-engineering-and-comput...
Or if you want more rigor you can go through these notes that cover the same material but in a more formal way (via sigma algebras and measure theory) http://ocw.mit.edu/courses/electrical-engineering-and-comput...
1) http://ocw.mit.edu/courses/mathematics/18-06-linear-algebra-... 2) https://www.khanacademy.org/math/linear-algebra/vectors_and_...
If anyone knows anything else (relevant to deep learning) could you please share :)
http://www-bcf.usc.edu/~gareth/ISL/
Is an excellent statistical learning reference.
https://www.youtube.com/playlist?list=PL5102DFDC6790F3D0
for a basic "Stats 101" course.
There's also this archived Coursera course. There aren't any active sections to sign up for, but the videos are still available:
* All of Statistics
* Doing Bayesian Data Analysis
Also the ML specialization on Coursera
For linear algebra I heard Paul Dawkins' document is great, it's not on his site anymore but you can find it online. I've read calculus 1 and 1/3 of calculus 2 and it's good material.
ESL ISLR doing Bayesian data analysis w/ jags/Stan bda3 - gelman prob graphical models Convex analysis - Boyd adv data analysis from elem pov - shalizi
Trying to build out my library. I have a background in prob/stats/analysis and measure theory/linear algebra and also knowledge of algorithms and data structures at the advanced undergrad level, So I'm not too concerned about technical depth just want to enjoy a good technical expository and gain intuition.
Related to printing - this made me chuckle:
> Printing seems to work best printing directly from the browser, using Chrome. Other browsers do not work as well. In particular, the Edge browser displays the "does not equal" sign as the "equals" sign in some cases.
Of all the printing bugs for a maths/logic heavy text! Can just visualise the head banging against the desk upon discovering this one, having struggled with understanding something for hours - well ok, either the universe is broken or... oohhhhhhh :-D
My favorite aspect of this book is that it provides a graphical models interpretation of DL methods, which is the most powerful perspective we have right now to reason about model design (instead of some large black box function that we train end-to-end without knowing what's in between).
It also explains some fairly recent models and techniques well (VAE, DCGAN, regularization) that form the basis of more complex architectures. If you understand the models here, you should be able to understand the design choices made in more complex architectures.
Thanks to Goodfellow, Bengio, and Courville for this excellent work.
It kinda looks like someone ran the original PDF through PDF.js and saved the rendered output to a HTML file.
<!-- Created by pdf2htmlEX (https://github.com/coolwanglu/pdf2htmlex) -->
"This format is a sort of weak DRM required by our contract
with MIT Press. It's intended to discourage unauthorized
copying/editing of the book. Unfortunately, the conversion
from PDF to HTML is not perfect, and some things like
subscript expressions do not render correctly. If you have a
suggestion for a better way of making the book available to
a wide audience while preventing unauthorized copies, please
let us know."The Concrete Mathematics book that resulted was a masterpiece in my opinion. I learned so much about how to computationally do discrete math from that book. And it's a very elegant package.
[Answer] Yep, Euler: https://en.wikipedia.org/wiki/AMS_Euler
Just submit a screenshot image of some glyphs of the font at a reasonable size, and let the service figure out the font!
Waste of time!!
mkdir dlbook;cd dlbook;wget --recursive --level=1 http://www.deeplearningbook.org/
cd www.deeplearningbook.org/contents
python
import pdfkit
pdfkit.from_file(
["TOC.html","acknowledgements.html","notation.html","intro.html","part_basics.html","linear_algebra.html","prob.html","numerical.html","ml.html","part_practical.html","mlp.html","regularization.html","optimization.html","convnets.html","rnn.html","guidelines.html","applications.html","part_research.html","linear_factors.html","autoencoders.html","representation.html","graphical_models.html","monte_carlo.html","partition.html","inference.html","generative_models.html","bib.html","index-.html"],
"Goodfellow-et-al-2016-Book.pdf")
Better solutions?edit: changed ".pdf" from a slightly longer approach to ".html" which actually exists in this workflow. Thanks @TheCabin!
edit(2): ... and gave a valid path to this listdir. Check before I post... check before I post...
edit(3): ... and removed the os.listdir line no longer needed in this approach. Gosh. Just ignore what I was saying and build your own approach. That'll probably be faster at this rate.
edit(4): Don't import os without using it.
Avoids the awkward file system structure by using pdfkit.from_url, but creates one .pdf for each chapter. I tried using a list of urls, but pdfkit failed because my version of wkhtmltopdf did not accept multiple input files.
Edit: pdfkit.from_file also fails on my system when passing multiple files. If that works for you, multiple urls are probably fine, too.
I think it was a mistake to try shaving off a couple of lines and a step at the price of a more convoluted install and a brittle process. Most people would probably be best off just converting one HTML file at a time to pdf (e.g. pdfkit.from_file(filename_in_html,filename_in_html[:-4]+".pdf") in some sort of iteration over the file names) and then concatenating the resulting PDF:s, for instance from command line by
pdftk TOC.pdf acknowledgements.pdf notation.pdf intro.pdf part_basics.pdf linear_algebra.pdf prob.pdf numerical.pdf ml.pdf part_practical.pdf mlp.pdf regularization.pdf optimization.pdf convnets.pdf rnn.pdf guidelines.pdf applications.pdf part_research.pdf linear_factors.pdf autoencoders.pdf representation.pdf graphical_models.pdf monte_carlo.pdf partition.pdf inference.pdf generative_models.pdf bib.pdf index-.pdf cat output Goodfellow-et-al-2016-Book.pdf
Sorry if I contributed to wasting your time.If I want this book in hard copy, then I will purchase it - I've done this regularly with free digital books - but when it is offered free digitally then in my opinion prohibiting to only certain file formats is futile (as evidenced here), and such constraints are ineffective attempts to encourage people to buy the hard copy through inconvenience.
And I must add that this is no slight to the authors, whom have my greatest appreciation for compiling their vast knowledge into a book and offering it for free. These guys are legends.
find . -iname '*.html' -exec wkhtmltopdf {} {}.pdf \;
There are also tools to convert the resulting files to a single pdf. The only problem I got is, that the woff fonts are not rendered by "wkhtmltopdf" :-/ Ideas?
For the lazy: http://www.deeplearningbook.org/front_matter.pdf
Personally, I was looking for an ePub version but no biggie.