Understanding Convolutions (2014)
colah.github.io
colah.github.io
In the book, it walks you through the convolution of an input matrix with a smaller filter matrix. You start with lining up the smaller matrix in the top left corner of the input matrix, and you multiply each overlapping square together and add them all up. That's the value of the top left element of the result (the result is a matrix). Then you slide the filter to the right, repeat the calculation, and that's the value one to the right in the result matrix. You repeat, panning and scanning across the input matrix and in this fashion fill in the the whole result matrix.
I thought I had it, but this blog post just confused things. I do remember learning about convolutions back in undergrad and being confused then, too, so that sounds about right. When I read the ML handbook, I had a "wait, is that it?" thought and was confused why I was confused back in the day. But I guess that's just one particular way of doing it, or one use case or something.
[0] https://www.amazon.com/Hundred-Page-Machine-Learning-Book/dp...
as mentioned at the end of the post, if you understand things in terms of functions rather than just multiplying with a sliding window, you can then make use of Fourier transforms, and this enables much faster algorithms to compute the same result.
You remember (a + b)(c + d) = ac + ad + bc + bd?
Or (2x^2 + 1)(4x^3 + x) can be calculated by doing conv([0,2,0,1], [4,0,1,0]) (the numbers represent the coefficients). The results are the coefficients of the resulting polynomial.
As the polynomials grow larger this gets harder and harder computationally. Convolution theorem lets you use Fourier to you convert this multiplication into a pairwise multiplication so that
fft(f * g) = fft(f) .* fft(g)
where * is convolution and .* is pairwise multiplication.
http://www.dspguide.com/ch6.htm
Pedagogically speaking, that book is one of the best technical books I have ever read.
This sounds like a big assumption. I think this is a deep-learning oriented blog post, and not for people with a background in systems & signals (basically information theory). I don't understand mathematical convolutions super well (I know the vague derivation, of integration of f(t)g(t-T)dT, and the metaphor of multiplication in frequency space), but I do get how convolutions work with neural networks, and maybe this is a failure on my part, but it doesn't feel like there's a super strong connection between the two. Is there something I'm missing out on?
I found the article a bit roundabout. Perhaps I'm not part of the intended audience; I understand convolutions, but have only limited familiarity with the internals of neural networks.
"How likely is it that a ball will go a distance c if you drop it and then drop it again from above the point at which it landed?"
p.s. I’m assuming of course that your method is functionally equivalent to the standard convolution. Is it?
Fouri-ous there's no mention of Fourier https://en.wikipedia.org/wiki/Convolution_theorem
why do we say convolution instead of total likelihood, which is so much more to the point the goal that is trying to be achieved?