Differentiable Programming
edge.org
edge.org
colah's blog post can be roughly summarized as taking the (well-known) representation of NNs as DFGs of modules (since the very early 90s, if not earlier), and then showing how common combinators in functional programming languages correspond to subgraphs in common modules. Since both FPs and DNNs can be represented as DFGs, this isn't a particularly surprising (or novel) correspondence. Several libraries used exactly these functional combinators for symbolically expressing NNs well before his blog post.
The trend of differentiable programming (as in, using gradient-based methods to train a network to learn an algorithm) is what I think the author is trying to highlight. This has been around since e.g. Das et. al, "Learning context-free grammars: Capabilities and limitations of a recurrent neural network with an external stack memory" (from 1992). There's a lot of history (and duplication) in this area, some keywords to search for being NTMs, RL-NTMs, Pointer Networks, Neural GPU, MemNNs, MemN2N, Stack RNNs, .... - the RAM (reasoning, attention, memory) workshop at NIPS 2015 had a bunch of recent work in this area.
The high-level idea is by augmenting a (trained) controller (e.g. an LSTM) with some external datastructure(s) (stacks, queues, memory, hierarchical memory) with the ability to read from and write to the external datastructure(s), and to propagating errors back to the controller (hence differentiable programming moniker), via hard (e.g. RL) or soft attention mechanisms.
In particular, optimization, being a branch of math, is itself a type of functional programming. Similarly, neural networks are functional programs. Programming with differentiable functions as a field on its own goes back at least as far as the 1960s. Closely related ideas like algebraic topology go back to the 1890s (this would be programming with continuous functions rather than differentiable).
The language of type theory is typically the province of computer science. The corresponding language in math is category theory. There are a few mathematicians who have been working on a categorical theory of neural networks. But this is pretty far from a mainstream research area.
Perhaps with people like Colah and Dalrymple pointing out these connections from a more applied point of view, these ideas will pick up steam.
[0] There is are technical caveats here that aren't interesting when talking about math that can be done on computers.
upvote for visibility and making sure colah gets credit for coming up with the idea first.
colah goes deeper and talks about how it relates to algebraic types from functional programming.
Holy fucking wow. Do they know what an Internet is?
No doubt colah's excellent blog post dives much deeper than this essay, but please note that this was wrtten for Edge.org's http://edge.org/annual-questions book, for a general audience, with very high level perspective, with very little space, and without the ability to get technical. Don't pit them against each other! They complement!
By the way, for one interesting example of the "conceptual agility" of scientists at the very edge, check out this interesting story about Feynman -- from a David Deutsch interview (with Sam Harris) -- in which Feynman derives months of Deutsch's work in a few minutes (listen for ~5 min):
https://www.youtube.com/watch?v=J21QuHrIqXg&t=6524 (1:48:44 - 1:52:40)
Scientists' minds are well primed to understand, derive, and formulate ideas -- they are super fertile memetic environments. And what's more-- they have the epistemic filters to annihilate untruths and converge on scientific knowledge (the best explanation not yet falsified). In math (inc. computer science), it's even better because we study and manipulate the objects themselves, not data from measurements of the things. The formulation, derivation, convergence, remixing, (and so on) of ideas is much, much faster.
Knowing both colah and davidad, I confidently assert they are both among the most brilliant young scientists alive. Humanity stands to gain much from the constructive interference of their minds. As readers and commenters, let's foster that.
Didn't get very good results though.
Maybe I'm a bit peculiar, but this article sent chills down my spine.
[0] http://www.damninteresting.com/on-the-origin-of-circuits/
It is indeed similar to some work from FAIR, where one of the tasks was binary addition: http://arxiv.org/abs/1503.01007
There is a lot of work around this going on at the moment. Google's neural Turing machine is a similar idea.
https://github.com/zenna/Arrows.jl https://github.com/wojzaremba/algorithm-learning
Is this true? My impression is that "deep learning" is a series of mathematical tricks and design and implementation concepts that get you to solve neural networks of depth greater than 3-4, which becomes mathematically challenging.
Also, we now have access to huge datasets to try our algorithms on. Progress in ML depends a lot on the training data that is available.