Frank Rosenblatt's perceptron paved the way for AI 60 years too soon (2019)
news.cornell.edu
news.cornell.edu
The unfortunate thing was not the Perceptrons book but the fact that Rosenblatt died prematurely soon after. He was very well-equipped to defend and carry on work on NNs, I think.
Olazaran argues that all the connectionists were perfectly aware of the _Perceptrons_ headline conclusion about single-layer perceptrons being hopelessly linear, which drafts had been circulating for like 4 years beforehand as well, and most regarded it as unimportant (pointing out that humans can't solve the parity of a grid of dots either without painfully counting them out one by one) and having an obvious solution (multiple layers) that they all, Rosenblatt especially, had put a lot of work into trying. The problem was, none of the multi-layer things worked, and people had run out of ideas. So most of the connectionist researchers got sucked away by things that were working at the time (eg the Stanford group was having huge success with adaptive antennas & telephone filters which accidentally come out of their NN work), and funding dried up (for both exogenous political reasons related to military R&D being cut, and just the lack of results compared to alternative research programs like the symbolics approaches which were enjoying their initial flush of success in theorem proving and checkers etc). So when, years later, _Perceptrons_ came out with all of its i's dotted and ts-crossed, it didn't "kill connectionism" because that had already died. What _Perceptrons_ really did was it served as a kind of excuse or Schelling point to make the death 'official' and cement the dominance of the symbolic approaches. Rosenblatt never gave up, but he had already been left high and dry with no more funding and no research community.
Olazaran directly asks several of them whether more funding or work would have helped, and it seems everyone agrees that it would've been useless. The computers just weren't there in the '60s. (One notes that it might have worked in the '70s if anyone had paid attention to the invention of backpropagation, pointing out that Rumelhart et al doing the PDP studies on backprop were using the equivalent of PCs for those studies in the late '80s, so if you were patient you could've done them on minicomputers/mainframes in the '70s. But not the '60s.)
I once heard Daphne Koller say that before big data, neural networks were always the second best way to do anything.
Which is why I wish there was a copyleft open source data analogue. If you train on everyone’s public data, your model should have to be just as publicly available.
Is this really common enough it's even worth mentioning?
It's not clear whether the model trained on some data is copyrightable at all (and if it isn't, it can't be a derived work and gets no protection nor restrictions from copyright law) as in general facts about a work - including things like word frequency statistics, which was a popular type of trained language models not that long ago (e.g. n-gram models used in statistical MT systems) - are not copyrightable and in that case making/distributing those is not an exclusive right of the author and needs no license or permission, that was settled long ago between publishers and e.g. dictionary makers.
There is also the notion that mechanistic transformations or difficult labor can't result in copyrightable work, there needs to be human creativity involved; in US copyright law (as in Fiest Publications Inc. v. Rural phone service Co, also see https://www.gutenberg.org/help/no_sweat_copyright.html) the fact that making some work required lots of work and cost you lots of money does not imply that it deserves copyright protection, no matter how much work was required; so it's irrelevant that someone spent millions of dollars worth of GPU time, and it could certainly be argued (case law would be useful here!) that the software used to train the model is copyrightable, but the model output by that software is not.
Perhaps case law or some new explicit law will settle otherwise, but currently all the research and industry is proceeding with the assumption that a trained model is not a derived work (in the copyright law sense) from the training data, and as far as I see this assumption is not being challenged in courts.
Too early to say that with a straght face yet.
"Deep learning" produced some amazing generative art, and some very flashy research papers, but it failed to drive business decisions. (And not for the lack of trying, that's for sure.)
You don’t have to fall for the hype and think deep learning will lead us to AGI in the next 5 years, but dismissing it as generative art and flashy papers isn’t any more accurate.
As for Alexa, et al - these things fall into the "generative art" bucket. The search results they give aren't any better than the Eliza-tier expert systems of yore. They just feel much better and more human when you use them.
(Which is also important, but doesn't drive business decisions, except as part of a marketing strategy.)
I personally own a startup that uses deep neural networks as part of its core product functionality and our business decisions would be drastically different if it weren’t for modern machine learning. We could technically ship some similar products, but they would either be far inferior or take orders of magnitude longer to engineer by hand.
Deep learning is a cool feature to differentiate Alexa from competing home electronic gadgets.
But if you want to forecast Alexa sales, or understand market segmentation, or the portrait of a typical Alexa buyer then you need something other than deep learning because neural nets utterly failed in this domain.
The business metrics problem is vastly, vastly more important than the "making cool gadgets" problem, and huge resources were poured into making deep learning a thing in this space. The money was mostly wasted. (A negative result is still a result, but still the misallocation of resources is staggering.)
FYI, deep learning isn’t a differentiating feature for the Alexa, it powers essentially all modern voice applications.
https://ai.googleblog.com/2020/11/improving-on-device-speech...
Before the ImageNet dataset, CNNs weren't worth it.
Ed: To be clear, the idea to use them to adapt the weights of NNs was also from the 70s but only rediscovered and applied to MLPs by at least two independent groups/individuals in the 80s.
Indeed. The book Talking Nets: An Oral History of Neural Networks[1] covers a lot of this ground. Read it and you'll see many people who were involved in the early history NN's mentioning how backprop was discovered and re-discovered over and over again.
[1]: https://www.amazon.com/Talking-Nets-History-Neural-Networks/...
Without a non-linearity depth doesn’t buy you anything.
Here's a 1963 paper, "System and Circuit Designs for The Tobermory Perceptron", which includes circuit diagrams. It would be intriguing to build one, a bit like building a working Babbage machine.
https://blogs.umass.edu/brain-wars/files/2016/03/nagy-1963-t...
All the pdfs (use filter) this blog has posted, http://web.archive.org/web/*/blogs.umass.edu/brain-wars/*
The analysis is modern because 1) unlike traditional statistics, it does not make any assumptions about the distributions of each class, and 2) is a non-asymptotic guarantee. It is also an online analysis, in the sense that it only compares performance against the optimal fixed classifier for the sequence that was actually observed, rather than any notion of an underlying distribution.
More details: http://www.argmin.net/2021/11/04/perceptron/
Makes you wonder if there are other abandoned techniques that might be worth circling back to nowadays...
https://blogs.umass.edu/comphon/2017/06/15/did-frank-rosenbl...
An MLP: Multilayer Perceptron can learn XOR and is considered a feed-forward ANN: Artificial Neural Network: https://en.wikipedia.org/wiki/Multilayer_perceptron
Backpropagation: 1960-, 1986, https://en.wikipedia.org/wiki/Backpropagation
> In fact, the approach laid out in PDP is very similar to the approach used in today's neural networks.
From: https://github.com/fastai/fastbook/blob/master/01_intro.ipyn...
The problem with your argument is we know too little about higher level brain operation to make any confident comparisons like that.
Regular expressions, perceptron, neocognitron, genetic algorithms, deep learning...