Neural Nets (1994)
teamten.com
teamten.com
A few years later SpamAssassin came out and I tried to use it to train a spam classifier without any luck. And I tried my hand at a few protein structure SVM classifiers (failng miserably, I didn't really understand it was critical to have a balanced false and true training set).
A few years after that I landed at Google (around 2007) and very few people were doing machine learning, other than ads and search and the work that was being done was far from what I knew about (mainly supervised training with SGD on batched data). Eventually Google adopted the paradigm I enjoy (synchronized SGD using allreduce).
Nowadays, we have ample CPU and data to train amazing models (with the most interesting working being around large language models). It took a lot longer to get there than I expected.
If you gave someone in 1994 the GPT-3 code and dataset, it would be impossible for them to regress and very difficult to run even if we regressed it for them. AI algorithms may not be limited by "what we tell them" but they ARE limited by the hardware we run them on.
NN models do converge to something (both in the sense of regressions converging, but also in the sense that adding more nodes eventually stops improving performance at a given task.) I suspect but cannot prove that in most cases what it converges to could be expressed more concisely and efficiently as something other than a NN. (I.e. that NNs can approximate any function does not imply they can do so efficiently)
So at the end of the day, the programmer needs to understand the algorithm well enough to know if a NN-based implementation of it would achieve sufficient performance on available hardware. If the answer is no, then the programmer still has to come up with something alone.
> I suspect but cannot prove
I think the rates of untrainable weights in a NN is proof enough, ie., that often around 30pc aren't being updated.
Also, what are some current technologies and ideas that are unpopular but which have some dedicated fans who insist there's major untapped potential (even if maybe it can't be fully tapped into today)? Specifically, little-known things rather than pretty popular stuff like quantum computing.
Then he got killed at work.
I didn’t hear about neural networks again for close to 20 years, but I can attest at least someone was stoked about them back then.
NN have been cyclical for decades now and I expect they'll cycle again.
I'm not complaining as I was always well paid in my career. But sometimes, being too early in a field can be as bad as being too late :-)
I am still getting about a new citation every month for one of my papers that was published 27 years ago!
BTW, I recently happened across a paper from 1967 (53 years ago) that mentions neural networks in passing as if it was a popular idea at the time - the original paper on the Medial Axis Transform! Scroll to the 2nd page (“364”) near the top: “Consider a continuous isotropic plane (an idealization of an active granular material or a rudimentary neural net) that has the following properties at each point”
http://pageperso.lif.univ-mrs.fr/~edouard.thiel/rech/1967-bl...
We had pretty high hopes back then. I and I think many others sort of assumed we would have had General AI by now. Or at least going in that direction. Another classmate said in 1992 that scientists had simulated the neural net equivalent to a worm or something. This topic came up re the neural net featured in the movie Terminator 2.
The current use of "neural nets" are both under- and overwhelming. But definitely over-whelmingly boring, to me. It's good, don't get me wrong. But from my viewpoint, it looks like we have found a trick, and now try to apply this same trick to as many new fields as possible.
Right, this is normal and expected. While scientific growth looks like an exponential curve, the lateral application path is equally important to maintaining that curve!
We may get to AGI one day (I have my doubts) but I have to remind myself that all theses clunky things we see today are v1 -- and virtually everything can improve!
It's like the difference between Kitty Hawk and a Boeing airliner. Yes, they are impressive feats of engineering, and the massive training hours tell the story.
But the Boeing Airliner is still less clever than a fruit fly, to merge the metaphors.
Edit:
to expand - it seems to get something as smart as a dog is almost the same as something as smart as a human.
1994 neural net
2020 GPT-3
.
.
.
.
.
unknown numbers of dots...
.
.
.
dog
humanThe funny thing to think about for people on this forum was what weak computing power was available to me at that time. As I recall I did this whole experiment in C (using the TLearn libraries) on a Masscomp workstation. The cool thing about the Masscomp was that it had a vector processing board that could do some parallel computations. Well I didn't use that but the lab had the hardware. So I was doing all these things on Motorola 68000 with less RAM and way less CPU than your watch.
There were at least some economists that had hope back then. Here's a paper my advisor published in 1995, but that he had written several years earlier for his dissertation:
http://econweb.rutgers.edu/nswanson/papers/selec3.pdf
Hal White was working with NN in the 1980s. Here's one of his publications from 1988:
http://www.machine-learning.martinsewell.com/ann/White1988.p...
I think distributed systems research is enabling personal user-level distributed networks (think like mastodon but to the point where a lay person can own their own node without much trouble). But will the economies of scale for commodities goods like motherboards make sense against behemoths buying up lots of hard drives? Who knows.
1994 is around the time I read Steven Levy's "Artifical Life"[1]. I definitely had the excitement the author has about genetic algorithms and classifier systems. They seemed much more approachable than neural networks, too. They appealed to the assembler programmer in me, I guess.
I'll echo what others are saying here. Brute force compute has opened the door for so many possibilities. Evolution has had so more "CPU" and "compute substrate" to play with than we can possibly imagine. I'm unsettled by the idea that we're training models we don't actually understand. I'm also not dismissive of the possibilities that they open up, though.
[1] https://www.penguinrandomhouse.com/books/100796/artificial-l...
I’d bet this holds in all sorts of fields. Somebody has a great idea that really needs good steel, and they’re unfashionable (“that’s dumb! It’ll never work (today)!”) until we invent a better steel process. Or ideas that just took a microscope to verify. So many cycles of debate and orthogonal progress rendered unnecessary when a relatively unrelated change happens.
GAs can be quite useful in situations where there are tons of very bad local minima/maxima and you desire the global minima/maxima. Unfortunately, neural networks have lots of really great local minima/maxima and the state space is so large that you'll likely never get the global minima/maxima (and you wouldn't know it if you did). This is why "Neuroevolution" of neural network weights hasn't really caught on.
One of the beautiful aspects of Evolutionary Computation is how you can use it as a generative mechanism. A fairly simple fitness function can produce complex and intriguing outputs.
Well, a learning algorithm is always going to be limited by something. In AI literature any such something is grouped under "inductive bias" and there is no machine learning algorithm that doesn't incorporate some kind of inductive bias. Bayesian learners have their priors, distance learners have their disance functions, Support Vector Machines in particular have their kernels, and of course neural networks have their intricate architectures painstakingly hand-engineered and fine-tuned to a particular domain, or even a specific dataset [1]. Indeed it is probably impossible to have learning without inductive bias [2] [3].
The opinion in the short piece above is representative of a current trend in machine learning, of taking "the human out of the loop", which loosely translates in trying to learn everything end-to-end, only from examples, while pretending that no attempt is made at any point to guide the learner to find a consistent hypothesis that explains the examples.
Unfotunately, in practicce, this ideal remains a fantasy. All the progress achieved with deep learning in the last few years would not have been possible without the discovery (by chance or concerted effort) of good biases that are conducive to learning in specific domains, e.g. convolutional layers for image recognition, or Long-Short Term Memory cells for sequence learning, etc.
And where did these good biases come from? Why, from human programmers. Humans ourselves most likely come equipped with very strong, very useful biases. We've used those biases and our ability to generalise to come up with powerful abstractions, such as the laws of physics or mathematics. It took us literally thousands of years to amass this fortune of knowledge (possibly millions, during our evolution).
Why would we not use those finely-honed biases of ours and all the knowledge we've collected to kickstart a new form of intelligence? After all, when we want fire, we don't sit around waiting for thunder to strike a tree, anymore. We can start fires on our own. Because of our intelligence- because to build upon prior knowledge is intelligent.
__________________
[1] See for example the Neural Network Zoo, a collection of neural net architectures:
https://www.asimovinstitute.org/neural-network-zoo/
Or most of the deep learning literature published in the last few years, where every little architectural tweak is presented as a major breakthrough.
[2] Famously argued by Tom Mitchell in "The Need for Biases in Learning Genearlizations":
https://www.semanticscholar.org/paper/The-Need-for-Biases-in...
[3] But also see the discussion between Yan Le Cunn and Christopher Manning:
>
In the days when Sussman was a novice, Minsky once came to him as he sat hacking at the PDP-6.
“What are you doing?”, asked Minsky.
“I am training a randomly wired neural net to play Tic-Tac-Toe” Sussman replied.
“Why is the net wired randomly?”, asked Minsky.
“I do not want it to have any preconceptions of how to play”, Sussman said.
Minsky then shut his eyes.
“Why do you close your eyes?”, Sussman asked his teacher.
“So that the room will be empty.”
At that moment, Sussman was enlightened.
Because priors are hard to embed in a model. For example, CNNs are great for translation invariance, but rotation and scaling don't come out of the box. Why don't they simply add the rotation invariance to the model? Because it's hard to express.
Also, human priors are limited. If AlphaGo was to be limited to human priors it would never have surpassed us.
The best approach so far is to make a network as free from priors as possible (like the transformers) and let it learn from the data, in essence let it rediscover the convolution or other efficient operations from mountains of data.
That's a limitation of neural networks, where inductive biases are very difficult to represent. Other approaches don't have any such problem, e.g. all the other approaches I listed above have well-defined, clean and tidy representations for inductive bias.
As to AlphaGo, this may come as a shock, but AlphaGo was not the first system to surpass humans in anything. It was the first system to surpass humans _in Go_ but e.g. the first computer system to beat a human grandmaster in chess was Deep Blue [1], in 1997, the first computer system to win a world championship against human players was the checkers (draughts) player Chinook [2], in 1990, the first computer system to outperform humans in medical diagnosis was MYCIN in the 1970's and so on. All those were systems that used strong inductive biases.
And of course, AlphaGo itself was limited by human priors- e.g. piece moves, checkerboard dimensions and structure were hard-coded into its architecture and a search for moves was performed by MCTS.
In any case, the lack of good enough priors is not a reason to not use any priors- it's a reason to look for better priors.
Edit: Has any end-to-end approach rediscovered an entire state-of-the-art architecture, like CNNs or LSTMs?
___________________
[1] https://en.wikipedia.org/wiki/Deep_Blue_(chess_computer)
[2] https://en.wikipedia.org/wiki/Chinook_(computer_program)
AlphaGo did surpass humans before AlphaGo Zero.
See eg also how classic chess engines work.
(Of course, you could describe alpha-beta search as some form of self-play. But that's probably going to far.)
So, its fine that we don't understand in detail what every artificial neuron firing means....
I’m personally tired of this ass backwards argument.
Of course it’s possible to understand how brains work. It’s called science. We (collectively) are well on our to understanding the biochemistry of neurons and brain structures.
Everything exists in the same universe and follows the same physical laws.
That we're so early in this cycle doesn't mean it's impossible.
The good news is that there is a lot of evidence that computation in the brain is more modular than that... It's not necessary to understand each neuron... Just like it's not necessary to understand each artificial neuron in a large nueral network.
However, I entirely agree with the sentiment: We don't necessarily understand brains now, but we don't dismiss their capabilities just because we can't explain precisely how they do what they do. By the same token, complex tools like deep learning shouldn't be dismissed just because we can't satisfactorily interpret them.
Recursion is a thing so it is not clear why this statement is self evident. Perhaps the word "truly" does some mysterious heavy lifting?