A Proposal For the Dartmouth Summer Research Project on A.I. (1955)
www-formal.stanford.edu
www-formal.stanford.edu
It's a highly recommended read [http://www.aiai.ed.ac.uk/events/lighthill1973/lighthill.pdf], or you could watch his presentation of the report (to Minsky amongst others) [http://www.aiai.ed.ac.uk/events/lighthill1973/] it's also on YouTube.
He was astute in identifying that AI has succeeded in A) Automation of well defined tasks, C) Investigation of problem solving processes, but failed to product much in the way of B) the combination of A and C, an independently intelligent artifact.
It's been a while since I read it, but I remember the video being entertaining, especially the exchanges between the Lighthill and the Minsky, and the analysis being relevent even to today's state of AI.
The Lighthill report destroyed the UK's lead in AI at Edinburgh.
Edinburgh's AI lab, founded by Donald Michie, a wartime colleague of Turings & Richard Gregory a vision & theory of mind expert.
Edinburgh had a Robot Arm that could assemble various wooden toys from randomly scattered blocks using vision & planning.
Edinburgh had produced POPLOG, a widely used European LISP (with less brackets :) )
Michie was a proponent of Machine Intelligence, his "trial and error" BOXES algorithm could learn to balance a pole - everywhere else used hand engineered Symbolic GOFAI.
Michie BOXES enabled learning robots that anticipated neural nets & SGD using reinforcement learning.
Edinburgh's unique vision was world leading at the time. Sadly European industry followed Lighthill's lead - the 1st AI winter.
Lighthill was a pure mathematician and not well qualified to vet AI, his criticisms proved wrong in hindsight - Edinburgh had automation, vision & learning on a PDP-11.
Pariah in Europe, Donald Michie went on to help develop Japanese robotic assembly lines and use BOXES for factory & satellite control. Here he laments : [18:24] http://www.bbc.co.uk/iplayer/episode/p0306rt1/micro-live-ser...
\sidebar:
POP-11 has Lisp-y list support with a Pascal-like syntax. It was pretty nice. It had assignment the right way around (if you have the stack as a mental model):
5 -> x
Grad students and undergrad keeners were advised to learn enough LISP to read papers from the MIT AI Lab :)https://en.wikipedia.org/wiki/Poplog
It was pretty well done, but not exactly fast on a 1990 Sun minicomputer shared by a hundred students.
Almost all of our undergrad assignments were set and submitted via the POPLOG system, including lecture notes and tutorials: 'teach texts' with hypertext links. You could highlight code snippets with the cursor and run them in the REPL. All pre-web on a VT-100. Great stuff.
Did you work with Margret Bowden's robotics group ?
She was a very respected figure in philosophy of cogsci and AI, but I don't recall her doing any practical robotics, which was barely present at Sussex at that time.
This "shower thought" within the report also caught my eye:
Incidentally, it has sometimes been argued that part of the stimulus to laborious male activity in creative fields of work, including pure science, is the urge to compensate for lack of the female capability of giving birth to children. If this were true, then Building Robots might indeed be seen as the ideal compensation. There is one piece of evidence supporting that highly uncertain hypothesis: most robots are designed from the outset to operate in a world as like as possible to the conventional child's world as seen by a man: they play games, they do puzzles, they build towers of bricks, they recognise pictures in drawing books (bear on rug with ball) -- although the rich emotional character of the child's world is totally absent.
It's somewhere in the Venn diagram between deeply perceptive and reductively essentialist, and I can't decide where.
When Google creates a program to learn Go, it learns go so well that it knows it (arguably) knows it far better than any human (even if it isn't flawless).
But what did we learn about go? Well, we learned a bit about the opening I guess, since Lee Sedol has become fond of the "AlphaGo opening," but other than that... not much, right?
That's the funny thing about neural networks. They can converge to a set of weights that, when activated, perform better than any human. But we can't look at Weight 483 and Weight 958 and say "Ah, that's where it decided the corner is very valuable!" or something.
It learns, we don't. We can only learn from what it can then show us it has learned.
People don't bother to because:
1) it's a very boring problem (we already have a high-level view of what networks learn through various visualizations, and what you'd learn would be specific to one network learned for one dataset)
and 2) it's very tedious and not repeatable (have to do it all over for each new dataset and each new model).
You could look at AlphaGo's weights for the entire neural network for ten thousand years. But you would become no better at Go. The only way AlphaGo can help us improve at Go is by showing us what it has learned in the games it plays.
Sure, humans wouldn't become better at Go. But that's a limitation of the human brain (we're not good at mathematical memorization and computation).
For all we know, what the network has learned about Go (a highly complex and interconnected set of statistical dependencies) is what there is to learn about Go. You're implicitly making the assumption that what the network learns about Go is guaranteed to be translatable to something humans can learn.
On the contrary, what the network learns is merely reducible, with loss of accuracy to what humans can understand. And that is an active area of research (feature visualizations and explanations), but that is tangential to your point.
They were so optimistic in the early days. And they had so little compute power.
Also see the graphs for atari environments: many games are played by rl agents "at human-level or above".
It also feels a little like reading Andy Warhol's diary and realizing all the famous people knew each other. Never realized they were so close.
I hope they got their funding.
"Perceptrons might be fired to the planets as mechanical space explorers."
This has led Solomonoff to investigation of a question of universal sequence prediction. A couple of years later Solomonoff wrote a paper about such a system for prediction that used algorithmic probability (he is cited later as the original developer of algorithmic information theory which was later independently discovered by Kolmogorov who later acknowledged that the Solomonoff was the first). This method, Solomonoff's induction, is proven to be the most optimal (though incomputable) machine learning method possible.
He has never abandoned this project and for the rest of his life he focused on making more sophisticated system designs that are computable while still being proven to be optimal.
His latest system is called "Alpha", and it is designed as a machine for solving a sequence of function inversion and time limited optimization problems (a majority of science/engineering problems can be formulated this way) in a way that exploits experience gathered while solving these problems. This system, again, is proven to be optimal in a certain sense. He also tried to implement this system with various practical optimizations, but it didn't converge fast enough on his training sequences and on the hardware of that time.
Still, with modern hardware it is a possibility that it could work. And the whole design is described in the papers, so people can (and actually do, though privately and perhaps without much success) implement this system.
Here are the relevant papers: http://world.std.com/~rjs/publications/IncLrn89.pdf http://world.std.com/~rjs/nips02.pdf
I've long had a bias against minsky because I thought he said some very silly things about AI back in the day, but I think I was probably wrong, or at least that he deserved more attention than I gave him. I watched some interviews with him in the Youtube channel 'Closer To Truth' and he's by far my favourite interviewee. Very incisive.
EDIT: Like I said, he did have good reasons! That's why Perceptrons was so influential; it's just the weird unfortunate luck of history that he ended up diverting effort from what's now become a much more promising field.
None of that makes Minsky right, but it's hard to see how much could even have been achieved on neural nets back in '55. Our architecture design today descends from experimental results that were not going to be available for many decades.
EDIT: well I basically made the same point as you.
There are two ways that Minsky could have (mathematically) looked past the linear separability issue:
1. If you add another layer to the perceptron, it can solve the XOR problem.
2. If you add a non-monotonic threshold function, it can solve the XOR problem.
So these are two rather simple solutions to the issue he brought up.
Minsky assumed a trained percepton with the weights already set to act like an AND or an OR gate. He wasn't dealing with the learning problem.
Anyway, IMO people overestimate Minsky influence to single-handedly shut down an avenue of research. The reason _Perceptrons_ conclusions caught up is because they were sound and reasonable to his peers at the time.
I also seem to remember his video interview from a few years back where he elaborates on perceptrons and how much of his original conclusions are applicable to the state of art ANNs. Can't quite find it though.
The role of Minsky in killing perceptrons is seriously overblown.
I see from the wikipedia article you linked to, that they did know about the multiple layers. I thought it was suspicious that they had somehow missed it since it is so simple (at least to us now), and these guys are so very smart.
I wonder if they also knew (or realized, rather), that a single layer neuron with a non-monotonic function could have also "solved" XOR.
It's great to be enthusiastic about breakthroughs, but the history of AI is littered with partial success stories.
1. The computational resources at the time. I tried to run 14 node multi-layer NNs in the 90s and I'd have to go to lunch and come back before a single run was done. (backprop looped several times to converge).
2. The whole symbolic vs statistical debate that was going on at MIT. You had Chomsky, Fodor and Minsky lining up on almost philosophical grounds. (In Fodor's case, explicitly so).
according to [1]: "The project was approved and brought together a group of researchers which included pioneers such as Newell, Simon, McCarthy, Solomonoff, Shannon, Minsky, and Selfridge, all of whom made seminal contributions to the field of Artificial Intelligence in later years. "
[1] http://www.asiapacific-mathnews.com/04/0403/0015_0020.pdf
"After fifteen minutes of searching with Google, the majority of web pages give a citation that the person who said this was Marvin Minsky and the student was Gerald Sussman. According to the majority of these quotes, in 1966 Minsky asked Sussman to "connect a camera to a computer and do something with it".
They may indeed have had that conversation but in actual fact, the original Computer Vision project referred to above was set up by Seymour Papert at MIT and given to Sussman who was to co-ordinate a group of 10 students including himself. [2]
The original document outlined a plan to do some kind of basic foreground/background segmentation, followed by a subgoal of analysing scenes with simple non-overlapping objects, with distinct uniform colour and texture and homogeneous backgrounds. A further subgoal was to extend the system to more complex objects.
So it would seem that Computer Vision was never a summer project for a single student, nor did it aim to make a complete working vision system. Maybe it was too ambitious for its time, but it's unlikely that the researchers involved thought that it would be completely solved at the end. Finally, Computer Vision as we know it today is vastly different to what it was thought to be in 1966. Today we have many topics derived from CV such as inpainting, novel view generation, gesture recognition, deep learning, etc."
[1] http://www.lyndonhill.com/opinion-cvlegends.html [2] https://dspace.mit.edu/handle/1721.1/6125
“[T]he major obstacle is not lack of machine capacity, but our inability to write programs taking full advantage of what we have.”
Which is more relevant now than ever. It's interesting that at the time they still didn't consider themselves to be taking full advantage of what they had—which was positively primitive—and I wonder how they planned to squeeze more power out.
http://www.in2013dollars.com/1955-dollars-in-2016?amount=135...
Q: What if the original title is too long to fit in the space allotted?