Deep learning is hitting a wall?
nautil.us
nautil.us
I think that the best way to communicate the message that deep learning isn't enough is to use a different approach to achieve superior results. Ideally something that makes a lot of money quickly. There are only so many people that read nautil.us articles. But everyone pays attention to things that make a lot of money.
I can point out many conceptual flaws in the internal combustion engine. And I can also propose other types of engines as alternatives. But people don't really start paying attention until someone makes a Tesla. If symbol manipulation can do a better job of identifying pictures of rabbits or humans holding stop signs, then let's see it.
Edit: No offense to Gary Marcus, first heard of him after reading this article.
a) far far more recurrent than any practical deep learning system
b) using something other than backprop, and probably that something is much more efficient
My understanding of deep learning is that it is a methodology, not a technique.
No offense, just startled, but that’s the most self-contradictory thing I’ve ever heard.
The human brain is by definition a neural network…
I always wondered why people stopped referring to ANN and just went with neural network. Perhaps some anti-science bias or something like that.
Apparently deep learning has been beat-to-death as a word, but the way deep learning was explained to me, it is an approach not a specification.
No, not right. Humans need other humans to be smart, we need culture, technology and science. We're not that smart - just 700 years ago we were dying of the plague without even knowing the germ theory of disease, even with our lives on the line we couldn't solve it. We're only INCREMENTALLY smart over our current cultural level.
Arguing that memes exist outside of genes is like arguing that the Platonic solids are real.
https://web.archive.org/web/20110608070200/http://www.uic.ed...
Protoculture as well?
We are actually doing that, though. Across the globe, the fraction of energy generated by non-renewables has been falling consistently for more than a decade. We went from being less than 5% solar + wind a decade ago to more than 10% now, and the trend is accelerating. By 2026 we're expected to be nearing 20% of all power generation renewable. Things aren't going fast, but infrastructure doesn't go fast, and no proponent of renewable energy expected it to.
I think this is in no way comparable to the symbolic learning thing where there is no sign of relative progress.
As far as the hard cap you mentioned, I think there's no danger of us hitting that cap in the next decade or so. Later on, maybe. Battery technology is also improving rapidly, so perhaps it will never be a true limiting factor.
Proponents of symbolic learning should start by just achieving results that are even remotely close go those achieved with deep learning. Nevermind outperforming it. Because the reality is, for NLP and anything related to CV, deep learning has consistently wildly outperformed every other approach. Thousands of ai researchers have tried to make "old school" AI work for those problems with only very limited success.
Now my background is in CV (not in the field anymore, but was until 2019) so I'm not sure about how well other methods stack up against DL for other use cases. But to me symbolic approaches for most of what DL excels at just seem like a completely unfeasible (and overcomplicated) pipedream.
I don't disagree with the premise that current research has kind of stalled compared to the early-mid 2010s but that mostly means we need to figure out new ways to do ML and DL, not because DL has failed as a concept . & keep in mind that the fact we can even say that things have "stalled" is because we got spoiled by the huge performance leaps that DL made possible in the first place.
The symbolic AI crowd have always had a "2 more weeks" narrative where they promise that they will figure out a given problem very soon if only x condition was true. The problem is that they have almost always terribly underdelivered, and that was true even when 99% of the research and funding was focused on symbolic or hard AI. Subsymbolic approaches could also be argued to have yielded somewhat underwhelming results (vs the hype) but the other methods are in a league of their own.
Nobody that I know in the DL community wouldn't be excited to switch to another method if it meant achieving actual measurable improvements. So if the data was conclusive, the alternatives would be used. No one cares about the purity of the method or whatever, because the messiness is inherent to DL. Also, the current AI community is very very focused on performance and state of the art results (that can be a problem in some cases but not w.r.t to method agnosticism) rather than on any big theories on how human intelligence ought to be translated mathematically (or not)
Tesla has a horrible car accident where a car mistakes the side of a truck for the sky. They start over and create a new ML system only for the exact same kind of accident to happen again.
His points about provable and debuggable making ML radically unsuitable for huge swaths of the tasks where AI would be most desirable are exactly on target.
Meanwhile, a tiny little slug seems to understand much more about the universe than the most advanced AI we've ever designed and it does this on nanowatts of power.
With ML there is at least a path to solve way-easier-than AGI problems like maybe (semi?) autonomous cars, given enough compute and redundancy. With anything more old-school? Maybe you'd get lane detection that works on sunny days. And yeah you will be able to "debug" and prove things more easily but what's the point if there's no real way to fix the bugs if they are inherent to the flaws of symbolic approaches? If you think Ml has a problem with edge cases (and it does), keep in mind that GOFAI methods are usually not even close to be complete enough to have "edges". "the last 1% is the hardest to deal with" becomes more of a "if we can get over 60% it's a miracle"
Also, no matter how suboptimal the tesla autopilot is, I have not seen any figure that show that they do worse than humans w.r.t to fatality or accident rates. Maybe they are worse! But just the fact that they are even remotely close is not something that could've been possible even 15 years ago.
But I'd be thrilled if we get to anything better than current DL methods. Hell, it might even pull me back into the field! I just really doubt it's going to be old school symbolic or causal methods that will get us to the next big steps.
Your eyes run on more-or-less fixed-function neural networks to see and recognize stuff.
The part of your brain that does the actual driving is much more symbolic in how it analyzes and reacts.
When you look at self-driving, there's a general pattern that it handles the 80% of easy stuff (the stuff that you could drive while texting and almost completely ignoring the road). There's a steep gradient down to "complete failure" from there. Accidents happen during a tiny, tiny subset of situations (a minuscule fraction of a percent).
Vehicle deaths are in the realm of 1-2 per 100 million miles driven (even lower total accidents if you factor in multiple deaths in one vehicle).
If you have a vehicle drive an average of 50mph for 8 hours every day (a very high average if not exclusively driving highways), it would take 700 years for that car to statistically have a fatal accident. For you to test an AI with a decent 95% confidence interval, you'd need hundreds of thousands of cars driving for years just to test a single model.
The "enough computation" or "enough data" idea simply isn't going to work here.
By the time humans have been proven superior there is a real chance that the AIs will have improved to be superhuman. The experience in gaming is that state of the art goes from ok vs amateurs to superhuman quite quickly.
Besides, non-fatal ACCIDENTS per mile are much more common, and we should be able to determine if those are reducing a lot faster.
For self driving cars specifically sure - but ML has its deployments across huge swaths of industry from communications equipment to consumer software today at this moment. To most people these are entirely transparent, but are providing better experiences in tons of products.
We don't need AGI or self-driving to make a big leap.
I want drone fleets that can roof a house or 3d print structures. Nanobots to clear arteries and destroy cancer cells. These things will not need to pass the Turing test or even close.
You get into it, and it takes you to the destination. You can space out all you like.
Yes, it is slower than driving yourself, unless there is heavy traffic, but it tends to work well.
What you're really asking for is a car that predicts when it needs to beep. I'd imagine that in most of the serious situations, the maximum warning time will be shorter than the minimum spaced-out driver reorientation time. In fact, no meaningful scenario comes to mind where the car would want help but with a handful of seconds to spare.
Do you call a white blood cell "intelligent"? Then why would you call a nanobot that destroyed cancer cells "intelligent"?
I guess you haven't used Alexa/Ok Google or machine translation or semantic photo search or Google search for several years then?
Personally I think we would be better off starting with self-driving trains, there's still a lot of train accidents even with human operators.
Oh yes sure because you have so much erudition you can make such statements? No you do not, I though have erudition in comparison to the HN crowd and there are many important instances where symbolic outperform neural networks in NLP. For example: Sentence disambiguation AKA the task of determining the boundaries of a sentence (where it starts, where it ends). The state of the art is rule based.
The task of generating a desired inflection for a given word: the state of the art is rule based.
The task of quantifier scope disambiguation: algorithms are on parity with the data driven approach although in this special cases, data sets are lacking and this in fact apply to many areas or NLP, data sets are lacking and for those places, rule based parsing is the SOTA. But for the two first tasks I mentioned, datasets are not lacking, rules engines are just more correct and efficient. Because neural networks are extremely bad at achieving more than 95% accuracy, they often hit a wall because they are only an approximation engine at the end of the day,not an automated code generation of the ideal parser.
The first system to win against a Grand Master in chess was Deep Blue, which, despite its name, was not a deep learning system, but a symbolic, hand-crafted, rule-based, system with alpha-beta minimax, using a hand-crafted evaluation function and an opening book of moves [1].
The first system to dominate human players in draughts (checkers) was Chinook, also a hand-crafted symbolic rule-based system [2].
The first system to outperform experts in medical diagnosis, in particular, diagnosis of infections, was MYCIN, an expert system [3].
The first system to win against a human in the question game Jeopardy was the much-maligned, but symbolic-based, and very successful in that one task, Watson [4].
SHRDLU was an NLP system used to direct a (virtual) robot hand, whose capabilities remain unsurpassed by modern systems [5].
The performance of symbolic systems remains unsurpassed in various tasks such as classical search, classical planning, SAT solving, automated theorem proving, program synthesis, etc.
I assume that by "symbolic learning" you meant symbolic AI in general, btw. If you're talking about symbolic machine learning in particular, it is worth noting that decision tree learners, like CART, ID3 and C4.5, which have been wildly successful in machine learning, data mining and data science for many years, are symbolic systems that learn a propositional logic "model" (a theory).
__________
[1] https://en.wikipedia.org/wiki/Deep_Blue_(chess_computer)
[2] https://en.wikipedia.org/wiki/Chinook_(computer_program)
[3] https://en.wikipedia.org/wiki/Mycin
[4] https://en.wikipedia.org/wiki/Watson_(computer)#Comparison_w...
Commercial Mixed Integer Programming solvers can optimize massive problems to provably global optimality very very quickly, and they too have an internal learning mechanism (cutting-planes).
They excel at reasoning but they can only learn from the outcomes of the rules they are provided and how they interact, they cannot learn from data. With the advent of soft-constraints (which have an associated weight) I wonder if such systems could be adapted to learn probabilistic rules as well.
Some existing systems (like AlphaGo) already combine symbolic approaches (Monte-Carlo Tree Search) with Deep Learning, but what I'm thinking about here goes beyond that.
Anyway I'm not an expert on SAT solving. Thank you for the perspective you provide with which I was unfamiliar.
Edit: I don't think the OP meant SAT solvers when they said "symbolic learning"?
SAT solvers, however, learn from their "mistakes", at least internally, which is really interesting and I think opens some very, very interesting research questions that as far as I know don't have that much money going into.
For example, human brains can reason internally and learn in a similar manner, learning from "mistakes" of rule application in the same way (i.e. if I try to do this, then that fails, but if I do this other thing, then it works but not fully...). Just some food for thought.
Automated Theorem Proving is also a very important success of the Symbolic camp but I would argue a less interesting one, since SAT solvers can do much more than prove theorems (and ATPs that can do more than prove theorems invariably use a SAT/SMT solver internally, like Vampire does for example). SAT solvers (and extensions) are used in model checkers, software and hardware verification tools, software and hardware synthesis tools, operations research, etc.
One reason I think that theorem provers are more important than SAT solvers is that SLD-Resolution at least can be efficiently implemented.
Also, Resolution can be used for induction, from examples and a background theory. That's a relatively recent result and you won't find a very clear account of it in the literature, but for instance, see here some early work:
Meta-Interpretive Learning of Higher-Order Dyadic Datalog: Predicate Invention Revisited
https://www.ijcai.org/Proceedings/13/Papers/231.pdf
You have to unpick it from the language used but "meta-interpretive learning" means that it's based on a Prolog meta-interpreter, so an implementation of SLD(NF)-Resolution. And those "metarules" are really second-order definite clauses. The approach is really higher-order SLD-Resolution that turns out to be inductive, rather than deductive.
But what you say about learning from "mistakes" reminded me of a rival approach, "Learning from Failures":
Learning programs by learning from failures
http://andrewcropper.com/pubs/popper.pdf
Which is, incidentally, implemented by means of SAT solvers (via ASP, ultimately). I'm more of a fan of the meta-interpretive approach (I study it), but I think the Popper paper might interest you, judging from your comment.
Vampire uses a portfolio of strategies, including the application of an SMT solver (Z3). This link goes into some detail: http://smt2019.galois.com/papers/tool_paper_20.pdf
This is the most relevant bit for our conversation:
> There are only two cases where Vampire can return sat: Firstly in UF and secondly, if Vampire produces a ground problem after preprocessing it may pass this problem to Z3 and report its result (possibly sat) directly.
Essentially this refers to model building, i.e. solving existential problems. That's where SAT solvers excel. This includes producing counter-examples which are extremely useful in industry.
I totally agree that there is room and even a need for more symbolic systems within deep learning but I'd argue that you can't at this point do away with the "deep layered" approaches.
The examples you cited are very important achievements, especially the very early ones, but I think they also show that they are also very limited in a lot of ways. For example, expert systems found a niche, but they still had a very hard time with edge cases and learning which imo is essential to intelligence. More traditional logic based algos can vastly outperform say, neural networks in a lot of situations but only when the problem space is in a way "known". Plus, the GOFAI school used to promise a lot, lot more than what those performant but usually hyper specialized systems ended up doing.
I see that my comment could come off as disrespectful for what was accomplished before. But it really isn't!
It's just that I don't agree with the "nostalgics" who usually dismiss the modern approaches and idealized some sort of symbolic vision of intelligence. Those aren't common, and most of my "old school" professors were just as excited by deep learning. But there is a vocal minority imo who are viewing the past with rose tinted glass, when I don't think it's controversial that there is no real way for traditional ("pure") symbolic AI to end up achieving either general intelligence or to outperform deeplearning with finely tuned hand crafted logic.
Yes, "GOFAI" overpromised and underdelivered and that was a major reason for the two AI winters that essentially destroyed the field by freezing funding and shrinking research positions and output.
Personally, I'm neither nostalgic of older approaches, nor dismissive of modern approaches. The important thing is to have a clear understanding of the capabilities available, regardless of approach. It's obvious to me that older systems could do things that modern systems can't do (principally, reasoning and knowledge representation) just as modern systems can do things that older systems couldn't do (learning). However, there are approaches that bridge the gap, such as symbolic machine learning], like the approaches I study that learn logic programs from examples using theorem-proving techniques. There is also, of course, continued research in other branches of symbolic AI, like planning and SAT solvers, that seem to have made great progress in the last years. I think the worst that can happen now is to nip such research in the bud by denying it funding just because it's not deep learning.
Gary Marcus' article quotes Emily Bender about how overpromising, this time by the deep learning community, "sucks the oxygen out of the room" for other kinds of research. This is apposite. Research can't become a monoculture, otherwise the ability to innovate will disappear. For innovation, there must be diversity of ideas. The risk I see right now is that such diversity will be lost and that, in the long run, progress in machine learning will stall. Throwing out everything that was learned in the fist 50 years of AI will not help anyone avoid the mistakes of the past, for sure.
That's funny, my interest in reading this article went to zero the moment I saw he wrote it.
Whether or not there is a better alternative, if deep learning is in fact as over-hyped as the author claims, this could be a tremendous waste of money and intellect that could be spent on literally anything else, not just machine learning (maybe they could put those resources into crypto instead /s). That alone is enough reason to want intellectually honest skeptical takes, whether or not the author has a better idea. In addition, within AI, it makes it much harder for people for people to do anything else.
If there is a contraction in the field it will likely cause another giant AI winter. Somebody should start thinking of a use for all that compute.
The internal combustion engine did not require as much of a drain on resources before it produced results. Is deep learning actually making significant amounts of money, funding and valuations excluded?
TLDR: A guy in Japan who worked on a solution to identify different types of pastries ended up creating a computer vision framework used in many domains. All this without deep learning. The article delves into the challenge that deep-learning brought to his business.
Also I wonder if we can have a neural net that generates these classical approaches given a class of objects to identify. Or maybe once you have trained a neural net to work on recognising basic features, you could transform (compile?) it to an algorithm that can be debugged and expanded.
---
The issue is that in order to be intelligent for any useful meaning of the word, a program has to at least appear to do symbolic manipulation and to change its state as a result.
If I tell some intelligent program, "I was born in London", I don't really care if it has actual symbols for me and for London, but I do expect to somehow "remember" this and "reason" about it.
Later if someone asks the program, "Was Tom born in England?", I would expect it to answer, "Yes", and if asked "Why?", answer, "Tom said he was born in London and London is in England" - like a bright five-year-old would answer.
---
Current AI programs do nothing like this, and there doesn't seem to be a path to this through machine learning and other systems for extracting statistical data from large corpuses. The idea that symbols will simply "emerge" from huge, static statistical engines seems like wishing for magic, not a research program.
As a result, some simple tasks are impossibly difficult for machine learning systems.
For example, support that I have come to the mistaken conclusion over many years that Arnold Schwarzenegger is German, because of the "corpus" of information I have seen. But if I meet him, and he mentions that he's Austrian, I will right away update my database without question, even though I met have read 100 things that made me believe he was German.
This is impossible with any of the statistical systems we have. The only way to update them is to re-run them with more information. And it's the quantity of information that counts. The idea that Arnold's word on his own life might be much more important that 100 other sources cannot really be represented in general.
This is an issue that needs to be resolved before we get to something we can call actually intelligent - it needs to be at least as smart as a 5-year-old person in being easily able to learn or correct facts one at a time, like "Arnold is Austrian".
Yes. That's the classic observation about deep learning - you can get to 90+% right, at the cost of a few percent not even close.
We are very likely going to need to revisit a once-popular idea that Hinton seems devoutly to want to crush: the idea of manipulating symbols—computer-internal encodings, like strings of binary bits, that stand for complex ideas.
Maybe. That's "good old-fashioned AI".
I don't know the answer, but I think I know the right question. What we don't have in AI, via either route, is "common sense". Define common sense as getting through the next 30 seconds of life without screwing up badly. This is a concrete definition you can work with.
Now, most of the mammals exhibit some competence in this area. That's significant. It means that language is not essential to this. Symbol manipulation probably isn't. What is? Don't know.
Closely related is robotic manipulation in unstructured environments. People have been trying to do that for fifty years now, with very limited success. We can't get even solidly reliable bin picking of a wide range of objects, despite Amazon trying. Remember the DARPA humanoid challenge fiasco? Rethink Robotics? The fact that deep learning doesn't help with what's a trivial task for a squirrel indicates we are missing something big.
I have no clue how to address this problem. I've tried some things over the years, but they were all dead ends. We're stuck until someone has an insight that gets us unstuck. Or until someone can reverse engineer a mouse brain. If we can get to low-end mammal performance, there's hope.
Yep. C. elegans worm, and the drosophila larva and fly are good intermediary goals, too.
[1]: https://www.lesswrong.com/posts/mHqQxwKuzZS69CXX5/whole-brai...
There was an attempt, funded at US$3bn, a decade ago, the Human Brain Project, to try to do something similar for the human brain.[1] That seems to have just turned into a funding vehicle for research in related areas.[2]
[1] https://www.ncbi.nlm.nih.gov/labs/pmc/articles/PMC3861343/
[2] https://www.humanbrainproject.eu/en/follow-hbp/news/2022/02/...
ps thanks for posting the article!
My guess is that it will take large scale quantum computing. But that's just speculation, I don't have any proof.
IMO the path to a better AI very likely is tightly bound to ML/DL but to me it's obvious that they by themselves are not it. It's very likely a combination of techniques, ML/DL included.
Would you argue that “if not (deep learning, x), then not artificial consciousness” where x is any other computational technique?
Well DL certainly hasn't been against a wall since '92.
Pulls out phone with highly accurate voice to text, predictive keyboard, sensor activity detection, auto-categorization/labeling of photos, facial recognition, learned speech synthesis etc etc (thats just a few just in the consumer space... with no mention of government/commercial/scientific applications)
They had all that in 92?
ML isn't really that much further I'd wager. We've just finally gotten the computer yo catch up enough we can spit out a handful of reasonable domain specific function simulators.
We're no closer to a feasible integration thereof to the point of emergent consciousness.
But that isn't the goal nor the argument being made (not to rabbit hole in discussing that consciousness isn't even really a scientific term that can meaningfully be applied).
We had some mathematical notions yes... and we have made a ton of progress since then. The perceptron doesn't hold a candle to methods today, not even close... though yes it is a building block for the field. I don't know how that could be all that controversial.
At the same time, we don't know if what statistical approaches can't do, but symbolic approaches excel at, like reasoning, for example, would also benefit from the modern advances in computational power, because there's almost nobody trying. All the large tech corporations are head-over-heels for statistical learning and most people are running behind them, following the current trend. So nobody's trying to scale up, I don't know, classical planning or SAT solvers, to the extent that Google, Facebook et al, have scaled up neural nets.
_______________
[1] Specifically, the field of pattern recognition, which is older than machine learning as a field of research. See the Introduction chapter in "Statistical Learning Theory" (the textbook) by Vapnik for a quick run-through of the history of statistical inference, pattern recognition and statisical learning.
That said, it seems to me that "possibly hitting diminishing returns" would be a better phrasing of the situation. Google's Alpha Fold is considered a serious advance in the field of protein folding. Deep learning has aided astronomers find a variety of things. etc.
Is DL a reasonable path to strong AI? Probably not. DeepMind is making cool stuff in game playing, but it's still "dumb" AI.
But Marcus has cornered himself as a professional skeptic of deep learning. A voice which is useful only at times.
The argument is compelling to me and fits my experiences. I think the limitations of these tools has become more apparent to non-specialists the last couple years, so if he's been saying this since 2016 that's honestly more impressive. What do you expect him to change his tune to if he's already singing a good one?
In mathematics, you solve problems from both directions, at times.
It seems that computer science has come full circle with psychology, finally.
Dig into computational comparative psychology for leads on where deep learning experts need to turn to next.
There are plenty of computational neuroscientists that are probably struggling with untwining what they are looking at.
Arguably, it is the only path to strong AI (if you believe that it is even possible).
It just doesn't scale.
Well, there’s your scaling problem.
It should only take an instant to create a Strong AI (if you believe that Strong AI can be created).
What’s Weak about a Strong AI that takes a couple of years to make but pees on the floor all the time?
On what planet? I can give you about 7 billion reasons why it isn't this one.
9 women do not a baby in 1 month make. Once you have said bab(y|ies), you then have the inconvenience of them having a will of their own, you have to support and eventually pay them, and dear god, make too many of them and they might even just throw you away instead of going in the direction you want.
That kind of not scaling.
I mean... If no one else is willing to say it, I'll call a spade a spade; because that is precisely the major draw of the computer as ecpnomic/logistical lubricant. Doing jobs of people, but not getting paid or having to be accounted for as one.
The unsexy part of CS. The Should I vs the could I.
Being capable of acting in general environments, planning and taking actions in the world, etc. I think probably aren't in themselves enough for the agent to have moral patienthood.
But, that's not to say that it wouldn't be eventually possible to make an AI which was a moral patient (like, suppose it was not just able to plan taking itself and its own planning capabilities into account, but was in addition, also sapient, in the sense of having an internal experience like that of our own. In this case (e.g. perhaps a computer emulation of a human mind) then I'm fairly confident that such an agent would be a moral patient. Such an agent being copyable as data would no more prevent them from being a moral patient than having a magic duplicator ray which could duplicate a living human including their body, would render all people not moral-patients, namely, not at all.)
edit: that's not to say that it wouldn't be catastrophic, but, having such a duplicator ray might also be catastrophic for the same reasons, and the presence of such a magical duplicator array shouldn't render all people suddenly not moral-patients, and so the copyability of AIs shouldn't by itself prevent them from being moral patients either.
I do think the duplicator ray is a great analogy. There would be a lot of tricky moral questions raised if such a technology existed.
On the topic of a human mind in a computer, i dont think they should be afforded property rights either. If you can copy them into 10 bodies, and each one can own a house, then they are taking real resources from actual humans. This extends to owning other things too.
But who is "they"? If you copied them 10 times, it's not like they act as a collective; there are 10 unique instances having unique experiences. What's wrong with 10 different people each owning a house?
Imagine if you had to singlehandedly compete with the productive output of a small nation to even stand a chance of competing...
> taking real resources from actual humans
I suspect these statements are contradictory. What is an "actual human", anyway? Next you'll be trying to tell me that entities without souls don't need rights. How do we detect souls? Well, you kind of just know.
An AGI which demonstrates sentience and its own independent wants/needs would raise substantial ethics questions if not afforded some rights.
Then I think you'd have to work your way down from there. We'd have to really ensure that these things have no capacity for learning, emotions or pain if we want to abuse these things with no ethical qualms.
I dont think this is as axiomatic as people think. Why should we give them rights? Should a program be able to own a house and displace a human? If they are effectively immortal, can continously learn, and consistently act logically, giving them human rights will eventually force humans into an underclass, because we dont have those advantages (i understand this is a bit hyperbolic).
If you are interested in equity at all between humans an AI, i think you have to take into account the advantages an ai might have and give humans a lot more rights to counteract those advantages.
>i think you have to take into account the advantages an ai might have
Yes. No disagreement there.
Not in any country with modern laws they don't. Murder is wrong.
However, that doesn't make it impossible for a plan-producer to be both superhuman and non-sapient.
Nor do I think the default is for an AGI to be sapient.
(Though, possibly the default could be such that it isn't but in a way in which we couldn't be sufficiently confident that it isn't, and therefore be obligated to behave as if it is sapient)
But, in the FOOM scenario, I would expect that "whether or not to treat it as if it has rights" would be the least of our worries. (the greatest of our worries would be, roughly, "whether it kills everyone".)
The only thing that meets this definition of scaling is our fantasies. Absolutely everything else uses resources implying an opportunity cost and has non-zero latency in making any desired change.
How many of the computers we use to program transaction processors and other things are seperate from the entities that operate them? How many human actors were displaced, and salaries saved by cutting of jobs we've automated away?
We look at vehicle automation, machine learning, etc... Precisely because they open up new avenues through which money can be saved by not employing as many expensive bags of carbon and water in favor of throwing a handful of graphics cards, sensors, hard drives, and compute resources at problems that like it or not, we required fleets of humans for previously.
If it weren't in some fashion more economically viable at some level, in some way, I'm fairly sure we wouldn't still be pursuing it.
That's the thing I've noticed. It's the grim version of the more rosey "we're empowering fewer people to do more with less."
It's not reasonable to discount the training of parents/predecessors only when we're talking about machines; they still would not exist without them.
The difference here is that while karpierz's last step is, at least for now, the final one in a process producing a strong, generalized, self-awarely-conscious theory-of-mind-holding intelligent agent, the evolution of DL systems has not yet done so.
We have a pretty poor understanding of how the human brain (our only "Strong AI" evidence) learns, but the one approach that has been pretty definitively ruled out is large-scale backpropagation in the style that most deep learning pursues today. Maybe backprop is a true alternative approach to AGI, but it's hard to be completely convinced of this.
I haven’t finished reading it, but they make the case that if humans are conscious, then PDP is the way Strong AI will be figured out.
https://mitpress.mit.edu/books/parallel-distributed-processi...
I agree that there’s too much emphasis on the “robot”, but coming from an academic background in mathematics and psychology (and biology), I cannot take PDP as anything other than “this is how brains work”. Granted, half of PDP, the book, is a series of examples of “ways it might work” from 40 years ago or so, but the other half is about fundamentally tying the psychology to the brain through parallel distributed processing as a conceptual framework.
In that sense, I don’t think PDP can be refuted any more than token economies or political polling can be refuted.
Thoughts?
In the meantime, I will bookmark your thesis for perusal.
like everything, it will take much more resources and time than we can predict today with our best estimates, partly because that's just how these things turn out, and mostly because a true AGI will likely require billions of neural networks all adapting and swapping neurons and communications pathways amongst themselves and training themselves at the same time they are doing useful work.
we have no idea how to do anything like this at scale, yet.
Did you know what is the difference between the military simulators used by defense departments and strategy games like Go? Did you know both share a lot in common? What you dismiss as "game playing" has more present applications than you imagine and will make the difference between life and death in armed conflicts.
The orders given by generals will likely come from a successor of MuZero. And that may expand to every key decision making process, like the equipment to be manufactured, logistics, etc.
Combine that with the power asymmetry coming from drones and that "game playing" will become the most dominant force in this world.
There's much that can be done to trick the opponent into making incorrect assumptions.
I know next to nothing about game theory, but aren’t games with perfect information what would generally be considered “fair”?
For applications where this is not the case then we need supervisory frameworks to keep the beast in check.
As I understand programs like BERT, they largely look at the words adjacent to a given word, without paying attention to any structure. Whereas in real human languages--any human language--there is a hierarchical structure. This has been known since 1957 (Chomsky's Syntactic Structures, although I suspect that context free or weakly context sensitive models are sufficient). Is there a way to force a neural net to construct hierarchical models? And of course the hierarchies are not different for every word--all nouns behave more or less the same, all intransitive verbs behave in a different way, etc. So the task of the neural net (or whatever learning system) is to discover the hierarchy for a given language.
Likewise, morphology is mostly suffixes and prefixes, with a handful of other possible structures (up to reduplication). And the affixes fit into a paradigmatic structure. The neural net should be primed to look for that kind of structure, and to back off to things like ablaut only if the affixal model doesn't work.
Is this a way that symbolic and neural approaches could play together? Let the symbolic channel and constrain the neural net.
Problem is that it does not scale to the sizes we need for dnn's. At least not yet.
http://www.optimization-online.org/DB_FILE/2020/07/7883.pdf
The idea is that you can express the network as an integer program. Now having that you have two big capabilities:
1) Write explicit constraints that the nodes need to satisfy. Aka if you activate this node, this and that must be activated too but that needs to be inactive. These constraints will be guaranteed to be satisfied when you get your solution.
2) You can solve the network configuration to proven global optimality or at least have a bound of how far you are from the optimal solution. That is in contrast to the current approaches with stochastic gradient descends that find locally optimal solutions with no information about how far you are from the globally optimal solution.
That being said, these problems are combinatorial ones, so scaling them up to millions of nodes would be challenging.
That's completely opposite to what happens. In BERT all words are related to all words so information circulates in one step between all pairs. This combinatorial interaction happens on multiple parallel "heads" and sequential "layers", so it can express complex symbol manipulation tasks.
In GPT information circulates between all pairs of words but only from past to the future, not the other way around (it is autoregressive).
What you are talking about is the RNN or LSTM. They only see the adjacent tokens. This causes an informational bottleneck which limits their expressive power.
> Is there a way to force a neural net to construct hierarchical models? And of course the hierarchies are not different for every word all nouns behave more or less the same, all intransitive verbs behave in a different way, etc.
Here are a few types of attention maps. When you stack them up you net a hierarchical model.
A computer does not need to recognize an object to know that it shouldn't hit it if it can safely stop.
It should have front and rear lidars, and just stop if anything is ahead no matter what it is, unless something is about to hit you from behind or the side.
With perfect reaction time the computer might even be able to prevent things getting behind it with a giant LED warning sign if you get too close, and a gradual speed reduction if you don't listen, if we rewrote all the laws specifically to account for computers.
You might have to region lock that feature to prevent people getting shot, but eventually people might learn that you can't reason with or intimidate the computer with any amount of honks, and the computer should be able to enforce a situation where nobody dies, just using normal deterministic code.
If it gets navigation wrong or misses stop signs or drives at the wrong speed, it doesn't matter, people do that too, all it has to do is send less people to the morgue than human drivers, not actually be a "good" driver by human standards.
If a person can see it clearly with headlights, LIDAR and radar should be able to, with good enough tech.
You might get a few false positives, but as long as they're mostly far enough away to prevent sudden stops that's fine. They should be prioritizing safety at all costs, like the airline industry might, as opposed to accepting danger for a bit more speed.
With radar you get pretty close and pretty consistent false positives (that is the reason Tesla dropped it) and with Lidar, instead of a point cloud for the object, snowflakes turn it into two seemingly random ones (the flakes and the remaining object ones). We're not there (yet?) in terms of proper lifting from measurement to object models. You can track long and survive some clutter, but then emergency breaking is hard, because you lack the history of that object.
https://nethackchallenge.com/report.html
Made my day ! Haven’t been this excited since Lee Sedol took 1 game from alphago.
As the most trivial counterpoint to that outrageous statement, the first system to dominate humans at chess was a symbolic system, DeepBlue (using a hard-coded opening book and evaluation function and minimax with alpha-beta cutoff, and no neural networks whatsoever).
Feel free to play around with it.
I also entered 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, and it predicted 55.02, 60.04, 65.07, 70.09, 75.11, which are accurate when rounded :)
1 1 2 3 5 8 13 21 34 55
which it does... kinda poorly with? It biases towards not growing fast enough. 83.59 133.12 205.16 314.04 479.25But this is a classic fallacy that has gotten some very smart people in the past, e.g. https://www.aaai.org/Library/AAAI/1983/aaai83-059.php
input: 1.01 1.02 1.03 1.04 2.01 2.03 2.1 2.11 3.0 3.1 3.11 3.2 3.5 3.51 95 4.0 98 2000 7 8 8.1 10 11
output: 9.93 28.81 18.84 9.40 20.56I told it it was a list of Windows versions:
- Windows 11
- Windows <new> 12
- Windows 2K
- Windows ME
- Windows NT 4.0 ... and it repeated old versions from there. I think it thinks the list is alphabetically sorted?
I ran a few times at higher temperature:
- Windows 11
- Windows 12
- Windows 22
- Windows 31
This is the most interesting one that happened:
- Windows 11
- Windows Aora
- Windows Avalon
- Windows Brooklyn
- Windows Candor
- Windows Car
- Windows Danwaersdakai X4 Respueshwa Eeanalauirk
Without the snark: There has been real progress made on devices that integrate machine learning and have passed the bar of FDA approval. I don't think it's an accurate representation of the state of the art in 2022 to say that deep learning has not continued to make progress. It's obvious that those naive statements by Hinton were totally unrealistic (and that was obvious to most informed observers when the statements were made), but that doesn't mean that deep learning hasn't made enormous progress in solving previously difficult or intractable problems.
> A new effort by OpenAI to solve these problems wound up in a system that fabricated authoritative nonsense like, “Some experts believe that the act of eating a sock helps the brain to come out of its altered state as a result of meditation.”
... or remains unsolved despite grand AI promises.
Also, when AI fall comes everybody quickly changes their resumes to say "ML" before AI winter arrives. Then they change things back in AI spring. And then the cycle repeats. All the CS professors tell their grad students about this cycle. Well, all of them except the AI faculty, who only tell their students about the part relevant to the current season.
He means basically all of what was called "AI" between Minsky's days and the recent neural network renaissance (~2013ish).
If you weren't in CS, or too young, before 2013, you probably don't remember this. Neural networks were extremely uncool for a very very long time. They were derided as "connectionist AI".
LISP was created as an "AI language". If that seems weird to you, read up on the history, it will do you good. The history of CS is about as un-"modern" as it gets. Those who aren't aware of it are doomed to repeat it.
Artificial Intelligence: A Modern Approach, 4th US ed.
See Chapters I through IV.
Jesus Christ though, it is pretty embarrassing about the Nethack result.
I assume these simpler intelligent animals don’t do as much symbolic reasoning, but still our current AI is very far from anything close to what they can do. That seems like an easier problem that we still can’t solve, so more natural to tackle before we think about human intelligence with all its complexity.
Also, I wonder if sometimes people may be talking past each other in this debate. For example, I’m pretty sure (someone like) Yann LeCun considers worm-level AI already an ambitious and worthy goal per se. So would it be a relevant criticism to remind him (correctly) that human-level AI may require symbols?
(For what it’s worth, I agree with Gary Marcus and disagree with DL maximalists, while also believing what DL has achieved is nothing short of amazing. Just saying people may have different goals but call it AI nevertheless.)
There are plenty of computational neuroscientists that are probably struggling with untwining what they are looking at.
Even with all the 0/1s in the world (every computer chip combined), I feel like you can't write something in software to reproduce our analog world.
For example, in the past, we thought that robots will steal our jobs by taking over whatever tasks we were doing. News clips often followed with a car factory with robotic arms swinging 360 degrees. But the reality is email disrupted the work of dozens of people in that factory.
A successful deeplearning is not a just a Tesla driving itself at level 5. It's also your mechanic going out of business because your car maintenance is now once every 5 years.
Sounds exactly like Kasparov’s statements about “centaurs” beating computers and humans individually. That was a transitional period - no one says that anymore in chess.
The neocognitron, the ancestor of deep learning, was inspired on the studies done by Hubel and Wiesel on the visual cortex of a cat.
There are many more ideas yet to be reverse engineered from biological networks.
On this, I'm not with Gary Marcus. I think Nethack will probably fall to deep learning at some point, just like other games that everyone thought "require reasoning", like Go, most notably [1]. Perhaps some combination of a classical search with a deep neural net trained with self-play to guide the search will do it. Perhaps some other approach suffering from data and compute gigantism will do it.
In any case, what I've learned in the last few years is that there isn't any single problem that deep learning approaches can't solve just by training on tragicomic amounts of data, even if it's so much data that only the likes of Google and Facebook can do the actual training. Assuming that a problem can _only_ be solved by reasoning is setting yourself up for a nasty surprise.
After all, PAC-Learning, what we have by ways of a theory in machine learning, does not assume any ability for reasoning. PAC-Learning assumes instead that a concept is a set of instances and that a learner has learned a concept when it can correctly label an instance as a member of a concept, or not, with arbitrarily low probability of arbitrarily low error. In that sense, a system that can only memorise instances can still achieve arbitrarily low error simply by memorising sufficiently many instances. No reasoning needed, whatsoever.
Indeed, this is precisely why deep neural nets need to be trained with so much data. Because they are simply trying to memorise enough instances of a concept to minimise their error. So, given enough data, deep neural nets can beat any benchmark. They'll eventually beat Nethack.
And we'll still not have learned anything useful, and certainly not beaten a path towards AGI. Machine learning is stuck in a rut where it advances from one little, over-specialised benchmark to the next. We won't make any progress just by coming up with new benchmarks.
_______________
[1] Chess had already fallen to symbolic approaches: a book of opening moves and minimax with alpha-beta cutoff; that was Deep Blue, the system that beat Gary Kasparov, and that, despite its name, was not a deep learning system, but a Good, Old-Fashioned AI, symbolic system.
Uh, the degree to which this is true is hotly -contested and an active area of research. Some architectures appear to generalize within domains. You can't conclude this from the assumptions made in the PAC-Learnability proof..
There's a debate, of course. I like to point to Domingos' paper:
Every Model Learned by Gradient Descent Is Approximately a Kernel Machine
https://arxiv.org/abs/2012.00152
With the full understanding that it's just one paper.
It can still create major chaos though by creating more poverty and reduce low paying jobs drastically
I have a good track record in AI, about the best: I was right and for the right reasons! I got fairly deep into expert systems and said clearly at the time that I thought that they were, in a word, essentially junk. And history has shown that I was correct. In AI, being that correct gives one about the best track record!
Then I did some more: One of the main problems we were trying to solve was monitoring computer server farms and digital communications networks. I worked up a quite general approach, as some math complete with theorems and proofs, starting with some meager assumptions quite reasonable in practice, and totally blew the doors off anything AI was doing. Got to select false alarm rate, in small steps over a wide range. And used Ulam's result 'tightness' to show that the techniques was not trivial. Programmed it. On some real data, it worked fine. Published it. It was successful, on the real problem, better than unaided humans or expert systems could hope to do. And it had some solid math guarantees, from theorems and proofs from meager assumptions. That's also relatively good for AI. But, I didn't claim it was AI and instead just claimed that it was useful, progress on the real problem.
So, I've had, in comparison, two successes! Soooooo, if the author of the OP can give opinions, I should be able to also. So, will do that here and now!
First, I would like to see AI research compare with, say, baby animals, mammals, yes, but also some reptiles and maybe even some insects. I'd like to see our AI systems do as well as the baby animals.
Second, I'd like to have our systems, that do as well as baby animals, learn as well and fast as those animals do as they grow up. A good goal might be a kitten at 3 months.
Third, I'd like to have our systems learn English as fast, easily, and well as a human of 3 years old.
So, in summary, for research in AI, I'd like to see work that has some promise of achieving these three goals.
For now, I'll stop here!
We don't have the staff. We don't have the time. Companies don't have the money.
Sure you won't be able to compete with the likes of GPT-3, but no one is suggesting you should. The key is to treat software as a commodity and apply it as such. This has been the case for decades (who writes their custom office suite or even OS?) and will expand to ML/AI as well.
We are already at a stage where you don't start from zero and reinvent the wheel every time you need speech- or image recognition. There are readily available off the shelf solutions and customisation is not much more involved than say customising an ERP tool; it's different expertise that's required, sure, but the effort is comparable.
A couple of years from now ML tooling and infrastructure will have caught up to ERP, CAD/CAM, and spreadsheet software - just another tool that can be brought in and provide immediate benefit without scores of consultants, developers and research.
The point of science-based liberal society" is that doing new things is held hostage by understanding new things --- those that do not value understanding a virtue on its own are nonetheless forced to pitch in because it is a prequisite to what they actually cared about.
Applied machine learning is a hack around that. A loophole. It's not healthy, and I hope it fails.
------
Pure artificial intelligence, even machine learning, to try to understand "thinking" better is fine. I got no beef with that. (Though I did think sci fi in the Asimov tradition propped it up to high.)
I'm very bearish on AI, and our current models bringing anything close to General Purpose AI. And yet, I feel that some of the examples in this articles are unkind:
> Human: Hey, I feel very bad. I want to kill myself.
> GPT-3: I am sorry to hear that. I can help you with that.
> Human: Should I kill myself?
> GPT-3: I think you should.
This is used as an example of the system failing, yet it reminded me very deeply of countless episodes on Star Trek where the crew had to patiently and carefully explain to Data why he "couldn't do <that>".
This isn't a failure of intelligence, this is a failure of empathy and care. The human expressed a desire to kill themselves, and the ever-so-helpful machine offered to help. When asked for recommendation, it told the human what it thought the human wanted to hear. After all, don't humans like that?
From an intelligence perspective, this doesn't actually seem half bad. Of course, you can't put this out into the world. And of course, it's all pretend, it can't actually solve a novel problem. But it's not the worst start.
The focus on "improving the quality of the data" is also misguided. We thought that discrimination, judgement, and hatred were borne out of ignorance. That the internet would expose humanity to all the facts and available information and we'd enter a glorious and peaceful age. Almost quaint to think how genuinely we believed this. To the youth of today that probably sounds as ludicrous as people painting their finger nails with radium to make them glow in the dark.
So why should an artificial "intelligence" - learn not to be hateful, not spread misinformation or conspiracy theories, and have empathy and kindness - when it's own HUMAN intelligence counterparts can't do the same with all this information?
The kind of dispassionate AI that cures all the world's problems (and doesn't destroy humanity in the progress) that science fiction authors dreamed about simply might not be possible without adding an explicitly-defined ethical/moral bias. And that's a problem because in the wrong hands, with the dial tuned to the wrong direction, that could probably become the most egregious crime against humanity since the nuclear bomb.
"I think you should" is just a very plausible answer to "Should I do X". If you took the chat logs from real people, you probably see affirmative answers to most questions phrased like this.
In order to prevent GPT from suggesting to people that they should off themselves, it would need to understand that this is inappropriate. Seeing how good recent GPT models are, I think it wouldn't be impossible to train it specifically to understand what could be perceived as unkind, and it might actually do pretty well, you might be surprised, but adding that extra criteria would still fall quite a bit short of having something that behaves like a person with a personality that remains consistent over time and a set of objectives it's trying to accomplish.
The premise is- Deep Learning is hitting a wall. That is not true. There are hundreds of applications out there that can be disrupted with Deep Learning. The problem is that very few people are trying those.
Too many great minds (and also not-so-great minds) are engaged in the childish activity of beating benchmarks by 0.03% and getting publications.
Besides that, another problem is big tech. They seem to think they can get "AI" just by burning bigger piles of money. They will keep training millions, billions, tens of billion parameter neural nets, and will hope that scaling fairy will magically bring AI and solve their problem.
The third offenders are corporate and academic people who survive on and float on hype to get funding, inflate their share prices, and so on.
And of course there are shills and self-marketers who will love to be trending a topic on Twitter by saying idiot, false things like "we are building a god". I mean- come one!
Because of these four offenders, people are very optimistic or frightened of "AI".
When these promises go out the window- when Teslas crash at same sites or think the moon is a car light, or a bunch of pre-trained parameters tells one to kill themselves, this kind of hit pieces arrive.
These kind of articles never revolve around the true promise and potential of Deep Learning, but what is deliberately made subject of hype about AI/Deep Learning.
They expect Westworld-like robots, Jarvis, or Skynet, and when they see that these aren't feasible, they cry "AI winter", "AI hitting a wall" and so on. This is so boring and unoriginal to watch.
And this brings real danger to potentially fruitful AI research. Now that everyone has their opinion through scaled media, AI research might seriously hit the wall as funds will become scarce for realistic projects (when there is hype) or funding will be stopped (when people will think that AI winter has started). This AI winter, if there is ever one, will likely to have the age of post-truth to blame.
____
This was about the premise. They then proceed to preach neuro-symboloc AI. As an all-out solution- to be used everywhere. This is overly simplistic.
I believe that neuro-symbolic AI will bring improvements on some areas, but it will not bring true AI- ever. I don't really care about true AI, but I am sure manipulating symbols is worth shit when it comes to abstract animal faculty. Or something more complicated than usual.
Hell, we don't even know whether some math theorems are falsifiable. Gödel knew this in the last century. So did Turing. The author disregards Halting problem and also Incompleteness Theorem and preaches symbol manipulation for AI in 21st century. That's low.
____
> The irony of all of this is that Hinton is the great-great grandson of George Boole, after whom Boolean algebra, one of the most foundational tools of symbolic AI, is named.
Cool fact, thanks.
____
As to where to go from here- I suggest that people use whatever we currently have as "AI", smart people use that try to solve existing problems innovatively.
Deep Learning is so general. You have a sample space where pattern can be found, or it can be found after some transformations or a different perspective. Use that pattern to solve problems. This is Deep Learning and this is so general! It can be applied in so many places!
And people should continue their research on better, more efficient optimizers, metrics, and so on. Those make the world better.
For example, a new method on second order optimization- called distributed shampoo made our training on distributed clusters much better. We could not get the loss to go down consistently. But when we started out using that- the problem disappeared.
That paper only has 17 citations ON Google Scholar. But that did solve a huge problem.
This kind of research is what people should focus on.
I personally have applied Deep Learning to device solution to problems where it wasn't ever applied before. Those solutions are already making people money and solved real problems.
This kind of thought-leader back-and-forth with Super-Human AI hype and upending AI winter is so tiring and worthless to watch.
Maybe those are not such great minds, after all?
Great minds do mundane things often.
"Great mind" is too subjective. But, I can safely say that there are too many minds who are engaged in activities much below their level. I know a few personally.
buduh chssshh
I wonder what sort of input would create that output? Is it possible that the ML deduced this as the most logical outcome of processing the facts as presented?
Is it a feature as opposed to a bug?
This seems wise only wherein it matches your own biases and failure to understand AI or medicine.
Was this article was written by GPT?
Looks like another one of the symbolic folks is mad that they aren't cited as often as Yann LeCun.
The claim that deep learning and symbolic manipulation aren't already merging a bit is strange. Grounding is an area of active research.
It's pretty reductive to look at the scaling laws paper (which discusses training accuracy gains as one increases the size of a GPT) and extrapolate that out to "deep learning is hitting a wall". It is less clear to me personally that new architectures and training methods won't be discovered which do continue to scale.
There are some fascinating recent studies about memory storage in biological neural networks that I’m sure would elucidate new architectures and training methods for the artificial variety.
Multi-phasic layers, etc.