We can’t trust AI systems built on deep learning alone
technologyreview.com
technologyreview.com
This right here is the soft underbelly of the entire “machine learning as step towards AGI” hype machine, fueled in no small part by DeepMind and its flashy but misleading demos.
Once a human learns chess, you can give it a 10x10 board and she will perform at nearly the same skill level with zero retraining.
Give the same challenge to DeepMind’s “superhuman” game-playing machine and it will be an absolute patzer.
This is an obvious indicator that the state of the art in so-called “machine learning” doesn’t involve any actual learning in the way it is normally applied to intelligent systems like humans or animals.
I am continually amazed by the failure of otherwise exceedingly intelligent tech people to grasp this problem.
Humans are also not a general intelligence.
In certain sense Deep Reinforcement Learning is actually more general than human intelligence. For example, when playing games you can remove certain visual clues. It makes it almost impossible to play for humans, while Deep RL scores will not even budge. It means that Deep RL is more general, because it does not relay on certain priors, but it also makes it more stupid in narrow domain of human expertise. Try this game to see yourself: https://high-level-4.herokuapp.com/experiment
Here is bike with reverse steering: https://www.youtube.com/watch?v=MFzDaBzBlL0 Here is flipped vision experiment: https://www.youtube.com/watch?v=MHMvEMy7B9k
Human brains are amazing, but they also require certain amount of time to retrain when inputs/outputs are fundamentally changed.
PS. I didn't hear about anyone testing different board sizes with AlphaZero-esque computer players. But I saw Leela Zero beating very strong humans, when rules of the game were modified so that that the human player could play 2 additional moves: https://www.youtube.com/watch?v=UFOyzU506pY
Playing chess well is a combination of both conscious and unconscious skills. However when deep learning systems play, it is all the unconscious, automatic application of statistical rules. They are playing a very different game from the human chess game.
Because there is no abstract reasoning involved here, these systems cannot apply the lessons learned from chess to another board game, or to something completely different in life, which humans can and do. So even though they are much stronger than human players, they aren't strong in the same way.
Interesting. Has this actually been shown? I would assume a lot of the strategies a human is familiar with would fall apart as well. I'm no chess or go player but I would have to learn new strategies in a tic-tac-toe game scaled to 10x10. I would certainly not be as proficient although I would still consider myself to have intelligence.
If you’re still not convinced, I’ll prove that skills transfer by playing bullet against anyone who can make a 10x10 variant playable online.
[Edited to come across less egotistical]
Source: am a master, rated 2500 in bullet.
Please don’t attempt to twist my words in order to support your own bogus position. Act like a chess master and just resign already.
I'm not saying I'm definitely going to do this, but is there a rulebook somewhere for 10x10 chess? (What are the initial piece positions, and how would castling work?)
What I really want is a 10x10 or even 8x10 board using the original set of pieces. This would be sufficient to prove that human chess masters can adapt in a way that machine-learning based algorithms cannot.
Despite the constant change of the card pool, and also the wording of the rules text on the cards, and the rules themselves, human players are perfectly capable of "picking up a card they've never seen before and playing it" correctly.
[1] Detail on bughouse in this comment from an earlier discussion: https://news.ycombinator.com/item?id=20831586
Meta learning is for solving similar problems from a distribution (like different sized boards in your chess example) and has taken off recently (only baby steps so far though). Modular learning is also becoming big, where concepts that are repeatedly used are stored/generalized.
Train it on variable spaces, and you'll get an agent that can play on variable spaces. In fact, you can probably speed things up drastically by using transfer learning from a model which already learned 8x8 space and modifying the inputs and outputs to match the new state and action space.
What part of this do you think "exceedingly intelligent tech people" aren't grasping? Something qualitative? Do you think people in machine learning think of "learning" as literally meaning the same thing as the colloquial usage? What, precisely, are you attacking here? All the harsh anti-machine-learning viewpoints with no clarity are becoming exhausting.
Certainly, someone close enough to the technical process of deep learning will admit that it essentially an extension of logistic regression without any "larger" implications - at least some deep learning researchers are always clear to distinguish the activity from "human intelligence" (and even if a given research never parrots the hype train's mantra, they know it's there and inherently play some part).
But more a minimum assertion of deep learning is that it "generalizes well". And what does "well" mean in this context? In the few situations where data can be generated by the process, like Alpha-Go, it can make a good average approximation of a function but in most situations of deep learning it means "generalizes like a human" - especially image recognition.
This comes together in the process of training AIs. Researchers take data that they hope represents a pattern of inputs and output in a human decision making process and assume they can construct a good approximation of a function that underlies this data. A variety of things can go wrong - the input data can be selective in ways the researchers don't understand (there was a discussion about a large database of images from the net being biased just by the tendency of photographers to center their main subject), there can be no unambiguous "function" - loan/parole AI that's inherently biased because it associated data that isn't legitimate, objective criteria for the decision sought), and so-forth. Some tech people are aware of the problems here to but this stuff is going out the door and being used in decisions affecting people's lives. Merely noting possible problems isn't enough here. These "exceedingly smart people" are still handing off their creations to other people are taking them as something akin to miraculous decision makers.
Please refer, precisely, to my earlier comment in this thread.
Humans have orders of magnitude more neurons, more complicated neurons, more intricate neural structures, and their training data is larger and more varied.
In contrast to “machine learning” which is merely a fancy way to say “data processing with massive compute”.
Maybe, but they’re certainly not described that way by whoever is in charge of publishing DeepMind’s research:
“A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play”
https://deepmind.com/research/publications/general-reinforce...
The learning algorithm AlphaGo uses is somewhat general, and can handle different games (e.g. you can put chess or Go through the algorithm and it functions well for either).
The output of this algorithm, however, is a specialised agent. The agent is not general. If I create a chess agent and give it Go or chess with different rules, it will perform very poorly.
Creating general learning algorithms is arguably a somewhat easier task than creating a general agent, since learning algorithms are typically run for a long time while an agent often has to make time constrained decisions.
The holy grail of AGI is to make the learning algorithm and the agent the same thing, and have them be general. Then you have an agent which can rapidly adapt to its environment and self-modify as needed. We are still a long way off a system that would do this in terms of current research.
Their “general learning” tech doesn’t even generalize to barely modified variants of the original games it has claimed to master. I call bullshit.
But the point I was making is precisely that the "general learning" tech is in fact somewhat general. AlphaGo and certainly AlphaZero's learning tech generalises to Go, chess, and a few other games. That's relatively general in the domain of board games, in my humble opinion.
The reason this isn't close to AGI is because it's not the agent doing the learning, and so while a relatively general learning algorithm produces the agent, the agent itself is not general even in the field of board games.
> AlphaGo can play very well on a 19x19 board but actually has to be retrained to play on a rectangular board.
It doesn’t even generalize to the same game with a different board shape. Whereas a human Go master could easily do so.
DeepMind is essentially hacking the common usage of the word “general” in order so that they can make claims about “general” intelligence. And it’s working!
How is that not general? Sure it doesn't work for all problems but in the domain of board games it definitely feels very general.
The agent the training algorithm produces may not be general, but out of what I've read I've only ever seen DeepMind claim generality of the learning algorithm, not the agent.
If the system designer has to know the parameters of the challenge the system is up again, it should be obvious you can always add another parameter that the designer didn't know about and get a situation where the system will fail. This is much more of a problem in "real world situations" which no designer can fully describe.
The human state of the art solution seems to be going on slack and asking questions about the data provenance, which will decidedly not work for an automated approach.
A primary reason I can do better a better job than a generic algorithm is because you told me where the data came from (or I designed the schema and ETL myself), while the algo can't make any useful assumptions because all that info is hidden.
I’m not aware of any work or even sci-fi that addresses AGI with regards to this question, and would be curious if there’s stuff out there?
Anyways, with some naive googling I found these references which seem interesting with regards to lineage and causality (for the query "lineage causal database schema"):
[0] Causality and Explanations in Databases. The "Related topics in Databases" section seems interesting.
[1] Duke's 'Understanding Data: Theory and Applications, Lecture 16: Causality in Databases' (by one of [0]'s authors).
[2] Quantifying Causal Effects on Query Answering in Databases. This has some interesting definitions.
[3] Causality in Databases. Seems like a more in depth version of [0].
[4] A whole course on "Provenance and lineage"
[5] Causality and the Semantics of Provenance. Defines "provenance graphs" and some properties.
-----------
[0] http://www.vldb.org/pvldb/vol7/p1715-meliou.pdf
[1] https://www2.cs.duke.edu/courses/fall15/compsci590.6/Lecture...
[2] https://www.usenix.org/sites/default/files/conference/protec...
[3] https://www.cs.cornell.edu/home/halpern/papers/DE_Bulletin20...
However, with a more general AI, we would be able to tell it "this is where and how the data were collected" and it could make the necessary inferences. Fully general AI would also be able to ask the right questions and make reasonable guesses on its own. Everything you do now, and nothing like anything that's been developed.
To me the weak point of this article isn't in the thesis (of course everyone can agree that general intelligence is more useful than narrow, all else being equal) but that there was nothing said about how to get there. The only reason ML is currently resurgent is because it we've figured out how to do something that works, while general intelligence has proven beyond our reach for 60+ years.
Edit: actually it seems to be "classical AI" and hybrid approaches and I guess for more details one would need to read the book.
I guess articles like this are worthwhile to temper expectations of what's possible with the current crop of technologies for those not in the field, which could help prevent another winter due to overinflated expectations.
Your point about AGI which needs to ask questions about data provenance is super interesting. Are you aware of the line of inquiry into active learning? It's fascinating and has a long history: https://papers.nips.cc/paper/1011-active-learning-with-stati...
https://en.wikipedia.org/wiki/Active_learning_(machine_learn...
Or if we're talking serious AGI level, just feed it your codebase/email/slack history and have it learn all of that hidden info.
I have earned over 90% of my income over the last five or six years as a deep learning practitioner. I am a fan of DL based on great results for perception tasks as well as solid NLP results like using BERT like models for things like anaphora resolution.
But, I am in agreement with Marcus and Davis that our long term research priorities are wrong.
I think it probably won't be one breakthrough, but several, over decades. Personally, I'm pretty happy that AGI is taking a long time to materialize. We likely won't see a "fast takeoff scenario" (the computer is learning at a geometric rate !!1). It will likely happen gradually over years (progressively more intelligent, more aware computer systems), and we may have a chance to adapt in response.
I think that’s fair, deep learning today has an issue with learning guide rails and obviously it is only as good as the data you feed it. I think it’s fair that our models need more
The cutting edge NLP stuff just showcased at my company was pretty lame, too. I barely saw any statistically significant results at all and yet they rather unscientifically proclaim success because they got any effect at all. Some of what we do in our field doesn't matter because it comes down to whether a customer got a 2nd call back and got converted to some minor sale or added to a program for them. It's throw away and creates good will at conferences and talks. We make a big deal out of it.
We are spending hundreds of millions on projects, trying to save money on generating leads, reducing interactions with customers and vendors through staffed phone banks, and so on. My company has hired all kinds of academics and research type people and has given them titles of "Distinguished this" and "Principal that" and honestly there's not that much to show for it, maybe zero direct outcomes so far. What galls me the most is in all the conferences and demos they are showing off things like High School robotics vehicles and AI parlor tricks and astonishingly little has translated into the business we do. Meanwhile, there are people in the company who do know how to reduce costs and get more done and have outstanding outcomes, but their techniques are not sexy and thus unimportant to the PT Barnum MBAs running our company. I'm sure that's true most everywhere, of course.
These Principals and Distinguisheds all keep proclaiming success while cashing fat paychecks. Meanwhile this year, our stock has had a tough go of it, so I'm curious whether these attempts will continue. The market takes no prisoners. Sure we get a lot of mileage out of looking cool for the recent grad crowd purposes of recruiting--kids want sexy, cool tech projects to work on and words like "insurance" turn them off, so there's that, I guess.
My take on that is that it won't be long before all those new recruits will figure out they got bait and switched pretty bad and that they aren't going to get to work on any of this sexy ML and AI stuff anymore than I am in my role. I got lured in by Data Science (because PhD), which just shows how gullible I am, but at least some of that traditional statistical modeling is having an impact here and there. The problem again is that even that is overblown by a couple of orders of magnitude! In my project, we're simply trying to get more real-time data out to people who need it without having to call in to get it and that is ridiculously difficult because of all the systems we try to knit together and how overall terrible our data quality is. And now my boss wants to build out an "analytics engine" to capture some of this sexy ML and AI stuff. It leads me to believe that the people involved are most interested in getting promoted and not much more.
Anyways, it is cool tech, but American taxpayers and people who are forced to buy our products are paying for it and I rather think they would prefer to spend their money in some better fashion.
That said, I also lived through and worked through the level of hype around expert systems. I think the high level of hype around expert systems in the 1980s was much more extreme and unwarranted that the DL hype levels. I base this on selling expert system tools for both Xerox Lisp Machines and for the Macintosh when it was released in 1984. Some of my customers did cool and useful things, but nothing earth shaking.
At least DL provides very strong engineering results for some types of problems.
Already computational resources are becoming prohibitive with only a few institutions producing state of the art models at high financial cost. If the goal is AGI this might get exponentially worse. Intelligence needs to take resource consumption into account. The models we produce aren't even close to high level reasoning and we're already consuming significantly more energy than humans or animals, something is wrong.
The scale argument isn't great either because deep learning is running into the inverse issue of classical AI. Now instead of having to program all logic explicitly we have to formulate every individual problem as training data. This doesn't scale either. If an AI gets attacked by a wild animal the solution can't be to first produce 10k pictures of mauled victims, intelligence includes to reason about things in the abscence of data. We can't have autonomous cars constantly running into things until we provide huge amounts of data for every problem, this does not scale either.
That is something of an illusion.
Obviously there will be some sort of uneven distribution of computing power; some institutions will have more, some less. The institutions with more power will create models at the limit of what they can do, because that is the best use of their power.
So if the thesis of more power = more results holds then truly cutting results will always be by people with resources that are practically unattainable by everyone else. Google's AlphaGo wasn't a particularly clever model, for example. It just had a lot of horsepower behind it to train it and the various ranging shot attempts Deepmind would have gone through. Someone else would have figured it out albeit more slowly in a few years as computing power became available.
Computational power is still getting exponentially more affordable [0]. Costs aren't really rising, so much as the people who have spent more money get a few years ahead of everyone else and can preview what is about to become cheap.
[0] https://aiimpacts.org/recent-trend-in-the-cost-of-computing/
Most of the article is describing past scenarios, only the last 3 paragraphs make the argument that the past is a good representation of the present
https://twitter.com/ylecun/status/1066568396177842176
i.e. gradient-based learning is the final word on the matter.
_
/ \
,----------------------------.
| Hey! I found your minimum! |
'------------ ------------' /
.----. \ / / \ /
/ \ \/ / \ /
/ \ @ / \ /
/ `'--' \ /
/ \ /
/ \ /
/ \ /
/ \ .-. /
/ \_/ `\ /
\.__.' \ /
\ /
`;._.-'The average human has extreme difficulty reasoning about politics, while usually being reasonable on medicine (anti-vax being one of many exceptions). And it seems strange to expect a skilled pianist to also be a skilled neuroscientist or a skilled construction worker. On the other hand these people all use similar neural architectures (brains). So he seems pretty off-track when he criticizes "narrow AI" in favor of "general AI", as if there's some magic AI that will do everything perfectly, and even more off track when he criticizes researchers for using "one-size-fits-all" technologies, when indeed that is exactly what humans have been doing for millennia for their cognitive needs.
And sure, ML models in publications so far are typically one-off things that react poorly to modified inputs or unexpected situations. But it's not clear this has any relevance to commercial use. Tesla is still selling self-driving cars despite the accidents.
Total straw man. He actually uses an intern as an example in the very next sentence after what you quoted, as you would expect them to be able to read and get up to speed on a new area regardless of what it was. Meanwhile SOTA in NLP is a system that can be built to answer a single kind of question but can't explain why it did so or do anything useful if given an explanation of why its answer was wrong.
But as I said, I don't see why an artist would suddenly get up to speed as a construction worker. He seems to overestimate the capacity of interns as well.
An artist understands the goals of construction work, and can pick up the skills necessary along the way, because we can understand a goal and have a wide variety of cognitive tools to let us know how we are doing. If you've worked closely with BERT you already know that interns have nothing to worry about, not just from the current crop of tools that includes BERT, but from the entire line of deep learning research, short of a sudden and dramatic shift in direction.
The opposite extreme is something like the Atari game system that DeepMind made, where it memorized what it needed to do as it saw pixels in particular places on the screen. If you get enough data, it can look like you’ve got understanding, but it’s actually a very shallow understanding. The proof is if you shift things by three pixels, it plays much more poorly. It breaks with the change. That’s the opposite of deep understanding.
Of course. There are an infinte way to make interpretations of perceptions and a finite subset of possible valid ones.
It's among those possible, that the AI will be a concrete implementation of an ideology.
To select which one is always done by humans.
There’s plenty of research on Bayesian neural networks for causal inference. But even more, a lot of causal inference problems are “small data” problems where choosing a strongly informative prior to pair with simple models is needed to prevent overfitting and poor generalization and to account for domain expertise.
Deep learning practitioners generally know plenty about this stuff and fully understand that deep neural networks are just one tool in the tool box, not applicable to all problems and certainly not approaching any kind of general AI solution that supersedes causal inference, feature engineering, etc.
This article is just a sensationalist hit job trying to capitalize on public anxieties about AI to raise the profile of this academic and try to sell more copies of his book.
I’d say, let’s not waste time on this crap. There are engineering problems that deep learning allows us to safely & reliably solve where other methods never could. We absolutely can trust these models for specific use cases. Let’s just get on with doing the work.
I agree with most of the article but I think this^^ skips over the different types of networks used to solve perception and language problems. A CNN is very different from say, word2vec, which isn't a very deep network at all.
It absolutely makes sense to use deep learning for both of these tasks.
In fact, one very effective thing to do is to use a Siamese network to learn joint representational spaces of text and imagery in the same network.
It’s really specious and disingenuous to say “boy, vision and language sure seem different but can you believe these DL researchers are using the same tools for both!?”
edit: after noticing this other hacker news article (https://news.ycombinator.com/item?id=21107706), I wanted to add that this line of thinking is applicable to understanding programs and proofs written by humans as well. Programs and proofs can be well-understood when their pieces, and the way those pieces compose, are well-understood. When the pieces, e.g. lemmata in a proof, are large or hard to decompose, the proof (i.e. the solution to a problem) is harder to verify and understand.
That said, in general I don’t expect that we could understand any particular solution produced by an AI, be it deep or otherwise, but I do expect it to be possible quite often.
The holy grail of neural nets has always been to build a simulation of the brain, figure out how it works, and apply that knowledge to how the human brain might work.
We're not there yet but progress has been made. Eventually we'll understand NNs well enough to explain not only themselves but also human brains. In any case we have no choice because we cannot deploy NNs in life critical situations until we understand how they work, because that's the only way to understand how they fail.
I'd say that's a goal for some people -- for those whose goal is to figure out how the brain works, rather than constructing a more ideal and powerful GI. Remember the brain is great at some things, but laughable at others -- such as a "7 +/- 2" items in short term memory, inability to immediately retain rote knowledge after one instance and in great numbers, etc. It's the merging of the fuzzy, goal-directed behavior of the mind, in conjunction with its ability to effect the "real world", and the super-human memory and computational capabilities of computers that makes possible future GAIs that are so powerful and possibly scary.
If symbolic AI is the right model, but difficult to build algorithmically. Vector based model just help to make it faster and better. Then we humans are fine. We simply proxy the lower level optimization to AI. Our functionalities will be shifted just like what happened when engine was invented hundreds of years ago.
It's certainly true that we think in symbols but they exist somewhere in the mushy goo of neurons, could symbolic thinking emerge from large ANNs in the same way?
I think ANNs implement only the word2vec function that translates images or sounds into symbols and vice versa.
People really interested in AGI should better look at Cyc and opencog
An SVM? A markov model? A large context free grammar with a dictionary?
Not saying its obviously possible but it doesn't seem obviously impossible and its a mistake to assume as such.
Sticky wicket.
Make a system that allows one to scan their face, and OPT-OUT OF ALL FACIAL RECOGNITION.