If you think there's something NLP can't do using machine learning, make a challenge dataset. That would be much more useful than yet another theoretical argument.
If you think there's something NLP can't do using machine learning, make a challenge dataset. That would be much more useful than yet another theoretical argument.
Recently I was impressed to see that the Winograd schema challenge -- a particular challenge dataset of the kind you mention -- saw huge performance gains from statistical NLP methods, which implies that the statistical methods encoded (something that at least sometimes works a lot like a substitute for) common sense about the world and the context of utterances.
The Winograd schemas were my favorite example of "XYZ" here. They involve pairs of sentences where the sentence is parsed differently based on assumptions from external-world knowledge. For example
Putting the barbell on the glass table isn't a good idea, because it's too (heavy/flimsy).
What is too heavy?What is too flimsy?
Or, from memory rather than finding the original version
The local authorities refused a permit to the demonstrators because they (feared/advocated) violence.
Who feared violence?Who advocated violence?
The compiler confirmed that that the program was buggy because it (enforced/failed) typechecking.
What enforced typechecking?What failed typechecking?
The doctor recommended surgery for the patient because she was (suffering from/experienced with) coronary artery blockage.
Who was suffering from coronary artery blockage?Who was experienced with coronary artery blockage?
In these cases, the most plausible referent of a pronoun changes depending on the semantics of the other word that you fill in, for a reason that has to do with the outside world, not syntax. (I don't each case absolutely has to include a "because" clause, but it's the simplest way to supply disambiguating context.)
I totally thought statistical techniques would be horrible at this. There's good reason to think that these require some detailed knowledge about the world. The statistical techniques were horrible at this, for a while, and now they're abruptly great at it (if I remember correctly, at or above typical human performance).
The reason we have not figured out NLP is not because we are incapable, it’s because enough of the right minds have not been looking at it as a puzzle worth solving, possibly because it and other AI related concepts are often introduced at or just before the PhD level and so most minds in science never encounter it (even though they could never encounter any other written concept without it)
Once this is clear, the discussion becomes healthier: linguists continue to try and understand fundamentals; ML NLP continues to try to get impressive, possibly lucrative results. Logical research can borrow from ML results, or re-use them, and so forth, but I agree with Chomsky that it can't be a shot in the dark always, because once you deploy it and pass it some tough decisions to make, well, what then?
I was pleased to see at least one idea, the recursive structures found in his theories, which is increasingly being accepted by people like Hinton, so that's great.
You can't understand why without considering the environment, the problem is that it's too expensive to train AI agents in reality and simulated environments are too simplistic.
But if such a simulated environment is available, agents can learn general skills.
> Deep mind: Generally capable agents emerge from open-ended play
https://deepmind.com/blog/article/generally-capable-agents-e...
You can't understand why without considering the environment, the problem is that it's too expensive to train AI agents in reality and simulated environments are too simplistic.
But if such a simulated environment is available, agents can learn general skills.
> Deep mind: Generally capable agents emerge from open-ended play
https://deepmind.com/blog/article/generally-capable-agents-e...
In this kind of setup the agents know why they act like they do.
Well, there was a good fifty plus years of linguists without access to good statistical methods who also failed to solve the problems as well... And they also didn't come up with the kinda functional answers that we now have from statistical methods, which are now able to do crazy things like translate reasonably well between nearly arbitrary language pairs.
Your statement sounds like a theoretician's sour grapes...
We do understand how and why it works to a certain degree (gradient based input to output approximation). But if you want the meaning of the third neuron in the 10th layer, then yes... But in the same sense we don't understand thermodynamics because we don't account for each particle, just a statistical aggregate of them.
And on the other hand, we don't understand much about how people work or how we are motivated, even though we have first and third person perspectives. Psychology seems to be no further advanced than AI. And yet we work with what we have.
Lots of brilliant people have looked, but the issue at the moment is that the last set of questions and challenges was surprisingly blown away by statistical approaches and we are waiting to find the ceiling on those approaches before being able to formulate the right questions for the next round of deep thinking I suppose. The Winnograd schema were kinda meant to prompt that....
The issue is that obtaining this understanding requires a lot of effort.
OpenAI invested in trying to understand how a simple CNN works.
You can read their results here: https://distill.pub/2020/circuits/
> In the original narrative of deep learning, each neuron builds progressively more abstract, meaningful features by composing features in the preceding layer. In recent years, there’s been some skepticism of this view, but what happens if you take it really seriously?
> InceptionV1 is a classic vision model with around 10,000 unique neurons — a large number, but still on a scale that a group effort could attack. What if you simply go through the model, neuron by neuron, trying to understand each one and the connections between them? The circuits collaboration aims to find out.
Either the thing you want to do is impossible or else a learned model can do it at least almost as well as... idk what the alternative even is, something not learned?
Taking this to the extreme, on an example where I have first-hand experience: I would never recommend replacing your compiler passes with transformers. The latter will be buggier, at best marginally faster, and it will take several orders of magnitude more effort to get them to work well enough for production. ML isn't the right tool. But, I mean, you can do it. You shouldn't. But you can.
To be fair to the article, the title is "won't", not "can't". And that's easier to believe, at least for me. It's not that X is unattainable using ML; it's that some other approach will get there first.
It is my understanding that a lot of current applications take a ton of training time / computing power and it's not so trivial to do in the moment. I guess this is less of a theoretical problem and more of a practical problem though.
Yeah, exactly. Theoretically, ML can do anything. Practically, not so much.
However, you can use ML in areas where compilers use heuristics. For example, how many times should you unroll this loop, or should you inline function B into function A? Your program is going to be valid with or without inlining/unrolling, and the compiler needs an "intuition" of what's going to be best for performance. Right now, in most cases, this uses simple hardcoded rules.
The same goes for math. You wouldn't ask a neural network to certify whether a mathematical proof is valid or not, knowing it has been 99% accurate on the test set. You couldn't trust it to verify all the steps. However, you could use a neural network to help you guide a formal theorem prover in searching for a proof, as in which branches of the infinite tree of mathematical expressions should you search.
And then, if ML is as good as we are, and clearly we're not good enough that we had to write a systematic proof system to compile things correctly, could ML have invented this same technique and implemented that compiler correctly?
In the end, we're just back at the same place, can ML be as intelligent as we are. Is it the holy grail of AI? What are its limits?
So a theoretical argument could useful if gave us an idea what are our mysterious something it. However, I'm not sure if this article makes a contribution here. Proofs of impossibility generally aren't useful, as you say.
That can be true without being demonstration that we're seeing unbounded progress in language use and understanding. A theory of the limit of NLP is still a theory of NLP and by no means better than other theories.
Umm, you just said it yourself, understanding. The model has no clue whatsoever what any of the output means, just that it scores good.
I don't understand Japanese. But given enough time and feedback I can try all the possible combinations of sounds I can come up with and answer a question in Japanese till I get a satisfying feedback from the Japanese speaker (ex: his facial expression) that asked the question. Is this satisfactory as for speaking Japanese in your opinion?
It's always seemed obvious to me (as an outsider) that it's missing reason and explainability. GPT-3 is a neat tool, but it seems like anthropomorphism to suggest that it's more than the best mimic humanity has been able to create so far.
I haven't. The Turing test has never even been close to being passed at the annual Loebner Prize. The best such industry NLP products I know of don't even use ML at their cores.
I think this is a valuable distinction. Unfortunately, though, it is one that the research landscape seems to have somehow forgotten about over the last decade or so.
More tangibly speaking using an example, while the capabilities of a language model like GPT-3 are amazing, Norvig's point is that science should not just stop there - it should ask the question "what does GPT-3 teach us about human language?"
Chomsky nowhere does this. That is at best a straw man constructed by Norvig of Chomsky's position. If you just look at the transcript of the Chomksy interview in the Norvig piece, you can see Chomsky describing a case where statistical analysis is useful.
I think I'm not gonna do it. [good in English]
I think not gonna do it. [bad in English]
(Yo) creo que no voy a hacerlo. [good in Spanish]
To linguists it's immensely frustrating that someone who's considered by the compsci crowd to be a leader in the field of NLP can't even get some basic facts of linguistic analysis correct. "Linguists can argue over the interpretation of these facts for hours on end", Norvig says dismissively. But any competent syntactician could have explained to him in five minutes why his examples are irrelevant to the point he's trying to make.
The rest of Norvig's post just consists in misunderstanding what Chomsky is saying, as far as I can see. Chomsky is saying that statistical analysis alone is not sufficient to achieve scientific progress. I don't think Norvig actually disagrees with Chomsky on this point, but he seems to think that Chomsky is saying that scientists should never use statistical methods.
The real kicker here is that Norvig is an engineer, not a scientist. He's contributed very little to the scientific study of language, and yet is lecturing someone who's contributed vastly more on how it ought to be done. Of course, that doesn't necessarily mean that he's wrong, but it does grate. Not to mention that the Chomsy/O'Reilly comparison is little more than trolling – it hardly seems calculated to stimulate an intelligent response.
I can write code to bubble up stats across text sources in any language.
I’m reminded of a behind the scenes video of Batman 1989, where they discuss a complicated layering of Joker makeup to achieve the effect in the scene where he wipes off fleshy colored face paint; white makeup on Nicholson, special coating to be able to apply the next layer, on and on.
No one thought to just have him wipe white face paint onto his forehead.
I can easily see a bunch of programmers making a mess out of an elegant problem given the spaghetti code I’ve worked in.
Maybe English means nothing about consciousness? Many linguists take the position it’s random sounds we’ve been polishing definitions of for years.
I'm not sure I understand your framing here, but in it, aren't you just describing what Google _is_?
Though, I’m confident that it should be possible to make something which, when unsure, asks for clarification.
In fact, I believe something at least rather similar is already done in, uh, I think it is called “delegative reinforcement learning”, where an agent takes actions, but when it is highly uncertain, it can instead choose to have an expert (who is assumed to be competent at the task, or at least has a tolerably low risk of very bad outcomes) to step in instead.
(Or, wait, maybe delegative reinforcement learning is still just used theoretically? Not sure.)