NLP's Clever Hans Moment Has Arrived
thegradient.pub
thegradient.pub
> AI trained to classify skin lesions as potentially cancerous learns that lesions photographed next to a ruler are more likely to be malignant.
> Agent pauses the game indefinitely to avoid losing
> A robotic arm trained to slide a block to a target position on a table achieves the goal by moving the table itself.
> Evolved player makes invalid moves far away in the board, causing opponent players to run out of memory and crash
> Genetic algorithm for image classification evolves timing attack to infer image labels based on hard drive storage location
> Deep learning model to detect pneumonia in chest x-rays works out which x-ray machine was used to take the picture; that, in turn, is predictive of whether the image contains signs of pneumonia, because certain x-ray machines (and hospital sites) are used for sicker patients.
> Creatures bred for speed grow really tall and generate high velocities by falling over
> Neural nets evolved to classify edible and poisonous mushrooms took advantage of the data being presented in alternating order, and didn't actually learn any features of the input images
Aka "heads, we win; tails, I don't lose"
Quite interesting
Well, this sounds like speedrunning. People have found arbitrary code execution vulnerabilities in SNES and used them to jump to the credits (which counts as completed the game) in less than a minute: https://www.youtube.com/watch?v=Jf9i7MjViCE
They also used the same technique to inject runnable code in the game: https://www.youtube.com/watch?v=jnZ2NNYySuE
In this case it’s just choosing an option that involves very large numbers because it’s learned that it’s opponent can’t handle large numbers. There’s no code injection
My point here is that there is similarity between (some) human players and some AI players. Even the discussion whether exploiting a glitch is actually 'winning' also looks very similar.
[edit] Just found out there's a second part: https://soundcloud.com/user-430897588-587876313/lstm-song-2
The learning agent discovers something that is true of the training set, but does not generalize to other examples of the problem outside of the training set.
The article talks about much of the problem being in data set construction, because it is very tricky to design data sets without accidental biases that the learner can use to correctly categorize the examples that have nothing to do with the actual problem you want the learner to solve. The traditional techniques to avoid over fitting, like holding out part of the training data, don't do any good if the entire data set is not representative of the real world in some systematic way.
Yeah. This is part of why DistilBERT (and the fact that you can do pretty well without BERT) is interesting, to me. It seems like for a very long time there have been people complaining about certain individuals and orgs throwing cash/compute at problems to look good rather than solve anything. The difference is nowadays it's starting to be less of a fringe view, thank heavens.
The most interesting thing about NLP (as someone who works in it), is precisely that it is very, very hard to get anywhere. And that in turn is why the field keeps turning up so many new NN designs: the flipside the author rightly points out is that this has to happen for data as well, if we aren't to fool ourselves about our progress.
Great read.
There's a reason we haven't seen a hard takeoff with BERT writing BERT+1 in even better Python and it probably isn't because nobody is trying, it's because BERT is too stupid to write a computer program.
Well put. The timing is good too, because NLU-is-just-around-the-corner hype is starting to have some really negative social consequences. Yesterday a disturbing article about automated test scoring in the states was trending on HN.
The human mind has many faculties, and the brain has correspondingly many functionally different elements. Plus years of brute force learning.
Compared to GPT2, AlphaZero and whatever we are waay more complex and had a lot more training done.
Probably the super secret sauce is in the hyperparameters that determine how to cobble all those functional components together to get something more than the simple sum of them. (But they are also learned and brute forced through evolution.)
AI needs better brains. (So R&D can proceed faster. We can test new ideas cheaper, we can optimize algorithms, so they'll learn and perform better.)
Human knowledge was created by humans. We might repeat it blindly too, but somebody had to create it first. That is fundamentally different from what Hans did.
e.g. highschool mathematics
But some people - communities of people - pored over this stuff for centuries to work it out and try to find flaws (and found them). You could have an AI that tries to find flaws - provided you have an actually correct statement of the problem.
That's the classic scifi problem with robots: they do what you tell them to.
It comes across as a warning that we don't really have any solid understanding of what we're doing when we create new and powerful entities, and that pride cometh before the fall.
Plus, any sufficiently advanced magic is indistinguishable from technology :P
> Look at the history - it's actually hard to find something that is not biased - false, biased believes are spread throughout the timeline of human kind. Given enough time, our knowledge gets closer towards truth
How do you think you know any of that?
Parents, family, teachers, community... ???
However, even on an individual level, haven't you ever had a new idea that didn't come from another person but just your own observation and contemplation?
It would be a copout. Instead of actually tackling AI's problem of common sense, claim that maybe layers of logistic regression and matrix factorization is all there is, we are its equal, just a few layers up in abstraction and evolution. Does one really stem and count tokens to decide if a movie review is negative sentiment? Or does one empathize with its writer and build a complete model inside your head?
The horse would be the AI researcher claiming reasoning and understanding from an activation vector trained on word co-occurence on Wikipedia, and the farmer giving clues, is the heated community and industry, mistaking impressive dataset performance for a solution for a problem they're starting to forget.
first problem: find the roots x^2 + 2x + 1
second try: find the roots of 4x^2 - 2x + 4
Or something like that.
> Our main finding is that these results are not meaningful and should be discarded.
As often happens, ML found a way to exploit a trivial bias in the dataset. Tip of my hat to the researchers for actually doing a good job! Also really enjoyed this read.
Example: For years, they have strived to make a ....... in the community.
Only a few possibilities out of 30,000-100,000 English words in use can fit properly in the blank.
Thus, many tasks, even those designed to test for real understanding, can be “solved” pretty well by putting together a number of relatively shallow cues. BERT and similar models learn from huge datasets (billions of bytes) and they probably capture millions of those correlations in the model. (These models are indeed wonderful accomplishments and are very useful for many things, but not ultimate solutions to true natural language understanding.)
For more info, see: https://super.gluebenchmark.com/
I agree with the author that we need even better datasets evolved under public scrutiny since some of the datasets we currently have are already designed very well by their authors but the problem of designing datasets that can withstand correlation detections by DNNs (and still amenable to standardized evaluation method) could be too challenging for any single team under limited time.
For years, they have strived to make a pizza in the community. (But since they lack flour for the crust, it has not ended well). For me, I can fit nearly any noun into that sentence and imagine a viable scenario where it would be reasonable. I honestly don't know what you were getting at. For others, I am sure that it is obvious and they can tell you the few words you were imagining.
My English language ability is quite good, so I wonder why I can't perform at these kinds of tests. I also wonder if knowing why I can't do this is useful for NLP.
If you are allowed two words, "real difference".
The way to find it is to find candidate words and try them one by one until you hear something that sounds like you've heard it a hundred times before.
> I also wonder if knowing why I can't do this is useful for NLP.
Yes. I wonder if the inability is acquired or learned or innate? Could you learn to do it? Do you have an aversion to catch phrases and well-used (hackneyed) clichés? Do you prefer to weave your own words into sentences?
I notice you use some odd prepositions in odd orders. For example in your second sentence most people would say "I have real difficulty with fill-in-the-blank tests."
You also wrote "perform at these kinds of tests," which is a place where almost all American native English speakers would use "on" rather than "at".
If I had to guess, you're much less sensitive than the average person to slight variations in word pair frequencies, and you could certainly create a test to test this hypothesis. For example you could choose any n-gram likelihood data derived from well-written English texts, and write a program to measure your ability to distinguish high- from low-frequency word pairs. This should be lower than other people in your peer group.
BERT and similar models do consider a wider context window (up to 1000 lemmas if I recall correctly), which is a reason they do quite well when trying to fill in blanks for a complete passage and usually less accurately for shorter texts.
If one thinks about loose associations that newcomers make compared to experts in a domain, this seems very similar to me.
When designing a vector graphic one can enable "snap to grid", similarily at some point we will have to "snap to makes sense" by means of verifiers or provers. "Why" type questions ultimately ask for a proof or derivation, which in the past (before the advent of logic, to which a philosopher would have to carefully abide) was not objectively verifiable, but at some point it is foreseeable that neural networks will be asked to construct (or append) a formal belief system, predict a conclusion, and justify it by supplying a proof.
Yes, I think this is how the human mind works, kind of, but I don't think the part that provides the grid is like the verifiers or provers we have implemented. There is something that provides a substructure, and I think that you can see in mental illness where it's not functioning properly, but even when healthy, it's not that logical. When humans do logical reasoning, that's a very high level activity, superimposed on top of the other layers, IMO.
I think that the "snap to" part is going to be very difficult to develop because we are not conscious of it. Where to even begin? It might be fruitful to study instances where it isn't working properly - like thought disorders in schizophrenia.
So yeah, I think human thinking may have severe flaws that are rather similar to the ones discussed in the article, but that doesn't prove that current AI has all the components needed to match humans.
I adamantly agree, there is probably a lot of generally applicable (i.e. not domain-specific) "tricks" or "implicit insights" that mammal brains use which we haven't discovered yet.
Another good example is, how it's often stated that humans have no trouble learning things "in one shot", but how on closer inspection that may or may not be true and is very hard to verify:
Consider how when we have a clear negative or positive experience (say being thrown out of class or standing in a corner facing the wall versus getting a compliment on your work etc), we afterwards typically replay the sequence of events and how it led to the current situation. It is unclear if this replaying occurs only for the highest-level abstract thoughts possibly centralized into a couple of regions recording and replaying episodic memory, or if this is actually happening in a more decentralized way throughout the brain while we are only subjectively aware of this episodic aspect at the conscious level. If this self replaying of the last lessons that apply locally is spread and operating independently throughout the brain, could this be an explanation of the brain wave patterns? Is this origin of the "aha" signals or is that just a fantastic concoction of the reproducibility crisis?
Local feedback signals (numbering on the order of number of synapses) vs global feedback signals (numbering on the order of kinds of hormones, diffusing neurotransmitters, blood sugar etc) ===
We still know very little of the feedback mechanisms in the brain, is it a low number of global feedback signals, or a large number of local feedback signals?
A) global feedback: the hardcoded feedback mechanism is primitive wetware (listening to the low number of chemical signals) while the high-speed feedback is learned by neurons influencing each other in the prograde direction as emergent behaviour resulting from this primitive wetware (the only feedback being through axons literally feeding back as opposed to feedforward networks); or
B) local feedback: if the high-speed feedback is hardcoded wetware too, i.e. retrograde signalling of adjoint derivatives across the synapse, which would mean that the reverse accumulation automatic differentiation we use is closer to biology than currently accepted.
For example in A) suppose there is a low bandwidth (low frequency) feedback reward signal, say blood sugar rising after eating a given sweet for the first few times, then the shortest path between the taste buds and the chewing and swallowing motor neurons would improve their weights for intake, a first level anticipation, but then later seeing or touching the sweet or actively seizing one might cause the previously trained neurons to train the newly recruited signal paths by some other neurotransmitter reward by anticipation.
Fast-approach-but-imprecise adaptation vs Slow-approach-but-precise fine-tuning ===
Another facet is that we currently have "1 training regime" for digital neural networks, with which I mean: we use gradient descent both to 1) adapt the weights from a totally randomized initial state to a somewhat acceptable state, and then after a nonexistent pause, 2) to continue improving the weights near the top of the hill (or bottom of valley...) of the scoring landscape. We have no guarantee that nature uses a single regime, let me concoct a hypothetical (and thus improbable) guess: perhaps it uses gradient descent to get the weights approximately where they should be, but then a per-synapse lock-in amplifier correlates (multiplies over time 2 signals while summing) the positive / negative feedback with its own positive or negative variation on the weight, smoothes this product-sum correlation signal (low pass filtering) to get a very high precision feedback signal which has higher precision than a single instantaneous "sample" of feedback (local or global). Throughout the brain synapes some would be closer to the instantaneous feedback signal regime (in response to new lessons. or for short term memory), while others would be closer to the LIA regime, for longer term memory and / or precision. That would be trillions of lock-in amplifiers in a single human brain...
EDIT: in fact SGD stochastic gradient descent can already be seen as a global LIA mechanism, but with all weights / synapses using the same low pass filter. In theory we could give each synapse / weight not only a weight but also each their own timescale (or example count scale, the tau in (alpha) and (1-alpha) multiplier factors when expressed as a filter), and adapt the timescale depending on the local feedback signal (adjoint derivative) in the backpropagation algorithm. To vary the weight just use weight times (1 + 0.001 times the randombit bit r), and multiply it with the feedback signal (the component of the usual gradient, i.e. the derivative that corresponds to this synapse / weight). A similar trick could be used to vary the timescale. One could also hardcode the timescale for certain synapses as a hyperparameter to forcibly locate short term vs long term memory pathways.
EDIT2: As an example of the function of weights that would settle (or were forced) at long vs short timescales at low or high levels:
low level, short term: adapting to reverb entering or leaving a church, adapting to noise when entering or leaving a party,
low level, long term: associating spectral peaks among frequency bins (the fundamental and higher harmonics for sounds of different timbres would remain in roughly the same place, regardless of background noise)
high level, short term: unimportant small talk, or important but quickly dealt with information (did I already pay this drink or can I walk away?)
high level, long term: what is my pin code?
Speak for yourself! There are plenty of other ways. Gradient descent is just currently in fashion.
I think we're more likely to create AGI than to understand how the brain works in our lifetimes.
Certainly not. But it is similar in that confidence increases once one decides to believe in something.
> I think that the "snap to" part is going to be very difficult to develop because we are not conscious of it. Where to even begin?
Begin by becoming conscious of it. Learn to recognize when you have made a decision, or a judgment, learn to feel what that feels like when it is just about to happen, and then pay attention to what leads up to it. Is it some kind of subconscious ratiocination? Then slow it down and figure it out consciously. Is it a subconscious review of some sense data? Then review it with higher awareness. Intuition and introspection will teach more than trying to watch it break.
Especially in how people think machines work: the user facing interface is optimized for these naive interpretations, while the true operation is hidden because of complexity that "would confound us".
> [Good test takers] are prone to adopting shallow heuristics that succeed for the majority of [tests], instead of learning the underlying [facts] that they are intended to [assess].
Semi-related, I wonder where one draws the line on which heuristics are deep vs shallow in test taking. If a question assesses vocabulary knowledge by asking you to pick an appropriate word from several choices, "these are words I've seen before even though I don't know what they mean" && "here's a choice whose Latin roots plausibly give it the right meaning" seems shallow to me. It allows the test taker to select the right answer without actually knowing it. At the same time, it's not shallow in that you're incorporating quite a bit of knowledge of the language as well as the context you've developed over years of reading.
---
As soon as the conversation here turned to test taking, the thread immediately reminded me of the scene in The Wire where the students pick the right answer from the blackboard by seeing one of the choices has a lot more stray chalk marks beside it, indicating it was pointed at during the previous period.
"The answer is B five. B five's got all the dinks"
The point is, good test design is actually hard.
Paul Nation's "Vocabulary Size Test" has a good approach to the challenge of determining how many words someone knows. Even with this relatively simple-to-state metric, there's a good bit of subtlety in getting a defensible result.
https://www.victoria.ac.nz/lals/about/staff/publications/pau...
I love this idea, thank you for sharing it.
EDIT: if you are into ML and like this idea, certainly check out MetaMath, the book is very accessible, and the author (Norman Megill) is extremely friendly and helpful. It takes perhaps a few days to a week to learn, study, and reimplement the verifier. One does not at all need to actually study the whole of the set.mm database (i.e. all the theorems and proofs) to implement the snap to makes sense I propose above. It simultaneously gets rid of the expensive labeling for correct vs incorrect proofs on one hand, and is a path to self-explaining AI on the other. I can not at all exclude that "soon" AI will be proving math conjectures faster than humanity can conjure them. That would be quite an interesting world! At first it would be trained on generating proofs for propositional logic statements (the challenges could be generated by a prover doing a "random walk"), then first-order logic / set theory / numbers.
Word2vec results in similarity scores you can use to find analogies. Great! Now how do you know which analogies to look for? Which ones to remember? Which ones to pursue? You need some kind of intuition for where to look. So you need (1) a way to move around the search space and (2) a way to know when to backtrack, or when a line of inquiry is unfruitful, and this last is precisely the capacity for boredom. If you can't come up with new ideas and you can't get bored, you'll suck at finding proofs just like you'll suck at math or problem solving in general. There's also a danger in getting bored by the wrong things. So how do you develop this boredom intuition?
We choose what to think about, but until we make conscious machines, we are probably the only creatures that have to make this choice.
extracting the analogies is not that hard (a naive brute force is looping over combinations of 4 words and testing how close the relationship holds), but more importantly, one doesn't need to extract the analogies, the neural networks utilize these analogies implicit in their embedding.
I have never seen the boredom of neural networks investigated, the closest concept that comes to my mind is surprisal, which is widely understood since Shannon & information theory...
Neural networks can't get bored because they don't decide what to think about. They are like a total functional programming language with no control flow constructs. This is why they always give an answer in the same amount of time, regardless of the input vector.
Surprisal is only a measure of information (given a distribution) and is only distantly related to what I'm talking about. However, if a neural network could choose to think more, choose to get new data and reconsider, choose to go back and look at that one from a couple minutes ago... then you'd also want it to have the ability to get bored. AKA to recognize when it is spending resources on an unprofitable line of inquiry.
The linked article suggests this does not hold (for the ACT), but doe suggest picking a single letter and sticking with it for all blind guesses would outperform a purely random guessing strategy.
To see that better, consider what I call "simplified Chinese room". It's a variation on a traditional Chinese room, where inside the room, there is only a pattern recognizer, which basically will match the input to arbitrarily many inputs (but not all possible) it learned before and chooses the output for the best match.
Now imagine I want to train this "simplified Chinese room" on solving satisfiability problem. Because in that problem, an arbitrarily small change in the input (introducing a contradictory clause) can completely change the output. It is therefore impossible, I believe, to learn the concept of satisfiability by just using pattern recognition (storing and comparing, according to some metric, previously seen inputs and corresponding correct outputs). Instead, you need to build a mental model which is internally self-consistent with these example pairs.
But it's true we don't precisely know how to "find and represent the better theory", if it's at all possible (most likely not) and what the exact trade-off is.
Also note that my argument doesn't necessarily rely on computational difficulty of SAT - even if we limit to some "simple" subset of SAT instances, we might see that the pattern recognizer is unable to learn these problems correctly.
The author is right that what we do for NLP is not enough for reasoning. But it does not mean that we are too far from it.
Even in your variation of the room experiment, if the person in it could reply something along the lines of "here is what I think is a part of solution", which you simply had to feed back to eventually get the whole thing, would you claim it is a weak abstraction?
<rant>That's where graph neural nets come into place. They can learn relations, scaling to a large number of objects. A traditional approach would have to learn all possible combinations, hitting the combinatorial explosion. Graph neural nets can solve problems such as shortest path, sorting and dynamic programming. I think in the future if we are to get closer to human level we need graphs as the intermediate representation. Graphs could represent the objects in an image/phrase and their relations, then answer about the attributes of an object, the relation between two objects or classify the graph itself. All simulators are evolving graphs as well, and code/automata could be represented and executed as a graph. The transformer could be considered an implicit graph where the adjacency matrix is computed from the nodes at each iteration. The closest to AGI in my view would be model based RL implemented with graphs.</>
But what bothers me is that the human animal is trying to create an artificial human mind, the most amazing piece of meat we have laying around.
If we could create the reasoning of a dog first, to then upscale to more complex logic, I'd be more confident research is going places (as in, we have biological evidence building blocks to understand first).
Yep, probably best to avoid. I still don’t think I’ve seen a convincing rebuttal.