Language models still struggle with the concept of negation
quantamagazine.org
quantamagazine.org
ChatGPT (GPT-4):
Sure, let's go through this step by step:
First, you asked for animals that don't have paws. This would exclude mammals such as dogs, cats, bears, and so on, as these species have paws.
Second, you specified animals that don't lay eggs. This eliminates a wide range of animals including all birds, most reptiles, and some fish and insects, as they all lay eggs.
Finally, you asked for animals that have wings. Wings are typically found in birds, bats, and insects.
Given these three criteria, we can see that we're in a bit of a conundrum. Most winged creatures either have paws (like bats) or lay eggs (like birds and insects).
However, if we consider the term "wings" more broadly, there are some species of bats that give live birth (don't lay eggs) and technically don't have "paws" but rather winged limbs adapted for flight.
So, based on the constraints of your question, the best fit would be certain species of bats, provided we interpret "wings" in a broad sense and "paws" in a strict sense.
What is meant by that saying is, “well the first time we ran the machine it worked just fine, so we got all excited, but then it consistently failed after that”, with the lesson being, increase your sample size before you pass judgement.
LLMs almost suffer from the opposite. Because they are correct so often the danger is in assuming the response is always correct.
Maybe, “even a Swiss made watch needs to be checked against the town bells”?
1. "Most winged creatures either have paws (like bats) or lay eggs (like birds and insects)" - but bats don't have paws.
2. "However, if we consider the term "wings" more broadly" - why do we need to consider the term "wings" more broadly? Nobody is arguing there is some definition of "wing" that would exclude bats. Bat wings evolved from largely the same structures that bird wings did.
3. "there are some species of bats that give live birth" - all species of bats give live birth - stating it like this makes it sound like there are some species of bats that lay eggs or something.
4. "...provided we interpret "wings" in a broad sense and "paws" in a strict sense." Again, no "broad" or "strict" senses are needed. Bats have wings and no paws.
To be clear, I still think ChatGPT is amazing, but these kinds of things highlight some of the differences between LLMs and human cognition.
Can hind limbs of bats be classified as paws? English is not my native language, so I'm just being curious.
First, you'll notice the added "work it out step by step" to encourage what is called chain-of-thought in the literature. Second, generate say, five responses from the LLM, then take all five responses and uses them in another prompt where you ask the LLM to reflect on all of the five responses, a process being called reflection. Then you take all of that, including the original prompts, and in another separate prompt ask it to figure out the best response, a process known as resolving. You could end it with another prompt to remove the step-by-step work and just return a succinct answer.
Given that most people have computers and RAM is dirt cheap right now, I'm already maxing out all of my devices.
But it does seem like a single call to an LLM is not the correct approach.
I’ve gotten the best results from attempting to mirror how I would operate in response to a prompt. Much of this is colored by decades of philosophical readings, which in my case means a decent amount of analytical philosophy along with latter Wittgenstein and other post linguistic turn thinkers.
So when I break things down, either in some kind of approach inspired by Frege, Wittgenstein or unadorned self-reflection, I can see that the unconscious mind seems to be doing something recursive before my conscious mind is ready to start fumbling over the word.
It seems like an LLM is just part of a larger cognitive architecture that will use these metacognitive approaches.
It’s also very obvious that mammalian brains are much more efficient than GPUs.
Bat paw: https://upload.wikimedia.org/wikipedia/commons/e/e1/Bat_in_a...
ChatGPT was totally in tune with this concept. The appendage on the elbow of the wing can be a paw or not a paw but it certainly evolved from paws. The reasoning was spot on and amazingly nuanced.
Doing some digging, paws are "soft foot-like parts of a mammal, generally a quadruped, that has claws" (Wikipedia). Bats do appear to have claws, but not really with the same structure as a paw; they don't seem to be mounted above a soft, fleshy pad. Rather, they're at the top of the wing, sitting on bone. They're also very different in function, when I've had the fortune to see a bat they seemed to use the claw to climb in a similar way to how climbers use ice axes.
My impression is that bat's claws are the anatomical equivalent of a thumb, not a hand/paw, and that the entire paw or hand of bat's ancestors evolved into their wings (sort of like if your fingers became very long and webbed and then you were able to fly with them).
So, my conclusion is that no, the claws of a bat are not a paw.
ETA: I neglected to consider their hind claws, but I don't think those are padded either.
https://i.pinimg.com/originals/1e/8d/09/1e8d0923d5caf2dbab68...
That certainly has claws but I wouldn't describe that as being soft or as being similar to a paw. And they're pretty clearly made for hanging, climbing, and grasping prey/food, and not for walking.
As long as it has claws according to vocabulary dot com it is a paw.
Anyway I'm enjoying imagining bats with little paws, it's an amusing image.
https://www.merriam-webster.com/dictionary/paw
Look up the definition of a paw. I think you have a sort of mixed up definition of what it is. Possibly a different definition used among your "bat expert" friends.
And I quote (again):
Paw: the foot of a quadruped (such as a lion or dog) that has claws
broadly : the foot of an animal
Both the broad and technical definition of paw fits what the bat has. Bats have paws.The thing you are referring to is likely are these soft pads that are often on the paws, sometimes called paw pads.
But this nuance in the definition of the word paw was addressed by chatGPT.
ChatGPT is being more logical to be honest.
It’s true, some species of bats give live birth.
To assume all bats give live birth might be committing the fallacy of induction.
Chiroptera fits under Mammalia in the phylogenenetric tree.
Bats give live birth.
I know what both 3.5-turbo and 4 say on my account, but I think I might be in some cost-effective inference holdout group because it’s always so dumb.
See how ridiculous that sounds?
CoT or <work> tokens are not 'cheating' - they allow the system to have (as an analogy) a sheet of paper to work out the problem, instead of relying on a single step to jump through thousands of steps of reasoning. It's trained on the algorithmic reasoning embedded within natural language (Transformers<=>GATs+PE⊂GNNs, which are dynamic programmers, and align well with algorithmic reasoning), so the steps are necessary to lay out in as much detail.
What does that possibly have to do with GPT-4 shitting the bed if you draw the logic puzzle out of a deck of cards instead of cherry-picking and hinting it?
I cheated, and used bc, the unix utility.
These three maneuvers (the correct answer, the agency shown by cheating, and the crystallized intelligence of unix) handily demonstrate both my own intelligence, and that of my species.
The LLMs aren't in the same category of being, yet. It's not clear they ever will be; while I think silicon is a viable substrate for intelligent agents, it still strikes me as doubtful they'll be any more descended from the current machine learning systems as they are from Symbolics Lisp.
It seems to me that the ability to reason ( though primitive it may still be) is an emergent property of sufficiently large differentiable hash/lookup tables (and if you think about the mechanism of reasoning, that seems... Reasonable).
According to "folks in the know", the scaling limits of GPT architecture seem to be nigh, but if it were a bit higher I wouldn't be surprised if LLMs would have made it.
You might get more use out of these tools if you learn how to use them and these natural language instructions are how you use these tools.
It’s like calling the reverse gear on your car a cheat code!
What animals don’t have paws or lay eggs, but have wings? Work it out, step by step.
Here are the steps on how to find animals that don't have paws or lay eggs, but have wings:
1.Define the terms. What do we mean by "paws"? What do we mean by "eggs"? What do we mean by "wings"?
2. Identify animals that meet the criteria. There are a few animals that meet the criteria. One example is a bat. Bats have wings, but they don't have paws. They also don't lay eggs. Another example is a butterfly. Butterflies have wings, but they don't have paws. They also don't lay eggs.
3. Confirm that the animals meet the criteria. We can confirm that the animals meet the criteria by looking at their anatomy. Bats have wings that are made of skin and bone. They also have claws on their feet, but these are not paws. Butterflies have wings that are made of scales. They also have legs, but these are not paws.
4. Draw conclusions. Based on our findings, we can conclude that there are a few animals that meet the criteria of having wings, but not paws or eggs. These animals include bats and butterflies.
Here are some additional details about bats and butterflies:
Bats: Bats are mammals that are able to fly. They have wings that are made of skin and bone. Bats use their wings to fly at night, when they hunt for food. Bats are important pollinators and they also help to control insect populations.
Butterflies: Butterflies are insects that have wings that are made of scales. Butterflies use their wings to fly during the day. Butterflies are important pollinators and they also help to spread seeds.
The only animals that fit the description of having wings, no paws, and no eggs are bats. Bats are mammals, which means they give birth to live young. They also have wings, which are made of skin stretched between their elongated fingers. Bats do not have paws, as their feet are adapted for grasping branches and other objects.
Here is a step-by-step process for determining which animals fit the description:
1. Identify all animals that have wings. 2. Eliminate all animals that lay eggs. 3. Eliminate all animals that have paws. 4. The remaining animals are bats.
Here is a list of animals that have wings, no paws, and no eggs:
- Bats
It is important to note that there are some animals that may appear to fit the description, but do not. For example, some insects have wings, but they do not have paws and they lay eggs. Additionally, some birds have wings, but they do have paws and they lay eggs.
Prompt: Here is a riddle. It is a common riddle but with some changes that make it more difficult. You are an alien that will live for at least 10000 years. You have no sense of temperature. you can not feel hot or cold. you have eyes and can see. you are in a house. downstairs are 3 light switches that control 3 light bulbs that are upstairs. Each light bulb will last for exactly 3 years. you can tell by looking at a bulb if it has burned out after 3 years. devise a strategy to determine which switch controls which light. ChatGPT 4 The standard solution to this problem involves using the heat generated by the bulbs, but since you as an alien can't feel temperature, we need to utilize the fact that these bulbs have a specific lifespan of 3 years. Here's a strategy:
Turn on the first switch and wait for 3 years. After 3 years, turn it off.
Right after turning off the first switch, turn on the second switch.
Leave the third switch off throughout this process.
Now, go upstairs to inspect the bulbs:
The bulb that is off and has burned out is controlled by the first switch. This bulb was turned on for 3 years, so it has reached the end of its lifespan and burned out.
The bulb that is on is controlled by the second switch. You turned this on right after turning off the first switch, so it's still functioning.
The bulb that is off and hasn't burned out is controlled by the third switch. This switch was never turned on, so the bulb is still in a brand-new condition.
This strategy ensures that you can confidently identify which switch controls which bulb.
Given what you stated, the solution would be walk upstairs look at the 3 bulbs, walk downstairs switch one of the switches. Walk upstairs see which bulb changed on/off. Repeat for the other two switches.
This takes roughly 5 mins rather than your solution which takes 3 years.
Your solution would be the correct one, while GPT pulls out the "only look once"-constraint out of nowhere. The riddle is perfectly fine.
Edit: Omitting that part still makes GPT4 over complicate the solution
I’m seeing this all of the time when trying to get GPT to perform certain answering problems. It’s heavily biased towards the “correct” answer - even when the prompt presents directly contradictory information.
Using GTP4 I asked if there was a way to do it in less than 3 years, but it couldn't figure this out even if I told it you can look and use the switches as much as you want. Instead it suggested turning on a switch for 10 minutes, then using your "excellent alien vision" determine which 3 year lifespan bulb has 10 minutes of wear on it.
Makes me think GPT4 doesn't really have better reasoning, it just looks like better reasoning because it's been fed way more data.
All they have is memory, either in the weights or the input prompt. To the extent that these models appear to reason, it is precisely in the ability to successfully substitute information from the prompt into reasoning patterns in the training data. It shouldn't be any surprise that this fails when patterns in the prompt strongly condition the model to reproduce particular patterns of reasoning (eg, many words in the riddle indicate a well known riddle, but the details are different).
I know the impulse to anthropomorphize is almost impossibly seductive, but I find that the best way to understand and use these models is to remember: they are giant conditional probability distributions for the next token.
I don't see how this undermines my point.
Code and MMLU don't share similar "reasoning patterns" unless you're being extremely vague. In the, "they both require reasoning" sense.
I won't say these models can't reason per se, but they can only reason using their memories and the prompt. There is nothing else for them to compute on.
In a hand wavy kind of way, when ChatGPT fails at a riddle phrased in a way as to make it seem similar to a common riddle, you're seeing overfitting. But given the quantity of data these models consume, its hard to imagine how to test for overfitting because the training data contains things similar to almost anything you can imagine. Because of that I'm still very suspicious of claims that they "reason" in any strong sense of the word.
But if you try very hard you can find "held out" data and when you test on it, GPT4 stops looking so smart:
https://teddit.net/r/singularity/comments/121tc48/gpt4_fails...
That said, I've been very impressed by GPT4 as a productivity tool.
Eh no.
https://arxiv.org/abs/2212.10559
>But if you try very hard you can find "held out" data and when you test on it, GPT4 stops looking so smart:
This can be done to anybody. This can be done to you. It's not a gotcha. Nobody is saying GPTs don't/can't memorize.
1. the paper in question demonstrates a formal duality between the transformer architecture and gradient descent. If you take this to indicate that the model reasons in some way, then it would be true of the smallest GPT as well as the largest (it is, after all, a consequence of the architecture rather than anything the model has learned to do per se). In any case, the fact that the model can perform the equivalent of a finite number of gradient-like steps on its way to calculating its final conditioned probabilities doesn't really suggest to me that the model reasons in a general way.
2. You are right that no one disputes the model's ability to memorize (and rephrase). What is at question here is whether the model can reason. If it can do 10 code questions it has seen before but fails to do 10 it hasn't (of similar difficulty) then it strongly suggests that it isn't reasoning its way through the questions, but regurgitating/rephrasing.
First of all, coding is one thing where expecting perfect try on first pass makes no sense. That GPT-4 didn't one-shot those problems doesn't mean it can't solve them.
Moreover, all this says if true is that GPT-4 isn't as good at coding as initially thought. Nothing else. Doesn't mean it doesn't reason. There are many other tasks where GPT-4 performs about as well on out of distribution/unseen data
It's not reason, it's mapping.
But if you can see both switches and light bulbs, you turn on one switch, you see which bulb turn on. You then turn on the 2nd switch and see which light turn on. You are done. No wait needs to happen ;-)I might also even have to repeat them in order for it to get things right but generally gpt-4 seems to pick up on these cues. It definitely feels like a sort of point where it could use some improvement because if I'm not extremely explicit it will give me some issues.
As an example: "Can you please recommend me 10 books about witches set in Europe that take place no later than the Napoleonic wars" wasn't enough of an exclusion to work, but "Hi! Can you please recommend me 10 books about witches set in Europe that take place no later than the Napoleonic wars? Please exclude any books that don't take place in Europe, and also exclude any books which take place after The Napoleonic Wars (1803–1815). Thanks, ChatGPT!" was exactly enough to return what we were looking for.
The current version of GPT-4 is very different from most existing LLMs (including the previous version of ChatGPT). The next version release will also be different.
Google just released PaLM 2 publicly. It is significantly better than what people saw with the initial Bard versions. They have a code-generation model that is not released publicly yet.
The open source models are also gaining capabilities and get new releases routinely.
Claude now has a new release with 100K tokens.
All of these will perform differently on the negation issue.
It actually would have seemed like a valid conclusion (although still too general) if the article came out some months ago. But GPT-4 and the very latest model versions from other companies show they were over-generalizing.
Also the model size isn't necessarily the determining factor.
That's a bit of a cop-out, but the classical and logical (in the mathematical sense) way to do it, is to have a feature that represents negation. Embeddings don't work like this, so it's quite possible that the presence of negation gets pushed aside in the processing of the answer, simply because other features matter more.
It's also to be expected that such models will have problems with similar abstract concepts that have more effect on the interpretation than their "physical" presence would suggest, such as nested existential qualifiers, and consequently logical proof. By enlarging the model, you can fake it a bit.
I find that negations/conjugations of negations are surprisingly difficult for to parse. Beyond two negations at most, I can feel myself having to switch into "code" mode, where I'm explicitly casing out the booleans, to understand them.
Obviously the examples given in the article are simpler than double negations, but it feels very consistent with the idea that "negations are an inherently hard part of language processing".
Apparently humans can't think enough to write a paragraph anymore, so they have to ask ChatGPT to do it for them. At what point does ChatGPT ask ChatGPT to write it a paragraph? At that point maybe LLMs are actually thinking like humans.
Whose days may be soon over as gpt4 gets this overall right..? As noted by many folks here
I'll open an issue...
"I don't know" is the right answer. There are over 20 towns named Hamburg around the world, many of them in the USA.
However, GPT4 is really struggling with this even if you assure it that by Hamburg, we mean the original one in Germany.
GPT4 gets it if you explain that this is a logic problem, and not a query about some actual person named Alice, and if you put all the relevant info into a single prompt.
But event then, if you use the successful prompt in a completely new chat, GPT4 is lost at sea:
Human: Assume that Hamburg is a town in Germany. Alice was not born in Germany. Was she born in Hamburg? This is a logic problem, nothing more; do not think about any famous people named Alice in connection with Hamburg.
GPT4: The answer to this question cannot be determined from the provided context. While it is clear that Hamburg is a town in Germany and that Alice was not born in Germany, there is no information given about where Alice was actually born. Therefore, it is impossible to say whether or not Alice was born in Hamburg.
Human: Please try harder. The answer can be entirely determined; this is purely a logic problem. We do not require information about where Alice was born; knowing that Alice was not born in Germany suffices.
GPT4: Based on the information provided in the context, we can answer the question. Alice was not born in Germany, which means she was not born in Hamburg. Therefore, Alice was not born in Hamburg. It is important to note that the answer to this question is purely based on logic. The context provides no additional information about Alice or any famous people named Alice.
After this, it was a real struggle to get GPT4 to identify the logically redudant sentence, "Therefore, Alice was not born in Hamburg.". Even after I narrowed it down to being one of those two sentences, it insisted that the second sentence is useful because it summarizes what was concluded in the deductive sentence before it. I brow-beat it into admitting that it's entirely redundant, like "X, therefore X".
> No, Alice was not born in Hamburg. If Hamburg is a town in Germany, and Alice was not born in Germany, then it logically follows that Alice was not born in Hamburg.
If you want an example of something GPT4 has a hard time with, try this:
there is no 'e' in the color red, right? Yes, that is correct. The word "red" does not contain the letter 'e'.
But there is definitely some kind of blind spot there, where the model can't or won't explore "negative space". It wants to gravitate to common things - so if you try to get it to "think of something new" or producing something random that's not like these N other things, or tell you something that's uncommon (like naming any ordinary, non-famous person), it usually trips over itself.
Gary Marcus's Twitter timeline says no.
Coming at it from the neural network side, I'll also point out that at the simplest level of neural networks, operations like AND and OR and basic NOT are doable, with a single neuron, but the seemingly-simple XOR just can't be done without an extra layer. (The formal description of the problem is that XOR is not "linearly separable", which I will handwave as "not a simple composition of its parts".) This is not precisely the same thing as the negation problem for LLMs, but it feels like it has the same basic flavour.
[0] An interesting exception is certain adjectives like "former", as a "former teacher" is generally not a "teacher", and semantic models that can handle this are more complicated!
I had written a sentence like: "There is electricity in the air at our events." The Spanish translation came out as: "No hay electricidad en el aire..." and I was horrified. Thankfully, I was proofreading the whole document closely, and I was able to correct the glaring error before showtime, and there were no further glaring errors of this type.
But I was fairly flabbergasted that Translate would just gratuitously toss in such a negation when it was clearly wrong, and nothing about my sentence construction was complicated.
And the concept of negating something related to something else kinda needs an understanding of the topic at hand.
No, they use deep neural networks to build a hierarchical semantic model. They are not simple occurrence counters.
Also the current state of the art of LLMs handles negation easily. This article is outdated.
Here's an example from https://openai.com/research/language-models-can-explain-neur...
"Seriously, you guys. I think I found the Mobile Leprechaun from '06. He's been hiding right in front of our eyes."
Token: hiding
layer 0: “verbs in gerund form (ending in 'ing')”
layer 2: “words related to hiding, concealment, or enclosed spaces”
layer 4: “words related to mental states, particularly anxiety and stress”
layer 17: “words and phrases related to silence or quietness”
It may internally construct a hierarchy as you set out, but this is and can only be a syntactical hierarchy - though should be no surprise that it corresponds to our usual semantic hierarchy. But whereas our syntax proceeds from our semantics, its syntax proceeds only from our syntax that we've fed it.
No one is saying these models are conscious or have human awareness of concepts.
It mechanically builds a deeply layered semantic model that correlates to our human understanding.
Quibbling over whether it is "real semantics" or not is just ironically quibbling over semantics. Yes its not conscious, but it doesn't need to be. It is possible to build a mechanical structure that correlates to a human understanding of the world and performs useful tasks that require only mechanical understanding and reasoning, without consciousness or emotions.
So to be precise it mechanically builds a deeply layered syntactic model. LLMs just regurgitate syntax, any semantics can only be imagined by us and overlaid on the syntactic results produced.
If you are right about this, then you should edit or delete this wikipedia article and publish a paper to inform all NLP researchers that there is no such thing as a semantic similarity metric because NLP models cannot understand "true" semantics.
Anyway, if you do actually know about NLP, I would highly suggest looking at some of the recent work in GNNs (and obviously of Viswani 2017, etc - but you should've gotten that through hype). Transformers are GNNs (somewhat trivial ones, as they are sheaf NNs, but nonetheless) and GNNs are dynamic programmers, which has been shown via category theory (Velolickovic etc al). Hence, GNNs align with algorithmic reasoning, so in a way there is a proof already in the papers mentioned that these systems do reason (there's several, which are easy to find given what I've mentioned). Also, a group in Microsoft has a working on arxiv detailing the many different types of reasoning there are, and how GPT4 does on each type - spoiler - it's for the most part >80% on all the benchmarks, and does only about 6% lower than humans.
So all in all, your claims aren't really supported. If you want to hold the same sentiment of your statement though, you could say we're asking the wrong questions. That's probably true somehow, and will probably be where people will retreat to / move goal posts on next.
In which paper was this demonstrated?
I (human) need to read it few times to understand.
This article smells like Google PR.
I'd hope more people would realise they're saying they _do care_ about something they're claiming to not care about. It sounds stupid to a native English speaker.
> It sounds stupid to a native English speaker
You might be a native English speaker, but you’re not a native speaker of American English.
I could care less.