for strawberry, it see it as [496, 675, 15717], which is str aw berry.
If you insert characters to breaks the tokens down, it find the correct result: how many r's are in "s"t"r"a"w"b"e"r"r"y" ?
> There are 3 'r's in "s"t"r"a"w"b"e"r"r"y".
The issue is that humans don't talk like this. I don't ask someone how many r's there are in strawberry by spelling out strawberry, I just say the word.
A token is roughly four letters [1], so, among other probable regressions, this would significantly reduce the effective context window.
[1] https://help.openai.com/en/articles/4936856-what-are-tokens-...
Yes, it's an issue. We want the convenience of sending human-legible commands to LLMs and getting back human-readable responses. That's the entire value proposition lol.
Problems can exist as instances of a class of problems. If you can't solve a problem, it's useful to know if it's a one off, or if it belongs to a larger class of problems, and which class it belongs to. In this case, the strawberry problem belongs to the much larger class of tokenization problems - if you think you've solved the tokenization problem class, you can test a model on the strawberry problem, with a few other examples from the class at large, and be confident that you've solved the class generally.
It's not about embodied human constraints or how humans do things; it's about what AI can and can't do. Right now, because of tokenization, things like understanding the number of Es in strawberry are outside the implicit model of the word in the LLM, with downstream effects on tasks it can complete. This affects moderation, parsing, generating prose, and all sorts of unexpected tasks. Having a workaround like forcing the model to insert spaces and operate on explicitly delimited text is useful when affected tasks appear.
Especially native speakers are prone to the mistake as they grew up learning english as illiterate children, from sounds only, compared to how most people learning english as second language do it, together with the textual representation.
Psychologists use this trick as well to figure out internal representations, for example the rorschach test.
And probably, if you asked random people in the street how many p's there is in "Philippines", you'd also get lots of wrong answers. It's tricky due to the double p and the initial p being part of an f sound. The demonym uses "F" as the first letter, and in many languages, say Spanish, also the country name uses an F.
No, I would actually be pretty confident you don’t ask people that question… at all. When is the last time you asked a human that question?
I can’t remember ever having anyone in real life ask me how many r’s are in strawberry. A lot of humans would probably refuse to answer such an off-the-wall and useless question, thus “failing” the test entirely.
A useless benchmark is useless.
In real life, people overwhelmingly do not need LLMs to count occurrences of a certain letter in a word.
If full artificial intelligence, as we're being promised, falls short in this simple way.
We are also not exactly looking letter by letter at everything we read.
However, this implies you need to know the word to begin with.
I can write "asdf" and you might be oblivious to what I mean. I can mention "adsf" to a JavaScript developer and he will immediately think of the tool versioning tool. Because context and familiarity is important.
"vox"
"il"
"ha"
"pen"
"tosh"
is how I read it.
A lot of schools teach kids to read with a syllabic method... so... super close to the tokenization concept.
https://chatgpt.com/share/4298efbf-1c29-474a-b333-c6cc1a3ce3...
User: Write “strawberry” one letter at a time, with a space between each letter. Then count how many r’s are in strawberry.
gpt-3.5-turbo: ASSISTANT s t r a w b e r r y
There are 2 r's in strawberry.
After some experimenting, it seems like the actual problem is that many LLMs can’t count.
User: How many r’s are in the following sequence of letters:
S/T/R/A/W/B/E/R/R/Y
gpt-4o-mini: In the sequence S/T/R/A/W/B/E/R/R/Y, there are 2 occurrences of the letter "R."
Oddly, if I change a bunch of the non-R letters, I seem to start getting the right answer.
You don't ask a human being how many r's there are in strawberry at all. The only reason you or anyone else asks that question is because it's an interesting quirk of how LLMs work that they struggle to answer it in that format. It's like an alien repeatedly showing humans an optical illusion that relies on the existence of our (literal) blind spot and using it as evidence of our supposed lack of intelligence.
Tokens aren’t the source of facts within a model. it’s an implementation detail and doesn’t inherently constrain how things could be counted.
Combinations of tokens get encoded, so if the feature isn't part of the information being carried forward into the network as it models the information in the corpus, the feature isn't modeled well, or at all. The consequence of having many character tokens is that the relevance of individual characters is lost, and you have to explicitly elicit the information. Models know that words have individual characters, but "strawberry" isn't encoded as a sequence of letters, it's encoded as an individual feature of the tokenizer embedding.
Other forms of tokenizing have other tradeoffs. The trend lately is to increase tokenizer dictionary scope, up to 128k in Llama3 from 50k in gpt-3. The more tokens, the more nuanced individual embedding features in that layer can be before downstream modeling.
Tokens inherently constrain how the notion of individual letters are modeled in the context of everything an LLM learns. In a vast majority of cases, the letters don't matter, so the features don't get mapped and carried downstream of the tokenizer.
What you're saying sounds plausible, but I don't see how we can conclude that definitively without at least some empirical tests, say a set words that predictably give an error along token boundaries.
The thing is, there are many ways a model can get around to answering the same question, it doesn't just depend on the architecture but also on how the training data is structured.
For example, if it turned out tokenization was the cause of this glitch, conceivably it could be fixed by adding enough documents with data relating to letter counts, providing another path to get the right output.
There is a lot of problems like that, that can be reformulated. For example if you ask it what is the biggest between 9.11 and 9.9, it will respond 9.9. If you look at how it's tokenized, you can see it restate an easy problem as something not straightforward even for a human. if you restart the problem by writing the number in full letters, it will correctly respnod.
Meanwhile, it helps make sure engineers and product designers who want to build a more targeted product around LLM technology know that it's not suited to tasks that may trigger those kinds of failures. This may be obvious to you as an engaged enthusiast or cutting edge engineer or whatever you are, but it's always going to be new information to somebody as the field grows.
[1] Suggestion from chatgpt3.5 for new fruit name.
There are tokens for individual letters, but the model is not trained on text written with individual tokens per letter, it is trained on text that has been converted into as few tokens as possible. Just like you would get very confused if someone started spelling out entire sentences as they spoke to you, expecting you to reconstruct the words from the individual spoken letters, these LLMs also would perform terribly if you tried to send them individual tokens per letter of input (instead of the current tokenizer scheme that they were trained on).
Even though you might write a message to an LLM, it is better to think of that as speaking to the LLM. The LLM is effectively hearing words, not reading letters.
Very large language models also “know” how to spell the word associated with the strawberry token, which you can test by asking them to spell the word one letter at a time. If you ask the model to spell the word and count the R’s while it goes, it can do the task. So the failure to do it when asked directly (how many r’s are in strawberry) is pointing to a real weakness in reasoning, where one forward pass of the transformer is not sufficient to retrieve the spelling and also count the R’s.
LOL. I would fail your test, because "fraise" only has one R, and you're expecting me to reply "3".
Here's a thought experiment: if I gave you 5 boxes and told you "how many balls are there in all of this boxes?" and you answered "I don't know because they are inside boxes", that's a fail. A truly intelligent individual would open them and look inside.
A truly intelligent model would (say) retokenize the word into its individual letters (which I'm optimistic they can) and then would count those. The fact that models cannot do this is proof that they lack some basic building blocks for intelligence. Model designers don't get to argue "we are human-like except in the tasks where we are not".
"A significant effort was also devoted to enhancing the model’s reasoning capabilities. (...) the new Mistral Large 2 is trained to acknowledge when it cannot find solutions or does not have sufficient information to provide a confident answer."
Words like "reasoning capabilities" and "acknowledge when it does not have enough information" have meanings. If Mistral doesn't add footnotes to those assertions then, IMO, they don't get to backtrack when simple examples show the opposite.
That doesn't even take into account what OpenAI has typically done to intercept queries and cover the shortcomings of LLMs. It would be useful if each model did indeed come out with a chart covering what it cannot do and what it has been tailored to do above and beyond the average LLM.
Me: spell "strawberry" with 1 bullet point per letter
ChatGPT:
S
T
R
A
W
B
E
R
R
Y
Me: How many Rs?
ChatGPT: There are three Rs in "strawberry".Never have been, never will be. They model language, not intelligence.
Meanwhile, for practical purposes, there's little arrogance needed to say that some things are preconditions for any form of intelligence that's even remotely recognizable.
1) Learning needs to happen continuously. That's a no-go for now, maybe solvable.
2) Learning needs to require much less data. Very dubious without major breakthroughs, likely on the architectural level. (At which point it's not really an LLM any more, not in the current sense)
3) They need to adapt to novel situations, which requires 1&2 as preconditions. 3
4) There's a good chance intelligence requires embodiment. It's not proven, but it's likely. For one, without observing outcomes, they have little capability to self-improve their reasoning.
5) They lack long-term planning capacity. Again, reliant on memory, but also executive planning.
There's a whole bunch more. Yes, LLMs are absolutely amazing achievements. They are useful, they imply a lot of interesting things about the nature of language, but they aren't intelligent. And without modifying them to the extent that they aren't recognizably what we currently call LLMs, there won't be intelligence. Sure, we can have the ship of Theseus debate, but for practical purposes, nope, LLMs aren't intelligent.
1,2,3) All subjective opinions.
4) 'Embodiment' is another term we don't really know how to define. At what point does an entity have a 'body' of the sort that supports 'intelligence'? If you want to stick with vague definitions, 'awareness' seems sufficient. Otherwise you will end up arguing about paralyzed people, Helen Keller, that rock opera by the Who about the pinball player, and so on.
5) OK, so the technology that dragged Lee Sedol up and down the goban lacks long-term planning capacity. Got it.
None of these criteria are up to the task of supporting or refuting something as vague as 'intelligence.' I almost think there has to be an element of competition involved. If you said that the development of true intelligence requires a self-directed purpose aimed at outcompeting other entities for resources, that would probably be harder to dismiss. Could also argue that an element of cooperation is needed, again serving the ultimate purpose of improving competitive fitness.
[0]: https://genius-u-attachments.s3.amazonaws.com/uploads/articl...
Seems to me like a legit question for a young child to answer or even ask.
They're not, but laymen shouldn't think that the LLM tests they come up with have much value.
They're pointless unless you have the expertise to check the output. Just because you can type text in a box doesn't mean it's a tool for everybody.
I'm seeing everyone and their parents using chatgpt.
Think about math and logic. If a single symbol is off, it’s no good.
Like a prompt where we can generate a single tokenization error at my work, by my very rough estimates, generates 2 man hours of work. (We search for incorrect model responses, get them to correct themselves, and if they can’t after trying, we tell them the right answer, and edit it for perfection). Yes even for counting occurrences of characters. Think about how applicable that is. Finding the next term in a sequence, analyzing strings, etc.
In that case the tokenization is done at the appropriate level.
This is a complete non-issue for the use cases these models are designed for.
I think that's a perfect description by the way, I'm going to steal it.
As an aside, I wish we would call all of this stuff pseudo intelligence rather than artificial intelligence
Are LLMs human? No. Can they do everything humans do? No. But they can do a large enough subset of things that until now nothing but a human could do that we have no choice but to call it "thinking". As Hofstadter says - if a system is isomorphic to another one, then its symbols have "meaning", and this is indeed the definition of "meaning".
You also indirectly answered my initial question -- so thanks!
I don't disagree they produce astonishing responses but the nuance of why it's producing that output matters to me.
For example, with regard to social mores, I think a good way to summarize my hang up is that my understanding is LLMs just pattern match their way to approximations.
That to me is different from actually possessing an understanding, even though the outcome may be the same.
I can't help but draw comparisons to my autistic masking.
When you read and comprehend text, you don't read it letter by letter, unless you have a severe reading disability. Your ability to comprehend text works more like an LLM.
Essentially, you can compare the human brain to a multi-model or modular system. There are layers or modules involved in most complex tasks. When reading, you recognize multiple letters at a time[], and those letters are essentially assembled into tokens that a different part of your brain can deal with.
Breaking down words into letters is essentially a separate "algorithm". Just like your brain, it's likely to never make sense for a text comprehension and generation model to operate at the level of letters - it's inefficient.
A multi-modal model with a dedicated model for handling individual letters could easily convert tokens into letters and operate on them when needed. It's just not a high priority for most use cases currently.
[]https://www.researchgate.net/publication/47621684_Letters_in...
So the blob wasn’t trained to do that (yeah low utility I get that) but it also doesn’t know it doesn’t know, which is an another much bigger and still unsolved problem.
(A quick demo of this in the langchain docs, using claude-3-haiku: https://python.langchain.com/v0.2/docs/integrations/tools/ri...)
It's also funny, since this strawberry question is one where a model that's seriously good at predicting the next character/token/whatever quanta of information would get it right. It requires no reasoning, and is unlikely to have any contradicting text in the training corpus.
Does something deal with separate symbols rather than just meaning of words? Then yes.
This affects spelling, math (value calculation), logic puzzles based on symbols. (You'll have more success with a puzzle about "A B A" rather than "ABA")
> It requires no reasoning, and is unlikely to have any contradicting text in the training corpus.
This thread contains contradictions. Every other announcement of an llm contains a comment with a contradicting text when people post the wrong responses.
For example, if we tokenize Welcome to Hacker News, I hope you like strawberries. The Llama 405B tokenizer will tokenize this as:
Welcome Ġto ĠHacker ĠNews , ĠI Ġhope Ġyou Ġlike Ġstrawberries .
(Ġ means that the token was preceded by a space.)Each of these pieces is looked up and encoded as a tensor with their indices. Adding a special token for the beginning and end of the text, giving:
[128000, 14262, 311, 89165, 5513, 11, 358, 3987, 499, 1093, 76203, 13]
So, all the model sees for 'Ġstrawberries' is the number 76204 (which is then used in the piece embedding lookup). The model does not even have access to the individual letters of the word.Of course, one could argue that the model should be fed with bytes or codepoints instead, but that would make them vastly less efficient with quadratic attention. Though machine learning models have done this in the past and may do this again in the future.
Just wanted to finish of this comment with saying that the tokens might be provided in the model splitted if the token itself is not in the vocabulary. For instance, the same sentence translated to my native language is tokenized as:
Wel kom Ġop ĠHacker ĠNews , Ġik Ġhoop Ġdat Ġje Ġvan Ġa ard be ien Ġh oud t .
And the word voor strawberries (aardbeien) is split, though still not in letters.This is Hacker News, we are usually interested in how things work.
To paraphrase how this thread started - it was someone testing different boats to see whether they can simply float - and they couldn’t. And the reply was questioning the validity of testing boats whether they can simply float.
At least this is how it sounds to me when I am told that our AI overlords can’t figure out how many Rs are in the word “strawberry”.
For instance you could ask it for a JavaScript function to count any letter in any word and pass it r and strawberry and it would be far more useful.
Having edge cases doesn't mean its not useful it is neither a free assastant nor a coder who doesn't expect a paycheck. At this stage it's a tool that you can build on.
To engage with the analogy. A propeller is very useful but it doesn't replace the boat or the Captain.
"create a javascript function to count any letter in any word. Run this function for the letter "r" and the word "strawberry" and print the count"
ChatGPT-4o => Output is 3. Passed
Claude3.5 => Output is 2. Failed. Told it the count is wrong. It apologised and then fixed the issue in the code. Output is now 3. Useless if the human does not spot the error.
llama3.1-70b(local) => Output is 2. Failed.
llama3.1-70b(Groq) => Output is 2. Failed.
Gemma2-9b-lt(local) => Output is 2. Failed.
Curiously all the ones that failed had this code (or some near identical version of it)
```javascript
function countLetter(letter, word) {
// Convert both letter and word to lowercase to make the search case-insensitive
const lowerCaseWord = word.toLowerCase();
const lowerCaseLetter = letter.toLowerCase();
// Use the split() method with the letter as the separator to get an array of substrings separated by the letter
const substrings = lowerCaseWord.split(lowerCaseLetter);
// The count of the letter is the number of splits minus one (because there are n-1 spaces between n items)
return substrings.length - 1;
}// Test the function with "r" and "strawberry"
console.log(countLetter("r", "strawberry")); // Output: 2 ```
This prompt "please write a javascript function that takes a string and a letter and iterates over the characters in a string and counts the occurrences of the letter"
Produced a correct function given both chatGPT4o and claude3.5 for me.
Just like Dall-E is not layering coats of pain to make a watercolor... it just makes something that looks like one.
Your LLM (or you) should run the code in a code interpretor. Which ChatGPT did because it has access to tools. Your local ones don't.
Without looking at the word 'strawberry', or spelling it one letter at a time, can you rattle off how many letters are in the word off the top of your head? No? That is what we are asking the LLM to do.
In the model the individual letters hold little meaning. Words are composed of letters but simply because we need some sort of organized structure for communication that helps represents meaning and intent. Just like our color blue/blau/azul/blu.
Not faulting them for asking the question but I agree that the results do not undermine the capability of the technology. In fact it just helps highlight the constraints and need for education.
Because they don't have intelligence.
If they did, they could count the letters in strawberry.
They fundamentally perceive the world in terms of tokens, not "letters".
Nor do they understand how intelligence works.
Humans don't read text a letter at a time. We're capable of deconstructing words into individual letters, but based on the evidence that's essentially a separate "algorithm".
Multi-model systems could certainly be designed to do that, but just like the human brain, it's unlikely to ever make sense for a text comprehension and generation model to work at the level of individual letters.
Of course. Because these models have no intelligence.
Everyone who believes they do seem to believe intelligence derives from being able to use language, however, and not being able to tell how many times the letter r is in the word strawberry is a very low bar to not pass.
"How many R letters are in the following? Keep a running count. s t r a w b e r r y"
They are terrible at counting letters in words because they rarely see them spelled out. An LLM trained one byte at a time would always see every character of every word and would have a much easier time of it. An LLM is essentially learning a new language without a dictionary, of course it's pretty bad at spelling. The tokenization obfuscates the spelling not entirely unlike how verbal language doesn't always illuminate spelling.
Iow, what makes you think that it’s exactly letter-tokens that help it and not the high-level concept of spelling things out itself?
Once it is just counting lists then it's probably drawing on a higher level capability, yeah.
Also there are some cases where regular people will stumble into it being awful at this without any understanding why (like asking it to help them with their wordle game.)
"Count how many rs are in the word strawberry. First, list each letter and indicate whether it's an r and tally as you go, and then give a count at the end."
Llama 405b: correct
Mistral Large 2: correct
Claude 3.5 Sonnet: correct
There's definitely tons of weaknesses with LLMs for sure, but i continue to be impressed at what they do right - not upset at what they do wrong.
Even some people will make this association, it's no surprise that LLMs do.
I believe this could be fixed and is worth fixing. Because it’s the only way LLM will be able to help math and physic researcher write proof and make real scientific progress
*) that is, except sometimes by making adjustments to the system prompt
I actually think the craziest part of LLMs is that how, as a developer or SME, just how much you can fix with plain english prompting once you have that intuition. Of course some things aren't fixable that way, but the mere fact that many cases are fixable simply by explaining the task to the model better in plain english is a wildly different paradigm! Jury is still out but I think it's worth being excited about, I think that's very powerful since there are a lot more people with good language skills than there are python programmers or ML experts.
Me: How many "r"s are in strawberry?
Them: What?
Me: How many times does the letter "r" appear in the word "strawberry"?
Them: Is this some kind of trick question?
Me: No. Just literally, can you count the "r"s?
Them: Uh, one, two, three. Is that right?
Me: Yeah.
Them: Why are you asking me this?
"No, I don't think I shall answer that. The question is too basic, and you know better than to insult me."
At that point it was easier to do it myself.
EDIT: Although perhaps it's even more important when dealing with humans and contracts. Someone could deliberately interpret the words in a way that's to their advantage.
Maybe the various chat interfaces already do this behind the scenes?
We should always put some effort into prompt engineering before dismissing the potential of generative AI.
And sometimes CoT may not be the best approach. Depending on the problem other prompt engineering techniques will perform better.
For example, the latter model answered with:
To count the number of Rs in the word "strawberry", I'll break it down step by step:
Start with the individual letters: S-T-R-A-W-B-E-R-R-Y Identify the letters that are "R": R (first one), R (second one), and R (third one) Count the total number of Rs: 1 + 1 + 1 = 3
There are 3 Rs in the word "strawberry".
When this question is asked (from what the models trained on) the question is NOT “count the number of times r appears in the word strawberry” but instead (effectively) “I’ve written ‘strawbe’, now how many r’s are in strawberry again? Is it 1 or 2?”.
I think most humans would probably answer “there are 2” if we saw someone was writing and they asked that question, even without seeing what they have written down. Especially if someone said “does strawberry have 1 or 2 r’s in it?”. You could be a jerk and say “it actually has 3” or answer the question they are actually asking.
It’s an answer that is _technically_ incorrect but the answer people want in reality.
It is evidence that LLMs aren't appropriate for everything, and that there could exist something that works better for some tasks.
By allowing a language model to do function calling, you are essentially allowing it to do specialized "subconscious" thought. The language model becomes a natural language interface to the capabilities of its "subconsciousness".
A specific human analogy could be: I tell you to pick up a pen off of the table, and then you do it. Most of your mental activity would be subconscious, orienting your arm and hand properly to pick up the pen, actually grabbing the pen, and picking it up. The linguistic representation of the action would exist in your concious mind (pick up the pen), but not much else.
A language model could very easily call out to a text processing function to correctly do things like count the number of r's in the word strawberry. That is a job that your concious mind can dispatch to your subconciousness.
This tells us that the model is unaware of its own limitations, otherwise it would have figured that this tasks needs it to split the letters first.
[Question: How many r's in strawberry? First split the word into letters, one letter on each line then answer the question.]
[Answer: Let's split the word "strawberry" into letters:
s t r a w b e r r y
Now, let's count the number of 'r's:
There are 2 'r's in the word "strawberry".]
[Question: Did you count the number of r's?]
[answer: Yes, I did. Here's the count:
r r
So, there are 2 'r's in the word "strawberry".]
I suspect the issue is these models have to be trained to accomplish tasks, and this is not a task it makes sense to train a model to perform. You might train it to do this with function calling one day, calling a python library to count the r's in a string, but actually manually counting r's just doesn't play to an LLM's strengths.
Of note, this happened because earlier I had asked it without CoT and it incorrectly stated two “r”s. Once I started a new conversation (so fresh context window), it gave the right answer (3) and no longer claimed the “still the second r” as shown below.
Let's think through this step-by-step:
1. First, let's spell out "strawberry" letter by letter: s-t-r-a-w-b-e-r-r-y
2. Now, let's count the "r"s as we go through each letter: s (0) t (0) r (1) - First "r" found a (1) w (1) b (1) e (1) r (2) - Second "r" found r (2) - This is still the second "r" y (2)
3. After going through all the letters, we've counted 2 "r"s.
Therefore, the word "strawberry" contains 2 "r"s.
Me: How many Rs are in the word strawberry? Think step by step
Claude: Let's approach this step-by-step:
1. First, let's write out the word: strawberry
2. Now, let's go through each letter one by one: s - not an R t - not an R r - this is an R (count: 1) a - not an R w - not an R b - not an R e - not an R r - this is another R (count: 2) r - this is a third R (count: 3) y - not an R
3. After going through all the letters, we've counted 3 Rs.
Therefore, there are 3 Rs in the word "strawberry".
I think it will be a while before we get there. An LLM can lookup knowledge but can't actually perform calculations itself, without some external processor.
According to multiple sources, including linguistic analysis and word breakdowns, there are 3 Rs in the word "strawberry".
I just researched "how many r's are in strawberry?" in a search engine, and based solely on the results it found, I would have to conclude there is substantial disagreement on whether the correct answer is two or three.
No idea if there's any sense to this, but an LLM could be instructed to formulate and continually test mathematical assumptions by writing / running code and fine-tuning accordingly.
[0]:https://riza.io [1]:mailto:andrew@riza.io [2]:https://discord.gg/4P6PUeJFW5
Training data will only "fix" it in the shallow sense that it will have seen a comment like yours before. (As opposed to the deeper sense of "learning to count.")
What happens if you ask the total number of occurrences of the letter r in the word? Does it still not get it right?
How many letters R are in the word "s-t-r-a-w-b-e-r-r-y"?
The word "s-t-r-a-w-b-e-r-r-y" contains three instances of the letter "R."
How many letters R contain the word strawberry?
The word "strawberry" contains two instances of the letter "R."
1. Play tic tac toe such that you never lose
2. Which is bigger 9.11 or 9.9
3. 4 digit multiplication even with CoT prompting
"how many r's are in strawberry?"
The funny thing is, I replied, "Are you sure?" and got back, "I apologize for the mistake. There are actually two 'r's in the word strawberry."
> How many times does the letter “r” appear in the word “strawberry”?
> The letter "r" appears 2 times in the word "strawberry."
But also:
> How many occurrences of the letter “r” appear in the word “strawberry”?
> The word "strawberry" contains three occurrences of the letter "r."
Using more 'erudite' speech is a good technique to help focus an LLM on training data from folks with a higher education level.
Using simpler speech opens up the floodgates more toward the general populous.
It was also interesting to observe how GPT (4o) even tried to prove/illustrate the result typographically by placing the same word four times and putting the respective letter in bold font (without being prompted to do that).
If you find a case where forceful pushback is sticky, it's either because the primary answer is overwhelmingly present in the training set compared to the next best option or because there are conversations in the training that followed similar stickiness, esp. if the structure of the pushback itself is similar to what is found in those conversations.
> If you say it's incorrect, it will give you another answer instead of insisting on correctness.
> When you push it, it responds with the next most common answer.
Which clearly isn't as black and white as you made it seem.
How many thoughts go through your brain when you read this comment? You can give me a number but it will be a guess at best.