"Count how many rs are in the word strawberry. First, list each letter and indicate whether it's an r and tally as you go, and then give a count at the end."
Llama 405b: correct
Mistral Large 2: correct
Claude 3.5 Sonnet: correct
"Count how many rs are in the word strawberry. First, list each letter and indicate whether it's an r and tally as you go, and then give a count at the end."
Llama 405b: correct
Mistral Large 2: correct
Claude 3.5 Sonnet: correct
At that point it was easier to do it myself.
EDIT: Although perhaps it's even more important when dealing with humans and contracts. Someone could deliberately interpret the words in a way that's to their advantage.
There's definitely tons of weaknesses with LLMs for sure, but i continue to be impressed at what they do right - not upset at what they do wrong.
Even some people will make this association, it's no surprise that LLMs do.
I believe this could be fixed and is worth fixing. Because it’s the only way LLM will be able to help math and physic researcher write proof and make real scientific progress
*) that is, except sometimes by making adjustments to the system prompt
I actually think the craziest part of LLMs is that how, as a developer or SME, just how much you can fix with plain english prompting once you have that intuition. Of course some things aren't fixable that way, but the mere fact that many cases are fixable simply by explaining the task to the model better in plain english is a wildly different paradigm! Jury is still out but I think it's worth being excited about, I think that's very powerful since there are a lot more people with good language skills than there are python programmers or ML experts.
Me: How many "r"s are in strawberry?
Them: What?
Me: How many times does the letter "r" appear in the word "strawberry"?
Them: Is this some kind of trick question?
Me: No. Just literally, can you count the "r"s?
Them: Uh, one, two, three. Is that right?
Me: Yeah.
Them: Why are you asking me this?
"No, I don't think I shall answer that. The question is too basic, and you know better than to insult me."
For example, the latter model answered with:
To count the number of Rs in the word "strawberry", I'll break it down step by step:
Start with the individual letters: S-T-R-A-W-B-E-R-R-Y Identify the letters that are "R": R (first one), R (second one), and R (third one) Count the total number of Rs: 1 + 1 + 1 = 3
There are 3 Rs in the word "strawberry".
We should always put some effort into prompt engineering before dismissing the potential of generative AI.
And sometimes CoT may not be the best approach. Depending on the problem other prompt engineering techniques will perform better.
Maybe the various chat interfaces already do this behind the scenes?