Claude 3 Opus suspects it is being tested from benchmark question
twitter.com
twitter.com
It's not "suspecting" anything except in the ridiculous, sloppy use of language, analogy ridden way AI proponents like to state things.
"suspect" had a very good english language meaning tied to intelligence and thinking. Polluting the meaning by applying it to a complex software system strongly implies belief the software IS thinking. It's not scientifically observationally neutral language.
If I was in peer review in this field, I'd fail any paper which wrote this kind of thing.
Claude 3 Opus can be used by people to distinguish use between test/benchmarking, and intentional use. Naieve use of Claude 3 for benchmarking may be led astray because the LLM appears to behave differently in each case. This is interesting. It's not a sign of emergent intelligence, its refinement of meaning in the NLP of the questions and their context.
To me there is a clear difference. Ai language is not as direct as flying.
Here, when tokens are predicted in sequence developers are suggesting a motive behind the predicted sequence.
But ships and submarines don't swim. Although submarines do dive.
Coming back from the finer points of the english language, the LLM in question is likely to be trained on text discussing LLM evaluation. So it ended up generating something about it.
You LLM "activists" are doing yourselves a disservice speaking about them in religious tones. You lose some credibility.
The word "suspecting" is not a word with religious connotations.
Did the Chinese Room suspect anything?
In LLM we're being taken to "the intentionality emerges, it wasn't coded in" Which sould be disputed, but it's how I see people inferring "behaviour" in this space.
Agents yes, I do have problems with, having just written of them in a different but related thread (different HN story) I begin to suspect the seeds of this linguistic problem lie deep. They were there in "clippy" and the personification of the system.
For reference, GPT4-Turbo has perfect recall to 64k and drops pretty badly past that until 128k: https://twitter.com/GregKamradt/status/1722386725635580292