Similarly for Apple’s “red herring” paper, simply adding a generic caveat to “disregard irrelevant factors” (without specifying which ones) restored performance even in the weaker local llama models back then.
The flaw was not in the reasoning; the flaw seems to be simply that the assumptions we make are often different from the assumptions it makes. I wonder if that might be a fundamental underlying cause of misalignment.
If you were home and a family member asked you that question, you'd probably criticise the question rather than answering. LLM are RLHF'd into being milk-toast helpers that just try to answer questions like that with no criticism.
This is all beside the fact that the world of AI has changed pretty dramatically in the last few months.
The only context that these LLMs have is whatever is in their context window. Not only that, the context is updated on each token emitted with the new token. It's a chaotic process, and a flaw inherent to the technology. It's not going to be fixed anytime soon.
Everything I said applies to that question as well. The only difference is that'd I'd probably also ask "why are you phrasing that question so strangely?"
It’s nonsense to test if a product that is marketed and sold as being able to provide generalised intelligence on demand, does what it says on the tin?
Check yourself