For those wondering why ask these questions, we're testing for overfitting. Yes, you can overfit even if your training and validation curves don't diverge. If the model fails these questions, they've clearly overfit. But the solution space is very large and complex, so overfitting in one regime doesn't mean it didn't underfit another.
GPT 3.5: https://chat.openai.com/share/328cc39d-7fb3-4726-92a2-29437c...
LLaMA 2 70B chat: https://hf.co/chat/r/KX5H3P2
Falcon 180B Chat: https://hf.co/chat/r/V8XxdYh
None of these models can consistently get the answer right. GPT used to actually explain to me the difference between a pound and kilogram (correctly) and then use that answer to justify why they are the same. Such an answer is very clearly a demonstration of a lack of understanding as it isn't even remotely self consistent.
LLMs are powerful and amazing technologies. But we can also critique them. Hyping up models hinders the ability to improve them because every thing (LLMs, humans, governments, whatever) has limitations and is worthy of critique. But criticism isn't saying something is garbage and too many people confuse this.
https://chat.openai.com/share/34df3eb0-c41c-4c5b-80d2-48f5d0...
I'll be honest, the only reason I knew to read your variation of the farmer-boat problem carefully is because you set me up for success - I knew it wasn't going to be the original.
Curiously, I haven't been able to get it to one-shot the question, even by prefixing it with instructions.
Here, I'll demonstrate it. First I'll prod it with vague responses about it just generally being wrong. Trying to leak no information to the model. We'll try to slowly add a bit more and then use your followups. Notice that the model cannot escape the overfit regime without your strong hints. You told the model specifically what to consider. You told it that it is a trick question of a trick question. This is not how a human would handle the situation. To my followups they wouldn't spit out the same answers. They'd actually likely followup with questions if they were confused. Which is a behavior I've never seen from an LLM: asking clarifying questions.
https://chat.openai.com/share/57ab9bca-326d-45cb-9257-7fb8c2...