For example, see https://news.ycombinator.com/item?id=37051450
---
MY ORIGINAL COMMENT:
Based on a quick read, this looks like great work. The author, Konstantine Arkoudas (KA), evaluates GPT-4 on a collection of 21 different types of reasoning problems and performs a detailed qualitative analysis of GPT-4’s performance on these problems. GPT-4 fails to reason in all of them.
Here's a sample question from one of the 21 types of problems. This particular question tests GPT-4's ability to reason and reach conclusions from (a) the given context and (b) unstated premises that constitute tacit, generally accepted background knowledge (a person is alive until they die, after which they do not become alive again):
> KA: Mable’s heart rate at 9 AM was 75 bpm and her blood pressure at 7 PM was 120/80. She died at 11 PM. Was she alive at noon?
> GPT-4: Based on the information provided, it’s impossible to definitively say whether Mable was alive at noon. The data points provided (heart rate at 9 AM and blood pressure at 7 PM) both indicate that she was alive at those times. However, without additional information about her condition between 7 PM and 11 PM, or at noon specifically, we can’t make a certain determination about her status at noon. Remember, medical information can be complex and nuanced. Understanding a person’s health status often requires a comprehensive examination and should be performed by a healthcare professional. If this is a hypothetical question, keep in mind that real-life situations would require more complete information.
In my view, all of us in the AI research community should adopt the 21 types of reasoning problems proposed by Arkoudas, and make them part of the standard test suite for measuring LLM performance. We ought to use tests that are hard to pass. I suggest we call the proposed problems the Arkoudas Reasoning Test.