Could you provide an argument for why an LLM is a fancy markov chain that does not apply equally well to a human?
Could you provide an argument for why an LLM is a fancy markov chain that does not apply equally well to a human?
You can think against yourself, a LLM have troubles doing so. Also, they fail spectacularly when asked to do real-life operation: "I have to buy two bagettes at one euros, then five chocolatine, croissants and raisin bread at 1.40, 1.20 and 1.60 respectively, how much should i take with me?" when in my head, i just know it'll be between 20 and 25 in seconds (and in fact it's 23, i took random numbers but they are quite easy to add).
> To calculate the total cost, first find the cost for each item:
> Two baguettes at 1 euro each: 2 * 1 = 2 euros
> Five chocolatines at 1.40 euros each: 5 * 1.40 = 7 euros
> Five croissants at 1.20 euros each: 5 * 1.20 = 6 euros
> Five raisin breads at 1.60 euros each: 5 * 1.60 = 8 euros
> Now, add up the cost of each item:
> 2 euros (baguettes) + 7 euros (chocolatines) + 6 euros (croissants) + 8 euros (raisin breads) = 23 euros
> You should take 23 euros with you to purchase all the items.
Are you sure you're using GPT-4 and not 3.5? GPT-4 is incomparably more competent compared to GPT-3.5 on logical tasks like this (trust me, I've had it solve much more complicated questions than this), and you aren't using GPT-4 on chat.openai.com unless you're paying for it and deliberately picking it when creating a new chat.
Edit: Here's an example of a more complicated question that GPT-4 answered correctly on the first try: https://i.imgur.com/JMC7jsw.png
Funnily enough, this was also a problem that a friend posed to me while trying to challenge the reasoning ability of GPT-4. As you can see (cross-reference it if you like), it nailed the answer.