Ask HN: Why is GPT-4 or Claude-2 so bad at tic-tac-toe?
But after trying for 2h to get it work (even with GPT-4V) it seems like a fundamental limitation.
I've found a HN submission [1] where someone used a brute-force prompt to get it to play correctly, but as the top/only comment points out, it's a limited action space and enumerating most of it seems moot.
I was hoping for a more reasonable prompt. After all humans are able learn tic-tac-toe rapidly.
current hypotheses:
1) tic-tac-toe requires "spatial reasoning" and LLMs train on sequences (somehow GPT-4V didn't elevate that constraint)
2) tic-tac-toe requires "search" of future scenarios
Would love to hear what you think/know!
---
Previous discussion about T3 and GPT-4: https://news.ycombinator.com/item?id=35216614 (7 months ago)
[1] https://news.ycombinator.com/item?id=37626918