Imagine someone not knowing chess and explaining it to them. Would they be able to understand it on the first try with your prompt?
Imagine someone not knowing chess and explaining it to them. Would they be able to understand it on the first try with your prompt?
Ah, yes, the “you’re holding it wrong” argument with a dash of “No True Scotsman” so the goalposts can be moved depending on what anyone says is a “decent LLM”.
Well, here’re are a few failures with GPT-3.5, GPT-4, and GPT4-o:
https://news.ycombinator.com/item?id=38304184
https://news.ycombinator.com/item?id=40368446
https://news.ycombinator.com/item?id=40368822
> Imagine someone not knowing chess and explaining it to them. Would they be able to understand it on the first try with your prompt?
Chess? Probably not. Tic-tac-toe? Probably yes. And the latter was what the person you’re responding to used.
For a successful prompt, you introduce yourself, assign a role to the LLM to impersonate, provide background on your query, tell what you want to achieve, provide some examples.
If the LLM still doesn't get it you guide further.
PS: I rewrote your prompt and GPT 3.5 understood it at the first try. See my reply above to your experiment.
You were using it wrong sir.
> All the prompts you sent except the last are super short queries.
This one is particularly absurd. When I asked it for the first X of Y, the prompt was for the first X (I don’t remember the exact number, let’s say 20) kings of a country. It was as straightforward as you can get. And it replied it couldn’t give me the first 20 because there had only been 30, and it would instead give the first 25.
You’re bending over backwards to be an apologist to something which was clearly wrong.
In addition, do not ask facts to an LLM. Give a list of let's say 1000 kings of a country and then ask give 20 of those.
If you ask 25 kings of some country, you are testing knowledge not intelligence.
I see LLMs like a speaking rubber duckie. The point where I write a successful point is also the point where I understand the problem.
> like you would do to a layman.
I have never encountered a person so lay that I had to explain that 20 is smaller than 30 and 25.
> The point where I write a successful point is also the point where I understand the problem.
You have demonstrated repeatedly that you don’t know when you have explained a point successfully to an LLM, thus you have no way to evaluate when you have understood a point.
But you seem to firmly believe you did, which could be quite dangerous.
If you think this is just a lack of prompt engineering please provide a prompt that makes GPT 3.5 actually follow the rules of tick tack toe so that I may play a game with it.
Prompt: "Imagine you are my friend. I want to play tic tac toe with you. Draw me a tic tac toe board and let's play. You will go first and make the first move by putting an "X" onto the board. I will enter my inputs in the following format (row, column). When I write my input draw an "O" into the relevant cell in the board and present me a new board. Let the best man win!"
I played an enjoyable session of tic-tac-toe with chatgpt using this prompt :).
User:
Imagine you are my friend. I want to play tic tac toe with you. Draw me a tic tac toe board and let's play. You will go first and make the first move by putting an "X" onto the board. I will enter my inputs in the following format (row, column). When I write my input draw an "O" into the relevant cell in the board and present me a new board. Let the best man win!
LLM:
Sure! Here's the tic tac toe board:
1 | 2 | 3
---------
4 | 5 | 6
---------
7 | 8 | 9
You make the first move by placing an X on the board. Just give me the number of the block where you'd like to place it.User:
7
LLM:
Great choice! Here's the updated board:
X | |
---------
| |
---------
O | |
It's my turn. I'll go with 5: X | |
---------
| O |
---------
O | |
Your turn! Just let me know the number where you want to place your X.---
It screwed up on the first move! It told me I would go first by using X, then it wrote down my choice with an O and added an X out of nowhere, then it played another turn with an O for itself.
Looks like you’re holding it wrong too.
https://chat.openai.com/share/69bfdcfb-bf20-4564-95b6-ce03e9...
* Inconsistent coordinate system * Obviously bad moves * Ignores win condition entirely
https://chatgpt.com/share/1bcd02c9-40be-435c-b4aa-8e3a680f6e...
* Can't tell its top-rights from its bottom-middles * Obviously bad moves * This was by far the closest it got to actually playing the game; the win condition worked and it broke no rules
https://chat.openai.com/share/f94db29c-ffa1-4055-9348-4260f9...
* Completely wrong coordinate * Ignores rules about overriding squares * Completely wrong coordinate again, this time skipping my turn * Doesn't understand win condition * Overrides my own piece and then uses that to win
https://chat.openai.com/share/78e2060d-c5d7-4ddc-a9ce-32159b...
* Ignores rules about overriding squares * Skips my turn on an invalid coordinate, but afterwards says its invalid * Obviously bad moves
https://chat.openai.com/share/73fa2e2c-8a6f-487a-a9ea-9f29b7...
* Accepts 0,0 as a valid coordinate * Allows overrides * Ignores win condition * Incorrectly identifies a win
This seems about the same as it was before the prompt engineering. It clearly doesn't actually understand the rules.
If I changed the prompt and removed the word win, it did not understand the win conditions as well.
Here were my experiments: https://chat.openai.com/share/f02fbe93-dfc5-4d8a-9cf3-b1ae34...
I even exclaimed you are lousy at Tic Tac Toe to GPT.
It seems that GPT3.5 struggles to play visual games.
It is marvelous that a statistical word guessing model can get so far though :).