OpenAI's o1 Playing Codenames
suveenellawela.com
suveenellawela.com
This test pairs o1 with itself, which means the serializer is the deserializer. So while it's impressive that it can link 4 words, most humans could also easily link 4 with as much stretching! We just don't tend to because we can't guarantee that the other human will make the same connections we did.
This task could probably be solved nearly just as well with old school word 2 vec embeddings
I've tried. This approach is well beyond awful.
It’s odd to me that you would confidently claim it’s “beyond awful”.
I haven't surveyed all the papers, although I have read some. And all the ones that I've seen that work okay -- do so by using a language graph or word association graph in their algorithm. Not just embeddings. Even then the results don't look good to me compared to human performance.
Why does it sound crazy that it wouldn't work well? Have you used word embeddings much? Maybe you have and have good reason to think this - I don't mean to imply otherwise. But it doesn't sound crazy to me that it wouldn't work well.
If I am wrong I would love to know it.
And don't get me wrong, it's still a fun experiment! It's just that that 4 would never have worked if a human played against another human—there are simply too many other words that would be equally strongly associated:
* Gum: Gum is often wrapped in paper, so 'GUM' is strongly associated with the word 'PAPER'.
* King: King is a type of face card, which are printed on paper, so 'KING' is strongly associated with the word 'PAPER'. (Repeat for JACK.)
* Light: Paper is a lightweight material.
That's 4 others right there that are at least as closely connected in my head as LAWYER or LOG. The only reason why o1 pulled up the same four when guessing as it did when clueing is that it's the same model.
Again, I didn't mean this as a knock, just a warning about drawing too many conclusions from the test!
I think "paper" is a great clue, and those 4 words lawyer/mail/log/line match better than gum/king/light.
There's an even better reason for "lawyer/-paper" than chatgpt gave: lawyers "serve papers".
- Line (Standing in the queue…)
- London (they’re all queued up, innit?)
- Log (*backend distsys handwaving*)
- Mail (what do you think an inbox is, anyway?!)
- Round (homophone “Q” is a typographically round letter)Now obviously it’s still pretty decent at finding the clues. Probably better than a random human who hasn’t played much. Just I find the post’s level of hype overstated. It feels like the author isn’t very experienced with Codenames.
It would be interesting to compare AI:human vs human:human games to see which does better. It seems like AI:AI will overstate its success.
When I play, it's mostly about getting a good 2 clue each time. Then if you can opportunistically get a 3 or 4, that's awesome.
Some tactics come in for choosing the right pairs of 2's so you don't end up mismatched, or leaving clues that might be ambiguous with your opponent's... But that's mostly it.
It'll be fun for multiplayer! Just like how in other online games you can add in a AI to play as one of the players.
Example: The clue is "places 4" and the guessers choose 1 correctly and then 1 wrong answer, but they had achieved consensus about 2 others (and are confused about only the remaining 1). So the turns ends but they inform the clue giver to inflate by 2 next turn. That clue giver (after the other team goes) will then say the clue is "people 5" and the guessers will know that they shall select 2 places and 3 people.
This can cascade beyond just a pair of turns.
Important: the clue giver cannot acknowledge the instruction during gameplay. That would certainly extend beyond giving a clue! The guessers must know that their clue giver can play this way prior to the game commencing.
Edit: I just consulted the rules and this is the most relevant section:
> If you are a field operative, you should focus on the table when you are making your guesses. Do not make eye contact with the spymaster while you are guessing. This will help you avoid nonverbal cues.
> When your information is strictly limited to what can be conveyed with one word and one number, you are playing in the spirit of the game.
The author's use of the pronoun "you/your" switches from field ops in that first paragraph to spymasters in that second paragraph, confusingly. With that in mind, it boils down to this: field ops cannot seek non-clue information from spymasters, and spymasters cannot convey non-clue information. The strategy I'm suggesting involves neither!
You really just need an algorithm to generate unique sets of 8 or 9 from the whole board, and identifies those sets by a word.
Of course, I never play this way in my own games
There's also all kinds of not necessarily intended communicaton from the guessers in the fact that you can listen to which words they were considering and didn't pick. Nothing in the game attempt to say that you should not consider, say, whether they were going in the right or wrong direction in their guessing, but it sure can make a difference in how to approach later clues. If they were being very wrong, there might be a need to double up on words that you intended, and that your guessers missed.
In the same fashion, nothing in the game saying that I cannot listen to those guesses as a member of the other team, whether guesser or spymaster, and then change behaviors to make sure we don't hit words they considered as candidate words without very good reasons. Let them double dip on mistakes, or not make their difficult decisions easier. It's not as if the game demands that everyone that isn't currenly guessing should wear headphones to be sure they disregard what the other team says or does.
The rule on giving clues is:
"If you are the spymaster, you are trying to think of a one-word clue that relates to some of the words your team is trying to guess. When you think you have a good clue, you say it. You also say one number, which tells your teammates how many codenames are related to your clue." (emphasis mine).
The rule states that the number should be the number of words related to the clue. There is later provisions allowing you to use zero and infinity, but outside of these carve-outs (and imo the "allowed" language is telling here, since it implies any other number not equal to the number of words is not allowed) I don't think this is legal.
Taking this a step further, given that it's well-known that a clue is deemed invalid when it pertains to cards in certain non-definitional ways (sounds-like, number of letters, etc.), it seems extremely reasonable to call a clue followed by N invalid if it doesn't pertain to N cards in a definitional way.
Another thing the guessers can do if unsure about one of the tiles from the last round, is to tell the clue giver which tile they think it was. The clue giver then tries to give a clue that either tenuously links to it or clearly excludes it. That can give the clue more scope for linking to several other words. It risks giving information to the other team though so is more of an final turn play.
Doesn’t the turn end if you hit the opponents word?
If you want to get nasty, you learn to abuse the fact that the tile layouts follow rules and that you can rule out certain tiles without considering the words.
Can you clarify? Isn't the card placement random?
I don’t know most of these rules but there’s never 5 in a row; or even 4 in a row if you’re the team with one fewer (second team to play).
Edit: because the game layout is determined by choosing one of a few dozen possible layout cards and randomly rotating it
Personally I'd find that kind of play style very unfun, and would rather switch to fully randomized boards if I played enough that it became a problem.
You have to give entirely different clues depending on the people you play with.
Sometimes you can also play adversarial and introduce doubt into the opposing team by giving topic-adjacent clues that cause them to avoid one of their own cards. It works better if someone on the other team tends to be a big doubter. It also can work when the other team constantly goes back and tries to pick n+1 cards that they think they missed from the last round, which gives you a lot of room to psychologically mess with them.
Sometimes you have a clue that only really matches 2, but because only 1 of the wrong matches is a neutral card and you could match 2 more by a massive stretch, you say “4.” Worse case, they get 2 right but then they pick the neutral card but in the best case, you stand to gain 4 for a clue that should only match 2.
I like Codenames because they are many meta ways to play the game. What makes Codenames unique is that, unlike a lot of other games (Catan, Secret Hitler, CAH, etc.), it’s an adversarial team game where the team dynamics and discussions are not secret so you can use them to your advantage.
(Submitted title was "I got OpenAI o1 to play the boardgame Codenames and it's super good".)
It’s basically the same brain playing with itself. Seems quite natural to link the code names to the same words.
Let different LLMs play.
The clue giver justifies the link of Paper and Log as "written records", and between Paper and Line as "lines of text". But the guesser model connects Paper and Log because "paper is made from logs" (reaching the conclusion through a different meaning of Log), and connects Paper and Line because "'lined paper' is a common type of paper".
Similarly, in the first example, the clue giver connects Monster and Lion because lions are "often depicted as a mythical beast or monster in legends" (a tenuous connection if you ask me), whereas the guesser model thought about King because of King Kong (which I also prefer to Lion).
No, it doesn't. It reaches the conclusion because of vector similarity (simplified explanation): these explanations are post-hoc.
Video series on the topic: https://www.3blue1brown.com/topics/neural-networks
Which is to say that “why” it gives those answers is because its statistically likely within its training data that when there are the words, “why did you connect line and log with paper” the text which follows could be “logs are made of wood and lines are in paper.” But that is not the specific relation of the 3 words in the model itself, which is just a complex vector space.
If an LLM is told to do reasoning and then state the answer, it follows that the answer is basically guaranteed to be derived from the previously generated reasoning.
The best available evidence suggests this is also true of any explanations a human gives for their own behaviour; nevertheless we generally accept those at face value.
GPT-based predictive text systems are incapable of introspection of any kind: they cannot execute the algorithm I execute when I'm giving explanations for my behaviour, nor can they execute any algorithm that might actually result in the explanations becoming or approaching truthfulness.
The GPT model is describing a fictional character named ChatGPT, and telling you why ChatGPT thinks a certain thing. ChatGPT-the-character is not the GPT model. The GPT model has no conception of itself, and cannot ever possibly develop a conception of itself (except through philosophical inquiry, which the system is incapable of for different reasons).
Interestingly this could be a way to potentially reverse engineer o1’s weightings
“ I read through codenames official rules to see if using "007" as a clue was allowed, and it turns out it is! To my surprise, I even came across a Reddit post where people were discussing and justifying why this clue fits perfectly within the rules.”
In case you're wondering, the prompts are available here: https://github.com/SuveenE/codenames-ai/blob/main/utils/prom...
In that instance, the clues all matched 2-3 words, and the winning team got lucky twice (they guessed an unclued word using an unintended correlation, and their opponent guessed a different one of their unclued words.)
You also see a number of instances where the agents continue guessing words for a clue even though they've already gotten enough matches. For instance, in round 2, for the clue "Japan (2)", the blue team guesses sumo and cherry, then goes for a rather tenuous followup guess for round 1's 007 with "ring" (despite having gotten the two clued matches in the first round). A sillier example is in the final round, where the Red Team guesses 3 clues (thereby identifying all nine of their target words), then going ahead and guessing another word.
(For what it's worth, I think "shark" would have been a better guess for another 007 tie-in seeing as there are multiple Bond movies with sharks, but it's also not a match, and again, I wouldn't have gone for a third guess here when there were only two clued words.)
Imagine the energy savings if more people didn’t just automatically reach for LLMs for their pet projects.
None of them were better than a human at giving hints for 3+ words though
[1] https://jdsemrau.substack.com/p/nemotron-vs-qwen-game-theory...
the idea for this came when we asked chatgpt how to connect the words 'carrot' and 'ray'. maybe you can give a try too!
Maybe one could try having two different models play together, to see if they are genuinely good at the game or simply able to infer their own reasoning, if that makes sense.
I'm kinda bad at word games like codenames, even in my native language (french). With carrot and ray, I'd try something like "striation"? But it's really convoluted.
“the toyota yaris can move faster than the average human”
even opt-125m from years ago can pull more facts than the average human.
Star => Twinkle => Twinkle Khanna => Married to Akshay Kumar => Canadian Citizen => Maple Syrup ( Leaf ? )
The digital version could and should do this, IMO. (I don't actually know if it does, though, as I've only played the digital version a few times.)