An AI system for solving crossword puzzles that outperforms the best humans
twitter.com
twitter.com
It feels like the challenges might be that most clues are not cross-referential, and even for those that are, most information in the puzzle is irrelevant - you only care about one answer among many, so it could be difficult to learn to find the information you need.
But maybe this sort of thing would also be helpful for theme puzzles, where answers might be united by the theme even if their clues are not directly cross-referential, and could give enough signal to teach the model to look at the puzzle context?
1. You mention an 82% solve rate. The NYT puzzle gets "harder" each day Monday through Saturday. Do you track the days separately? If so I'd be curious how much of the 18% unsolved end up on Fridays and Saturday. (for anyone who doesn't know the Sunday puzzle is outside of the M-Sat range since its a bigger puzzle).
2. Related to the above Thursday puzzles usually have "tricks" (skipped letters and what not) in them or require a Rebus (multiple letters in one space) - do you handle these at all?
3. Is this building an ongoing model and getting better at solving? Or did you have to seed it with a set of solved puzzles and clues?
Sorry didn't have time to read the whole paper.
1. Monday puzzles are the easiest for our model, and Thursdays are the most difficult. You can see a graph of day-by-day performance here: https://twitter.com/albertxu__/status/1527704535912787968
2. Our current system doesn't have any handling for rebuses or similar tricks, although Dr. Fill does. I think this is part of why Thursday is the hardest day for us, even though Saturday is usually considered the most difficult.
3. We trained it with 6.4M clues. As new crosswords get published, we could theoretically retrain our model with more data, but we aren't currently planning to do that.
Our model seems to perform well despite this "time generalization" split, but there are a couple instances where it struggled with new words. For example, we got the answer "FAUCI" wrong in a puzzle from May 2021. Even though Fauci was in the news before 2020, I guess he wasn't famous enough to show up in crosswords, and therefore his name wasn't in our training data.
I think evaluating performance by constructor would be really interesting! But we haven't done that.
One thing I was curious about - the ACPT is a crossword speed-solving competition, with time spent solving a major aspect of total score. How did you approach leveling the playing field between the human and computer competitors?
The answer to that clue is included here:
Firstly, all "serious" British crosswords are "Cryptic" ie once you figure out what the answer is, it's apparent why that's the correct clue, but figuring out the answer from the clue involves lateral thinking and some skills learned from years of staring at such clues.
e.g. Private Eye's crossword 726 (back in April), clue 23 down,
"He finally gets to penetrate agreeable person (relatively) (5)"
The correct answer is "Niece". "Nice" can mean agreeable, the final letter of "He" is E, and so by having the letter E "penetrate" the word nice you produce "niece", a person who is a relative.
[ and yes, Private Eye is a satirical magazine, the crossword clues are, likewise, intended to make you a little uncomfortable while you laugh ]
Secondly, British crosswords are arranged with black "dead" squares between letters to produce more of a lattice, in which many letters only take part in one word, as a result longer answers are common
e.g. same crossword, clue 26 across is
"Figure on getting your teeth into our statistical revelations (6,9)"
The answer was "Number Crunching".
Puzzled, I didn’t move and set about figuring out why. Eventually I realised that I had solved, in my sleep, a crossword clue that I had not even gone to bed thinking about. I’d read it at my grandma’s house earlier the previous day.
Tiny picture makes computer work on time (8)
The brain is amazing. I’m not even any good at the cryptic crossword!
Definition: Tiny picture
Wordplay: Computer [MICRO] + work [DO] + time [T]
I wonder how the dead-square terminology reached Brazil. I think the popularity of crosswords there was originally due to Italian immigrants, who might not have said "dead squares".
I've heard people call the "letters [that] only take part in one word" unches (an abbreviation for "unchecked squares").
Cryptic crosswords in the UK often have hidden prepositions and clues like this.
There may be other sites that allow it -- the software seems to power a few diff crossword sites (with certain features enabled/disabled).
Anyway, it would be interesting what the AI would do with this, would there be two hotspots in the solution space, one for each variant?
If they trained the AI on the NYT archive then they would have the results of testing it on this one.
A thesaurus will get you far, but will never get you OREOCOOKIE and CHESSBOARD as answers for "It's all black and white" (from today's puzzle).
Sure, it can do crosswords well but the average human that does crosswords well can also do a zillion other things and this type of AI is not getting us any closer to that.
Also, I'm not knocking this paper at all. I think it's a great applied paper! Stitching together techniques to actually do something is a Herculean task. Hell, it's mostly what I do too.
Not necessarily, there is good evidence that a single model can work for many tasks - for example, the recent Gato system https://www.deepmind.com/publications/a-generalist-agent is a good example. It's just that we usually don't do that because for most practical purposes we want an agent for a specific purpose, and for most research experiments we want a simpler experiment to isolate some factor, so we usually train single-task models and don't try to make general systems.
That is not obvious at all.
I can just imagine if evolution was a side spectator event, people commenting: "Broca's area just regulates breathing. And that Wernicke's area is just pattern recognition in sounds. Those aren't going to get us to anything important."
Point me to actual large generalized models in nature that aren't composed of smaller specialized functions and you might have a leg to stand on.
(Oh wait, no, those legs things are pretty specialized too, and each have their own specialized parts. Bad analogy.)
Well, good luck with your identifying an example of complex generalization without subspecialties!