1. You mention an 82% solve rate. The NYT puzzle gets "harder" each day Monday through Saturday. Do you track the days separately? If so I'd be curious how much of the 18% unsolved end up on Fridays and Saturday. (for anyone who doesn't know the Sunday puzzle is outside of the M-Sat range since its a bigger puzzle).
2. Related to the above Thursday puzzles usually have "tricks" (skipped letters and what not) in them or require a Rebus (multiple letters in one space) - do you handle these at all?
3. Is this building an ongoing model and getting better at solving? Or did you have to seed it with a set of solved puzzles and clues?
Sorry didn't have time to read the whole paper.
1. Monday puzzles are the easiest for our model, and Thursdays are the most difficult. You can see a graph of day-by-day performance here: https://twitter.com/albertxu__/status/1527704535912787968
2. Our current system doesn't have any handling for rebuses or similar tricks, although Dr. Fill does. I think this is part of why Thursday is the hardest day for us, even though Saturday is usually considered the most difficult.
3. We trained it with 6.4M clues. As new crosswords get published, we could theoretically retrain our model with more data, but we aren't currently planning to do that.
Our model seems to perform well despite this "time generalization" split, but there are a couple instances where it struggled with new words. For example, we got the answer "FAUCI" wrong in a puzzle from May 2021. Even though Fauci was in the news before 2020, I guess he wasn't famous enough to show up in crosswords, and therefore his name wasn't in our training data.
I think evaluating performance by constructor would be really interesting! But we haven't done that.
It feels like the challenges might be that most clues are not cross-referential, and even for those that are, most information in the puzzle is irrelevant - you only care about one answer among many, so it could be difficult to learn to find the information you need.
But maybe this sort of thing would also be helpful for theme puzzles, where answers might be united by the theme even if their clues are not directly cross-referential, and could give enough signal to teach the model to look at the puzzle context?
One thing I was curious about - the ACPT is a crossword speed-solving competition, with time spent solving a major aspect of total score. How did you approach leveling the playing field between the human and computer competitors?