What are the "worst" spelling bee pangrams?
notes.billmill.org
notes.billmill.org
Source: helped build SB at NYT.
There's no actual answer to the question, given that the word list can be set inconsistently, so I had to choose _something_ to go off.
The best word list I was able to find is the one hosted by https://www.sbsolver.com/ , but unfortunately they don't distribute it.
I got a somewhat better word list for part 2: https://notes.billmill.org/blog/2024/03/mitzVah_-_the__worst...
But it's still imperfect. However a lot of the words I expected to be invalid have actually been in puzzles before, so it's not easy to guess which are going to be good and which aren't.
edit: there have been puzzles with as few as 16 words before: https://www.sbsolver.com/stats/count/low
edit 2: I modified the program to print puzzles with at least 16 words, and the "worst" puzzles it found with that constraint are:
unbEknown, jawbonE, monadnocK, woRkgRoup, daGlock, moonwalK, confLux, buLLhorn, yOkOzuna, Fraught, hogliKe
At least 5 of the proposed pangrams wouldn't make the cut, either.
Then you could take a very permissive wordlist and filter it using the historical data. For all words of six distinct letters or fewer, you could determine whether they were allowed, not allowed, or indeterminate (no puzzle ever appeared that would have allowed them). My gut feeling is that you'd be left with very few indeterminate words, though jouk and qajaq might well be among them - review those manually.
Makes me want to make a free clone that includes science words, and isn’t afraid of the letter S.
[1] https://www.oed.com/search/dictionary/?scope=Entries&q=Whing...
That's not to mention common phrases with a built-in redundancy such as "Cease and desist"? "Face mask"? "Free gift"?
Idiomatic English has lots of redundancies of one kind or another.
I can't think of a single use case for whinge that wouldn't be equally satisfied by whine. Can you?
English (as all modern languages) has tons and tons of exact synonyms and other types of redundant words. It's just a normal part of usage.
In this specific case I personally prefer whinge for emphasizing the complaint and whine for emphasising the noise, so I don't really think they are redundant - I think they are slightly different.
You can see 'whinge' gaining ground very recently at the expense of 'whine' here:
https://books.google.com/ngrams/graph?content=whine%2Cwhinge...
This is an outrage, and must be stopped. :-P
Wow. Another example of HN at its finest.
Great work on the game btw. My gf introduced me to the game and we love it. Though, we play a variation of it against one another in which we open the game on a single screen and whoever finds a pangram first wins.
The most annoying missing wordlist words are naphtha and caracal. An objective measure of word-use frequency should determine the words included. Probably super-obscure articles of clerical costume should not be.
> MORTIFY, FORTIFY, FIFTY, FORTY, MOTIF, FIRM, FOOT, FORM, FORT, FROM, IFFY, MIFF, RIFF, RIFT, ROOF, TIFF
The original Bees allowed only one vowel? That must have made it really tough to get long Bees!
I think the only listed words I'd think would get approved are jukebox, quixotic, and gimmickry.
It’s in beta right now and so I believe is accessible to everyone. Make sure to read the instructions.
I love puzzles but for some reason I’ve never been into word/dictionary games (Spelling Bee; Boggle; Scrabble) but since hearing about Strands via https://www.theatlantic.com/technology/archive/2024/03/stran... I’ve played every day.
Similarly I was fascinated by LetterBoxed, another NYT game and took a crack at a solver. https://hlfshell.ai/posts/letter-puzzles/
The NYT’s solutions are always two words, which I rarely get on my own. But once, a couple of years ago, I discovered a one-word solution to one day’s puzzle: LEXICOGRAPHY. Very elegant, I thought to myself smugly.
I happened to remember that solution a couple of months ago, and I decided to see if I could find others. I am not a programmer, but by asking ChatGPT 4 for help I was able to create a Python program and run it in Google Colab using a large list of English words that I had compiled from various online word lists.
Here is the beginning of the resulting list of one-word solutions to LetterBoxed:
acetylcholinesterase [a c e] [h i l] [n o r] [s t y]
acetylcholinesterases [a c e] [h i l] [n o r] [s t y]
achondroplastic [a c d] [h i l] [n o p] [r s t]
acknowledgement [a c d] [e g k] [l m n] [o t w]
acknowledgment [a c d] [e g k] [l m n] [o t w]
The code that ChatGPT 4 wrote for me and the full list of solutions are here:
https://gally.net/temp/20240318onewordsolutionstoletterboxed...
https://gally.net/temp/20240318onewordsolutionstoletterboxed...
But all I did was report the errors to the LLMs and paste their revised code back into Colab. The core logic of the search algorithm was created entirely by the LLMs based on my natural-language description of the LetterBoxed rules and the solutions I was looking for. I could not have written that code myself.
#include <array>
#include <algorithm>
#include <bitset>
#include <iostream>
#include <string_view>
#include <vector>
#include <unistd.h>
#include <sys/mman.h>
#include <sys/types.h>
#include <sys/stat.h>
#include <fcntl.h>
struct Word {
using Bits = std::bitset<32>; // A bit for each a..z, plus an "error" bit at index 26.
using Rules = std::array<Bits, 27>; // Map from letter indices to forbidden letters (plus a spare).
std::string_view str;
Bits bits;
static unsigned ix(char c) { return c - 'a'; }; // 'a'..'z' -> 0..25; others -> >25
bool ok() const { return !bits.test(26) && str.size() > 1; }
explicit Word(std::string_view s, Bits accept = {0x3ffffff}, Rules const& rules = {}) : str(s) {
for (unsigned i = 0, prev_ix = 26, bit_ix = 0; i != str.size(); prev_ix = bit_ix, ++i) {
bit_ix = std::min(ix(str[i]), 26u);
if (accept.test(bit_ix) && !rules.at(prev_ix).test(bit_ix)) { bits.set(bit_ix); }
else { bits.set(26); return; }}}
};
int main(int ac, char** av) {
auto usage = [av](){ std::cerr << "usage: " << av[0] << " abcdefghijkl [<wordlist>]\n"; };
if (ac != 2 && ac != 3) { return usage(), 1; }
char const* const name = (ac == 3) ? av[2] : "/usr/share/dict/american-english-large";
const int fd = ::open(name , O_RDONLY);
if (fd < 0) { return std::cerr << av[0] << ": failed to open " << name << '\n', 3; }
const std::size_t file_size = ::lseek(fd, 0, SEEK_END);
void const* const addr = ::mmap(nullptr, file_size, PROT_READ, MAP_SHARED, fd, 0);
if (addr == MAP_FAILED) { return std::cerr << av[0] << ": failed to open " << name << '\n', 3; }
const auto target = Word{av[1]};
if (target.bits.test(26) || target.str.size() != 12 || target.bits.count() != 12)
{ return usage(), 3; }
const auto rules = [w=target.str, ix=Word::ix](Word::Rules rules = {}) {
for (int i = 0; i < 12; i += 3)
rules.at(ix(w[i])) = rules.at(ix(w[i+1])) = rules.at(ix(w[i+2])) = Word(w.substr(i, 3)).bits;
return rules;
}();
const auto candidates = [&](std::array<std::vector<Word>, 26> candidates = {}) {
for (std::string_view in = {static_cast<char const*>(addr), file_size}; !in.empty();) {
const Word word{in.substr(0, in.find('\n')), target.bits, rules};
if (word.ok()) { candidates.at(Word::ix(word.str.front())).push_back(word); }
in = in.substr(word.str.size() + 1); // Skip past '\n' if present (might not be, at EOF).
}
return candidates;
}();
for (auto const& firsts : candidates) {
for (const auto first : firsts) {
for (const auto second : candidates.at(Word::ix(first.str.back()))) {
if ((first.bits | second.bits) == target.bits) {
std::cout << first.str << ' ' << second.str << '\n'; }}}}
}You need HTML IDs set on each, but they have to be different ID-anchors (obviously, and also because it'd be bad/invalid HTML to have duplicate IDs). So you actually need to track not just the other article but the ID inside the other article, and a way to generate those link IDs to begin with (since you definitely don't want to do it by hand). Gets tricky.
def calculate_backlinks(
pages: Dict[str, Page], attachments: Dict[str, Attachment]
) -> None:
for page in pages.values():
for link in page.links:
linked_page = find(pages, attachments, link)
if not linked_page:
info(f"unable to find link", link, page.title)
continue
linked_page.backlinks.append(page)
No deduplication performed. Since he's parsing the backlinks directly out of the markdown, you don't have to worry about a recursive loop where the backlinks section on one page appears as links in another. A simple solution would be to change the datatype of backlinks from list to set.[0] From OP's static site generator: https://github.com/llimllib/obsidian_notes/blob/main/run.py
Spenser coined it to use in the Faerie Queene in 1590:
His horses backe, yet to and fro long shooke,
And tottred like two towres, which through a tempest quooke
I confess that it's hard for me to get excited about solving puzzles to find obsolete nonce words.
So you might still be interested in the spelling bee, they mostly don't allow words like that.
(And thanks for the eytmology!)
Note: 'Nonce' is British slang for paedophile, though did you mean 'nonsense'?
Nonce is also used in cryptography for a number that is arbitrary and only used once: https://en.wikipedia.org/wiki/Cryptographic_nonce
Until the 1970s nonce meant "appears only once" and referred to figures and terms which never found common use after being coined.
grep -v '[^victmze]' /usr/share/dict/words |
grep .... | grep e
A very short program to generate all possible puzzles given a word list is at https://github.com/ncm/nytm-spelling-bee .When “ullage” was not in the list, but a week or two later “doggo” was, I considered giving up.
Fun game, but frustrating and annoying at times.
grep -v '[^victmze]' /usr/share/dict/words |
grep .... | grep e
A program to generate all possible NYTM SB puzzles is at https://github.com/ncm/nytm-spelling-bee . The alterations to match the online version are trivial. It runs in well under 100 ms on a modern CPU.There are bigger dictionaries packaged, e.g. wamerican-huge.
I just wrote two quick test programs to find the score of the set of words for every pangrams, and my approach takes 6s vs 15s for your proposed approach: https://gist.github.com/llimllib/cc01daa8be8ced13ddeb6c76cf1...
I’m almost certain that wasn’t true for at least one puzzle early this year, but haven’t been able to come up with an exact date.
https://www.nytimes.com/puzzles/spelling-bee -> click on 'How to Play'
> 1: 1642, 2: 364, 3: 90, 4: 24, 7: 2, 5: 12, 6: 5, 8: 1
That 8 panagram day was December 16, 2021