I found the best anagram in English (2017)
blog.plover.com
blog.plover.com
misrelation / orientalism; superintended / unpredestined; incorporate / procreation (don't mind if i do!); predators / teardrops (a cause and effect); counteridea / reeducation (a bit synonymous); streamlined / derailments (quite opposite!); truculent / unclutter; colonialist / oscillation; renavigate / vegetarian; persistent / prettiness; paternoster / penetrators (hmm); obscurantist / subtractions; nectarines / transience (a story of ripeness); definability / identifiably; indiscreet / iridescent; excitation / intoxicate; discounter / reductions (how logical!)
One small suggestion I have: add a point for pairs with different starting letters, and another point for pairs with different ending letters.
... and I don't use it, because unjumbling the word myself is satisfying, but typing the letters into a computer and getting the answer isn't.
how does that work?
Then go through a word and find the values for each letter and multiply them together, e.g. "tab" is 20 x 1 x 2 = 40 and hopefully an anagram that just rearranges the letters gets the same answer because multiplication doesn't change if you shuffle the numbers around, e.g. "bat" is 2 x 1 x 20 = 40 which is the same, "bat" and "tab" are anagrams... but with the integers it doesn't always work and different words can clash e.g. "fab" 6 x 1 x 2 = 12 and "cad" 3 x 1 x 4 = 12 have the same answer but are not anagrams.
Prime numbers help because the Fundamental Theorem of Arithmetic[1][2] says that there can't be any clashes when you multiply Primes, every number breaks down into a unique product of Primes (I can't prove that myself, but it is apparently true). So give the letters Prime numbers A=2, B=3, C=5, D=7, E=11, F=13, etc. and now "fab" 13 x 2 x 3 = 78 and "cad" 5 x 2 x 7 = 70 no longer clash. The only way to get the same answer is to have the same primes (in any order), so anagrams will have the same answer and non-anagrams will not.
Why it drops the overhead of sorting is that the time for sorting any collection requires looking at each item and comparing at least some of them, and swapping positions of at least some of them, generally O(N items x log(N)). Lookup the letter in a Prime value array and multiplication once per letter doesn't need any comparisons or any swapping positions, so it is O(N items) time, that gives this approach less work to do for each word, so it can finish faster.
It looks like (Python, assuming lowercase ASCII letters where 'a' starts at code 97):
primes = [2,3,5,7,...]
ascii_a = 97
product = 1
for c in word:
product *= primes[ord(c) - ascii_a]
Do that for the incoming word, and for every word in the wordlist, and see which have matching products, those are the anagrams. Or pre-compute for all the words in the wordlist and only do it for the incoming word and then lookup the matching ones.[1] https://en.wikipedia.org/wiki/Fundamental_theorem_of_arithme...
[2] https://www.varsitytutors.com/hotmath/hotmath_help/topics/pr...
And the technique of Prime Combinatorics for an "alphabet" can be used to solve Poker hands, Blackjack hands, Match 3 puzzle games, slot machine reel positions, and a slew of other similar problems where you would ordinarily have to build a very complex logic table/switch-case/if-then-else decision tree.
You could just do counting sort if the big-O is important, but I'm a bit suspicious about that big-O anyway. A model where multiplying arbitrarily big numbers is constant time is a bit unrealistic, kind of feels like it's getting off the hook for a log factor for free.
It's also all such small values that I'm not sure the big-O matters either, I'm not confident which would win without just trying it. I'd _guess_ that usual sort probably just wins though, or counting sort if you really went out of the way to optimize it.
Looking what was happening, I find C# doesn't throw overflow exceptions by default and instead wraps around. That takes 0.05 seconds for all 172k words compared to 0.1 seconds for sorting all words so it ran a lot faster, yes, but it was wrong and may have clashed words. Same with Rust and overflows, it seems. Turning it to System.Numerics.BigInteger in C# makes it take about the same time 0.1s for both. In Python which handles any size integers it takes a lot more time, 0.14 seconds for sorting, 0.4 seconds for products.
That means my embedded arrays of products are all overflowed as well. What a good example for arguing that programs should be correct first, and fast second. (I remember being suspicious of overflows, but I don't remember what I tried to conclude that all the word results would fit into int64, but it must have been bad).
(I'm curious,if it would be possible to shuffle the primes around to bring all words in my full wordlist down to UInt64; the highest is 2810298024552111657086849344270208275 - microspectrophotometries which is some quadrillion times too big. Rearranging the by letter frequency in the wordlist brings it down to 24474928756352445162715195352080 - electroencephalographically which is still a trillion times too big, so I doubt it).
I proudly boasted to a friend about my winning Boggle solver and they said it was the pettiest thing they had ever heard of.
Saddam Hussein = He damns Saudis
Charles Manson = Slasher con man
David Letterman = Dead mitral vent
Mary Jo Kopechne = My joke chaperon *
Benito Mussolini = So, I bout Leninism
Lee Harvey Oswald = Oe, why ever Dallas? *
* "Chaperon" is a valid alternate spelling of "chaperone"
** Yes, "oe" is a word
> I guess I’ll be known for nothing more than being the man who realized that “Spiro Agnew” was “grow a penis.” Gore Vidal said, “It could be ‘grow a spine,’ too, but yours is better.”
Boy, they're really socking it to that Spiro Agnew guy again.
I also like the mathematically correct
ELEVEN PLUS TWO = TWELVE PLUS ONE
especially because it's also an numeric anagram
11 + 2 = 12 + 1
Goedel's incompleteness theorem fans would be drooling all over this
Some of those are almost creepy in how they make so much sense.
Probably because looking at evolution it's more helpful to see a predator that isn't there, than to miss a predator that is there. It intuitively makes sense that we would tilt towards false positives over false negatives in our perception. Since one has a higher cost attached than the other.
Which is why I find some of the discussion arround whether GPTs are intelligent interesting. I see the argument quite often that they merely match patterns and combine things. Which to me seems very much akin to us. They lack the ability to interact with the world and there is no feedback loop as of now, but it seems to me that there is something very human to the AI we programmed.
Atlantic casino resort spa / Carter assassination plot
You can find the documentation, a Win32 executable, and the source code here: https://www.kmoser.com/anagrams/
In terms of this ‘chunk scoring’ method this scores very high (15 I think?), which definitely confirms its value as a way of rating anagram quality.
Similar to this, he produces a standard form for each word, but breaks each letter into letter pieces or 'atoms' which gives much more freedom for moving between words.
Definitely give it a watch. If you are not familiar with Tom7's videos, he has a hilarious whimsical style while also bringing to life completely out there ideas with some brilliant technical skill.
human parties | puritan shame> 7 admirer married > 7 admires sidearm
> This was easy to do, even at the time, when the word list itself, at 2.5 megabytes, was a file of significant size. Perl and its cousins were not yet common; in those days I used Awk. But the task is not very different in any reasonable language:
You wrote "Prior to Perl doing this was painful.".
I highlighted how you contradicted the author's claim that it was easy to do using awk. Awk existed long before Perl.
I tried to avoid modern awk features, like asort, to be something that would have worked in the 1980s:
{
# Convert to normal form:
# 1. Fold to lower case
# 2. Bin the letters to get frequency counts
# 3. Only consider lower case ASCII letters
split(tolower($0), letters, "");
# Ignore asort() in modern awks and do a bin sort instead.
for (i in letters) {
c = letters[i];
repeats[c] = repeats[c] c;
}
normal_form = "";
for (i=97; i<=122; i++) {
c = sprintf("%c", i); # no chr() in a 1980s awk
normal_form = normal_form repeats[c];
}
table[normal_form] = table[normal_form] "," $0
delete repeats;
}
END {
# Only show the ones with at least 5 matches
for (i in table) {
match_str = table[i];
split(match_str, matches, ",");
num_matches = length(matches)-1
if (num_matches >= 5) {
# print the number of matches, then the match string
printf("%d%s\n", num_matches, match_str);
}
}
}
When I try it on a word list I have handy, here are the most common words: % awk -f anagram.awk < words_alpha.txt | sort -n -t, | tail -5
13,elaps,lapse,leaps,lepas,pales,peals,pleas,salep,saple,sepal,slape,spale,speal
14,anestri,antsier,asterin,eranist,nastier,ratines,resiant,restain,retains,retinas,retsina,stainer,starnie,stearin
14,apers,apres,asper,pares,parse,pears,prase,presa,rapes,reaps,repas,spaer,spare,spear
14,arest,aster,astre,rates,reast,resat,serta,stare,strae,tares,tarse,tears,teras,treas
15,alerts,alters,artels,estral,laster,lastre,rastle,ratels,relast,resalt,salter,slater,staler,stelar,talers
Certainly Perl is more succinct, though note that even up to Perl 4 in the early 1990s you would need to use the string concatenation method to store the list of matches in the table.But, "painful"? No. Not to someone who knew how to use awk.
The simplest way to do it, is to convert all the words to (normal_form, orig_word) pairs, write the list to a file, then sort it.
It will be trivial to find the words with common normal form after the sort.
(Of course, you wouldn't catch me trying to implement that with C if perl is an option...)
No hash table needed, just splitting the line into the two fields, equality comparison, and appending values to a list.
https://news.ycombinator.com/item?id=13696196
Response by the author:
> Productive diminutives are infrequent to nonexistent in Standard English in comparison with many other languages.
https://en.m.wikipedia.org/wiki/List_of_diminutives_by_langu...
This is the TXR Lisp interactive listener of TXR 285.
Quit with :quit or Ctrl-D on an empty line. Ctrl-X ? for cheatsheet.
TXR's sound system features 120 dB separation between quarreling audiophiles.
1> (flow "/usr/share/dict/words"
file-get-lines
(group-by sort)
hash-values
(keep-if cdr)
(sort @1 : [chain car len]))
(("ho" "oh") ("am" "ma") ("em" "me") ("no" "on") ("ah" "ha") ("it" "ti")
("mu" "um") ("eh" "he") ("ate" "eat" "eta" "tea") ("bar" "bra")
[...]
("certification" "rectification") ("accouterments" "accoutrements")
("peripatetic's" "precipitate's") ("amphitheaters" "amphitheatres")
("colonialist's" "oscillation's") ("enumeration's" "mountaineer's")
("imperfections" "perfectionism") ("antiparticles" "paternalistic")
("broadcaster's" "rebroadcast's") ("impressiveness" "permissiveness")
("conservation's" "conversation's") ("tablespoonfuls" "tablespoonsful")
("certifications" "rectifications") ("amphitheater's" "amphitheatre's")
("certification's" "rectification's") ("impressiveness's" "permissiveness's"))(⊢⌷⍨∘⍋∘≢¨)↑1<≢¨⊢⌸(⍋⊢)¨⎕NGET'/usr/share/dict/words'1
Why shouldn’t my programs be intense neutron stars of weird symbols, if there’s a superhuman intelligence always at hand to explain and improve the code?
With the new AI, you just specify the "how". Then you get a buggy program, in one of the above specialized languages for the old AI, and the rest of the prompts in the chat are edit instructions on how the AI should fix the code to make it work.
It's idiomatic to pick ⊃ the first result of ⎕NGET to get just the lines to work on, and not the other things like file encoding.
Then grade-up-right-train-each (⍋⊢)¨ is redunant, it's the same as grade-up-each ⍋¨
The grade is an array of which indices to take to put the argument in sorted order and I don't think it makes sense to group ⌸ by that since the grade isn't the same for different arrangements of letters, so the whole approach breaks down there. I think the words have to be sorted, e.g to pick out 'bat' and 'tab' as the same letters:
{⍺, ≢⍵}⌸{⍵[⍋⍵]}¨words ← 'bat' 'dog' 'tab' 'cow' 'wok'
abt 2
dgo 1
...
Then 1<≢¨ would be "1 is less than the count (tally) of each" which fits somewhere in the solution, but not there and won't work on my array. We don't need to know the actual sorted letters so we can inline the bitmask of which answers matter or not in the first column: {(1<≢⍵),⍵}⌸{⍵[⍋⍵]}¨words ← 'bat' 'dog' 'tab' 'cow' 'woka'
┌→────┐
↓1 1 3│
│0 2 0│
│0 4 0│
│0 5 0│
└~────┘
Any with a 1 in the first column are anagrams, and 0 in the first column are not. And the other columns are indices into the wordlist where the matching words are, so words[1,3] picks out 'bat' and 'tab' and it's probably possible to filter rows with 1 in the first column: {⍵[;1]⌿⍵}{(1<≢⍵),⍵}⌸{⍵[⍋⍵]}¨words←'bat' 'dog' 'tab' 'cow' 'bta' 'racecar' 'carrace'
and then drop the first column: 1↓[2]{⍵[;1]⌿⍵}{(1<≢⍵),⍵}⌸{⍵[⍋⍵]}¨words
Then uhh filter out the zeros to avoid index error, and index into the word list, so: {words[⍵/⍨⍵>0]}¨⊂[2]1↓[2]{⍵[;1]⌿⍵}{(1<≢⍵),⍵}⌸{⍵[⍋⍵]}¨words←⊃⎕NGET 'wordlist.txt' 1
┌→───────────────────────────────────────────────────
│ ┌→────────────┐ ┌→────────────────┐ ┌→────────────┐
│ │ ┌→──┐ ┌→──┐ │ │ ┌→────┐ ┌→────┐ │ │ ┌→──┐ ┌→──┐ │
│ │ │aah│ │aha│ │ │ │aahed│ │ahead│ │ │ │aal│ │ala│ │ [...]
│ │ └───┘ └───┘ │ │ └─────┘ └─────┘ │ │ └───┘ └───┘ │
│ └∊────────────┘ └∊────────────────┘ └∊────────────┘
└∊───────────────────────────────────────────────────
It can probably be done shorter and cleaner with more skill than I have. It's a lot quicker to execute than the PowerShell version I commented, but took a lot longer to code.(That expands to {⍵[⍋⍵]}¨words to sort each word, then use that on the left of Key ⌸ with words on the right, and Key feeds the count and indices into the custom function which is ⊢∘⊂ that takes the indicies with right-tack ⊢ and throws away the count, and encloses them with ⊂. That gives nested arrays of words which sort the same, including invididual words that sort like nothing else. Then (⊢⊢⍤/⍨1<≢¨) added to the left is counting the words in each nesting and filtering out the single ones, and I think ⊢⍤/ is a bodge to use compress in a train without it being misread as reduce when both use the same symbol / ?)
The : symbol causes argument 2 of sort to be defaulted, as if it were omitted. That's the comparison function, which defaults to less. In TXR Lisp, you can explicitly default optional arguments with : which enables you to give arguments to later optional arguments to the right of those.
We then specify argument 3, the function for selecting the key to sort by for each element. Each element is a list of words (anagrams fo equal length). We pick the first one with car, and take its length: chain will compose those functions for us.
We use square brackets because the function names are in the function namespace: square brackets provide a function application language in which variables and functions are in one namespace.
PS C:\> get-content wordlist.txt
| group { -join ([char[]]$_ | sort) }
| where count -ge 2
| sort {$_.Name.Length}
Count Name Group
----- ---- -----
2 ho {ho, oh}
2 do {do, od}
2 ay {ay, ya}
[...]
2 aeghhiiloooppssty {pathophysiologies, physiopathologies}
2 aacghhiilloooppsty {pathophysiological, physiopathological}
2 aceghhiimooopprrst {microphotographies, photomicrographies}One of the top players in my club _instantly_ replied "cinematographer," and added, "but megachiropteran isn't in the OSPD [Official Scrabble Players' Dictionary], so it doesn't really count."
What's the longest word ever played in an actual game of Scrabble? And though this is hard to get data for, what's the longest word that ever could have been played but missed?
I believe that the gentleman in this case instantly saw the word 'cinematographer' hiding in 'megachiropteran' because his brain is a highly trained anagramming machine, and knew that it wasn't allowed because he knew all of the allowed 15-letter words, just in case he might get the opportunity to play one.
High-level players memorize staggering numbers of words. All 2- and 3- letter words is entry-level. All X-J-Q-Z, all Q-without-U, all words you ending in -MAN, all 70+ 7-letter words of the form SATINE+... top-level players know the 4000 4-letter words, the 5000 5's, and way, way more.
This (https://www.yahoo.com/lifestyle/scrabble-records-highest-sco...) lists cases in which 15-letter words were played by adding prefixes or suffixes to other words.
And here are the top scorers:
11 actinometro / cortamiento
11 aeronáutico / ecuatoriano
11 anemometría / mareamiento
11 apriscadero / esparcidora
11 arremetimiento / meritoriamente
11 atrasamiento / metatarsiano
11 camastronería / sacramentario
11 entropezada / panderetazo
11 importación / piromántico
12 anemométrica / maceramiento
ChatGPT
astronomers = moon starers = no more stars
My full first, middle and last name is long enough with common letters that there are several uproariously serendipitous and embarrassingly obscene (even for me) anagrams of my full name.
So bad I would never post them here. Much worse than you could possibly imagine. Take my word for it: you don't want to know.
The only advice I'll share is that parents should carefully screen their baby names with the advanced anagram server, and choose short names with unusual letters that have lower chances of backfiring.
https://github.com/dfhoughton/ranagrams
I find it a good time killer. Note, there's a word list in the repo, but `crate install ranagrams` starts you off without a word list.
Qui suis-je ?
Me voici Sultan !
Nuitisme vocal
Silice mouvant
Io, vent musical !
Le voici musant
Mot inclus à vie
Ce motival insu
Là vous émincit
Cultivons amie
Si il vaut ce nom
Son val muet ici
Indice:
Si nul vice à mot
Vu ici slame ton
Nom c'est via lui
Vaincu tel, omis
Who am I?
Here I am, Sultan!
Vocal nightism
Moving silica
Io, musical wind!
Here it is, musing
Word included for life
This unknown motive
There, it slices you down
Let's cultivate, friend
If it's worth that name
Its silent valley here
Hint:
If no word vice
Seen here, slams your
Name, it's through it
Defeated as such, omittedI was born in 1982, and a lot of the humor went over my head til much later. But the first ten seasons or so are still indelibly etched in my memory.
cholecystoduodenostomy / duodenocholecystostomy
https://gist.github.com/anonymous/431b163b2a2d532bfd0a3bdcc7...
See https://blog.plover.com/lang/anagram-scoring-3.html for more details.
Now that's something...
11 + 2 = 12 + 1
Eleven plus two = twelve plus one.
Eyeballed the entire list*. This is the most 2023 one I could find.
(*no)
http://www.fun-with-words.com/palin_panama.html
I've always been partial to pangrams: sentences that use every letter in the alphabet at least once, typically shooting for a short sentence.
My favorite is "Pack my box with five dozen liquor jugs". Not minimal, but delightful and uses only words that most mortals know.
Yes, this Guy Steele: https://en.wikipedia.org/wiki/Guy_L._Steele_Jr.
Old West Action
Rewrite all words so the letters in each word are rearranged alphabetically. Don't do this other thing suggested on StackExchange
Look at the list of anagrams
Recognize that the list is actually boring
Get silly with words and nerd out
The end!
cinematographer --> megachiropteranI wonder if the author realizes how condescending this is.
All those thoughts about what a stupid question it is, and how stupid someone must have been to ask it—all those are thoughts you had, not me.
Given the tone of the rest of the article I assumed that I'd either read something into it that wasn't there or was missing some context - for example, perhaps the person on SO had insisted that this was the correct approach, and you'd gone to some lengths to show that it was not and were exasperated by it by the time you wrote the article.
At any rate - it didn't really take away from your article for me, but I did see it as condescending.