Soundex – a phonetic algorithm for indexing names by sound
en.wikipedia.org
en.wikipedia.org
- Recovering English spelling from English pronunciation. Soundex and Metaphone are good at this.
- Undoing the torments of Eastern-European sibilants in various transliterations of surnames, e.g. Tchaikovsky, Tschaikowski, Tchaïkovski, Ciajkovskij, Tsjaikovski, Tjajkovskij, Tsjajkovskij, Csajkovszkij, and Czajkowski. Here, Double Metaphone or Daitch–Mokotoff Soundex do a better job.
[1] https://en.wikipedia.org/wiki/Metaphone [2] https://en.wikipedia.org/wiki/Daitch%E2%80%93Mokotoff_Sounde...
Solr, for example, supports quite a few phonetic matching algorithms by default[1].
[1]: https://lucene.apache.org/solr/guide/7_4/phonetic-matching.h...
I was curious about whether or not there was an IPA version (which would then work for any language) -- one of the first results I found was Eudex, which looks a little different (since it's a hash of the pronunciation) but more versatile.
$ espeak -q --ipa -v fr foie
fwˈaI found this quite interesting (and comprehensive) project that includes regular string distance (fuzzy search) methods, fingerprinting methods and a huge number of phonetic metrics including 7 different soundex methods: https://github.com/chrislit/abydos
A few others worth mentioning:
Scala: https://github.com/rockymadden/stringmetric
Java/Scala: https://github.com/vickumar1981/stringdistance
Postgres: https://github.com/eulerto/pg_similarity
After that I build a search engine programme that would look up swear words that were phonetically similar to input terms, as I thought that my fellow students might get a laugh looking up their own names. In the end, neither Phonix, double metaphone nor Soundex really produced any funny results.
plugs:
- Blogpost: http://olsgaard.dk/phonixpy-phonetic-name-search-in-python.h...
- Github repo: https://github.com/olsgaard/phonetic_search
[1] Gadd, T. N. “‘Fisching Fore Werds’: Phonetic Retrieval of Written Text in Information Systems.” Program 22, no. 3 (1988): 222–37.
In a similar vein, Levenshtein distance is also built-in which is very useful for fuzzy searching https://www.php.net/manual/en/function.levenshtein.php
But Harbour, an open-source implementation of Clipper superseded it with Metaphone - https://harbour.github.io/doc/hbnf.html#ft_metaph
https://www.neilvandyke.org/racket/soundex/#%28def._%28%28li...
[1] https://www.community-insurance.com/learn-how-to-read-your-d...
It does make it harder to fake an id, perhaps that was the motivation.
Edit: Some more interesting info for the states that implement this. Also, TIL Illinois drivers can appear to have the exact same DL number. http://www.highprogrammer.com/alan/numbers/dl_us_shared.html
> Soundex is a phonetic algorithm for indexing names by sound, as pronounced in English.
[0] https://en.m.wikipedia.org/wiki/List_of_dialects_of_English