Soundex
en.wikipedia.org
en.wikipedia.org
One big place was on all Kraftfoods sites for search in recipes, products and brand sites. One use was for ingredient lookups from misspellings and search 2000-2008ish, still there at http://www.kraftrecipes.com/ on the search function. When you put in 'chiken' you'll get 'chicken' for instance. Pretty useful for misspellings back then and even today.
Fun fact we later also used Alta Vista search and even had a Google appliance, back when they made those, for aggregated site searching that tied into all search results across their brand sites. Search would check misspellings which part of that was SOUNDEX() then also aggregate search ingredients, recipes, products and content across their enterprise brand sites using the AV or Google boxes.
Another fun fact, kraftfoods sites were one of the first Microsoft .NET production uses. We worked with Microsoft in .NET 1.0 in 2001 to coincide with the release in 2002. We switched them from a combination of Perl sites and Java sites from Java / ATG Dynamo 10+ servers and 20+ Oracle servers to .NET with 3-4 web/app servers and 3 Microsoft SQL Servers.
As I'm now a full time Python developer, I guess I owe it to ATG Dynamo...
PS: Inspired by Metaphone, I wrote MLphone [2] a phonetic lib for the Malayalam (South India) language. The phonetic keys the algorithm produces are Roman characters though.
http://www.highprogrammer.com/alan/numbers/dl_us_shared.html
I do search related stuff and we use phonetic algorithms for names (in a rather interesting way as well which I haven't seen employed elsewhere) and will occasionally get reports or inquiries of weird unexpected matches, or questions about small typos not producing any of the expected results.
I feel these approaches maybe were a better fit for a time where talking was absolutely the main means of communications, but in an era where people communicate more and more by typing things into their phones, any input is frequently a) copied over instead of transcribed or b) first seen written and then typed out by the user, on a small touchscreen keyboard with 1 to 2 typos of letters close to the actual intended letter.
I wonder is there such an approach that takes this key distance into account? (ie. in a search for Nock results containing Nick should be higher than Neck)
Most likely, but keep in mind that one of the design goals of Soundex was that it be easy for a human to work out and that it be indexable. It was developed for the US Census at the start of the 20th century, after all…
Followups to soundex such as metaphone are encumbered by license issues as far as I know, but Caverphone is free and clear AFAIK.
[2] is an insanely great overview of many of these algorithms, be sure to check it out if you are into this stuff.
[0] https://caversham.otago.ac.nz/files/working/ctp150804.pdf
[1] https://gist.github.com/kastnerkyle/a697d4e762fa8f53c70eea7b...
[2] http://ntz-develop.blogspot.ca/2011/03/phonetic-algorithms.h...
http://bible.conman.org/kj/genasys.1:1
and it would redirect to the proper page: http://bible.conman.org/kj/Genesis.1:1
You have to really misspell something for it to not work properly.For some use-cases n-gram [3] based string comparison might be an option too. It is in no way phonetic (therefore universal for many languages), but when it is just about finding similar words it often produces better results than the original Soundex (mostly due to its length limitation).
[1]: https://en.wikipedia.org/wiki/Cologne_phonetics
But Python has since removed the soundex module.
1.6.1 declared it "obsolete": https://www.python.org/download/releases/1.6.1/
Obsolete Modules
...
soundex. (Skip Montanaro has a version in Python but it won't be included in the Python release.)
Looks like it was finally removed in 2.1: https://www.python.org/download/releases/2.1.3/notes/ matches the NEWS file in https://www.python.org/ftp/python/2.1/ - Removed the obsolete soundex module.
https://pypi.org/project/Fuzzy/ has a modern implementation if you want to play with it.Totally subjective, but in my domain I've had better use either using cheaper string distance/similarity metrics (hamming, jaro/winkler, etc), or if you're looking for some kind of resource-saving/fuzzy indexing/blocking type use, an application that uses or extracts ngrams has worked pretty well for me. Your mileage may vary...
[0] https://github.com/DJMelksham/SAS-Data-Linking-Functions
For those interested I'd highly recommend the work of Peter Christen [1], who does a ton of research in this space. If you want to see some code, check out the implementations of several of these algorithms I wrote a while back [2].
(Disclaimer: source code author here.)
Disclaimer: I'm the author of the tool.
Interesting. Isn't this similar to how Hebrew works (or at least the one used in the Bible worked)? I wonder about the rationale (in either case).