At first look it doesn't pass the search engine test. https://duckduckgo.com/?q=I%27ve+bien+aloune+whyth+jou+eensi...
At first look it doesn't pass the search engine test. https://duckduckgo.com/?q=I%27ve+bien+aloune+whyth+jou+eensi...
https://duckduckgo.com/?q=eyeve+bien+aloune+whyth+jou+eensid...
Set region to Austria, and Ritchie is result nr 1. set Region to "All Regions" and he's result nr 10.
e.g. replacing one letter with three that sound similar but are different adds 4 to the distance - which is quite a significant distance already.
What if you want to use it in the opposite direction? Let's call it the CIA/NSA direction, where you know what unobfuscated phrase you want to find, but not how someone may have obfuscated it. This is arguably much harder, especially if you do not store some sort of representation of the content "as sounds/phonemes".
Something I've had in the back of my mind for a while is the idea of swapping out a regular Lucene tokenisation & analysis for one that treats phonemes as tokens instead of (stemmed/analysed) words being the tokens... with a similar arrangement on the query side. I think it'd have interesting capabilities... this being one.
What's the advantage to this?