Some words may also be difficult to pronounce/hear/spell by non-native speakers. Unlike regular sentences, there's no context to disambiguate.
Some words may also be difficult to pronounce/hear/spell by non-native speakers. Unlike regular sentences, there's no context to disambiguate.
Even for my attempt at the problem, I did various experiments on the word list, but an ideal attempt would check for similarity across common accents, etc and I certainly wasn't able to do that.
Having said that, I think it's a valid and realistic goal for good word encoder systems to aim for good roundtripability via voice or memory.
More seriously, English is such a terrible language for this, because it's so full of ambiguities.
English is no more prone to the problem of "some words sound exactly the same as other words" than any other language.
For the second point, -teen/-ty numbers sound different, are spelled differently, and mean different things. How are they supposed to support your point?
For the fourth, it's just false; "ghoti" in the pronunciation /fɪʃ/ does not come close to being valid written English. There is no such thing as syllable-initial "gh" /f/ or syllable-final "ti" /ʃ/.
The -ough suffix is a real case of one sound diverging into two sounds, but that is obviously not relevant to the problem of determining, from the sound of a word, which word you just heard. It comes up in the opposite problem of determining how to pronounce a word from the spelling, which we aren't talking about here.
The spelling bee is a cultural artifact; every language whose writing system is not extremely recent exhibits the phenomenon that the spelling of a word cannot be predicted from its sound. (In China, where spelling is much, much tougher, they don't have spelling bees. They do have traditional dictation exercises.)
You might find this wikipedia article interesting: https://es.wikipedia.org/wiki/Homofon%C3%ADa