English Letter Frequency Counts: Mayzner Revisited
norvig.com
norvig.com
* http://ktype.net/wiki/research:articles:progress_20110209#tw...
* http://ktype.net/wiki/research:articles:progress_20110228?s#...
I would love to redo this analysis with newer Tweets but alas, don't know where to get a usable corpus. Any suggestions? My goal was to explore many of the concepts from norvig's http://norvig.com/spell-correct.html using the Tweet dataset to build a better one-finger-keyboard and word-prediction engine for iOS.
So, if you wanted to tune your text prediction software for your phone...
It is actually much more difficult to model finger strain that the english language (in terms of n-grams). Subjective assessments vary a lot, and the quest for the optimal layout is bathed in controversy. Beyond switching away from qwerty, the most significant gains will be made by hardware solutions, like using an ergonomic keyboard (such as the very promising ErgoDox [2]). Other optimizations may come in the form of chorded keyers and better predictive technology. In this last case, the Google data may prove useful.
[2] http://deskthority.net/workshop-f7/split-ergonomic-keyboard-...
Redesigning a keyboard should be based on what people need to type, rather than on how English words are structured.
I believe it's beneficial to see a version that's based on the dictionary words alone as that would ensure no duplicate words exist to effect the n-grams, acting as a control group.
I don't trust his corpus compares to the original English corpus.
Also, at least 100,000 times isn’t so frequent, given Google’s huge corpus.
Here is the relative frequency of “Forschungsgemeinschaft” over time: http://books.google.com/ngrams/graph?content=Forschungsgemei...
Here are results on Google Books when searching for “Forschungsgemeinschaft”: http://www.google.com/search?q=%22Forschungsgemeinschaft%22&...
Google doesn’t promise that their English corpus only contains English books, just that it contains mostly English books. Given the raw size of their corpus this algorithmic solution seems like a reasonable tradeoff to me.
The frequency of “Forschungsgemeinschaft” is also easy enough to explain: The “Deutsche Forschungsgemeinschaft” (German Research Foundation) provides lots of funding for research in Germany. It is consequently mentioned in lots of research papers, many of which are written in English.
And seing that:
http://en.wikipedia.org/wiki/Deutsche_Forschungsgemeinschaft
is a research institute, one would expect tons of mentions for this word in _english_ scientific papers.
But that is an outlier, and given the vastness of the corpus those would have been sorted out.
That said, there would have been an easy way to filter non English books out of the way automatically: do a statistical analysis on each book (e.g on letter frequency) and reject the ones that stray too much from the norm -- or send them to a secondary filtering stage, e.g by word presence or a verification by a human. Done carefully that filtering would not harm the actual results at all (e.g by presuposing a specific letter frequency, because it would only reject extreme outliers that would indeed by non-english works).