I did something similar with a 10m Tweet dataset couple of years ago:
* http://ktype.net/wiki/research:articles:progress_20110209#tw...
* http://ktype.net/wiki/research:articles:progress_20110228?s#...
I would love to redo this analysis with newer Tweets but alas, don't know where to get a usable corpus. Any suggestions? My goal was to explore many of the concepts from norvig's http://norvig.com/spell-correct.html using the Tweet dataset to build a better one-finger-keyboard and word-prediction engine for iOS.