How to Write a Spelling Corrector
norvig.com
norvig.com
http://news.ycombinator.com/item?id=42587
http://news.ycombinator.com/item?id=327897
Many great thoughts and comments already posted, so it's worth reading the thoguhts of HN contributors as well as this classic from Norvig.
http://googleresearch.blogspot.com/2006/08/all-our-n-gram-ar...
One of the painful things about commercial NLP work is lack of good datasets without restrictions. There are some treasures like WordNet (which is a lexical db). For tagged data I found the American National Corpus but the excerpt they release without restrictions is too little data and it contains only a few styles of writing.
I had to spend a lot of time gathering, fleecing, and marking up my own data to make After the Deadline.
Perhaps the UCI Machine Learning Repository would accept it. If not, I'm sure I could find a way to have it hosted at my university.
To preprocess wikipedia, I have used the following software: http://sourceforge.net/apps/mediawiki/wikiprep/index.php?tit...
To remove boilerplate from gutenberg requires painfully constructed heuristics. It would be great to have software released to do that.
- not have the luxury to have such a large ecological footprint (taking into account all the externalities too)
- are not always connected and
- are not granted access to the full UN corpus etc.
then you can still do quite good, cheaper and smarter.
http://hunspell.sourceforge.net/
Hunspell is the default spell checker of OpenOffice.org and Mozilla Firefox 3 & Thunderbird. Gőg hasn't beaten that yet.
Oh and AtD does grammar, style checking, and misused words. Embed it in your application today :) http://www.afterthedeadline.com
execution speed is more important than lines of code in this case