Yeah, it's because I'm trying to use a whitelist rather than a blacklist for the words, and although nltk (python lib) has a pretty good corpus, it's still missing a whole bunch.
The plan is to manually ok those words that people add and I don't have, and hope that this doesn't constitute the majority of the words :)
edit: Sorry to immediately disappoint though.