http://radimrehurek.com/gensim/
It has good implementations of various algorithms, some of which support streaming or dirstribution, and it allows loading and dumping data in various formats.
I've used it for building content based recommender using tf-idf, lsi and similarity index. After the index is built, queries to it are really fast. It can handle quite large corpuses with little memory.