I don't know the python libraries off the top of my head, but the Python implementation of R (RPy) is decent and R is heavily used in most circles to have great literature written for it.
Two good first steps to look into, depending on your needs, are Bayesian classifiers and SVD (reduction of high dimensionality, the application to text processing was patented as Latent Semantic Indexing/Analysis, LSI or LSA, by IBM, I don't knnow if that's lapsed).
Basically, you need a reasonable feature to match similarity on. N-words are pretty easy to construct, a 2-gram would be every pair of words used in a document.
Tf-idf is a good metric with that kind of feature, because it handles well the bias of frequent words like "the"