Regarding finding similar documents what is the state of the art nowadays, LDA, word2vec, something else? What do you normally use?
Some examples of use-cases: are you searching for "semantically similar", or "near duplicate"? You can compare documents under different metrics and different _representations_. Some representations are: LSA, PLSA, LDA, TF-IDF, and Set representations, along with metrics such as Jaccard Distance, Cosine Distance, Euclidean distance, etc.
Doc2vec is the Word2vec analog for documents.
There is an implementation in Textacy.