1) Do words in a generic corpus (such as Wikipedia) actually form well-separated clusters?
2) Is it correct that you find word clusters in the corpus as a preprocessing step (as opposed to at indexing or query time)?
3) Do I understand correctly that you use all words in clusters as synonyms and pass them to Solr at query/indexing time? Is it query time, index time or both?
4) Given a language where words have many syntactic forms (e.g. buy-bought-buying), how does it work with clusters? Do both syntactic forms and synonyms end up in the same cluster? Wouldn't it be beneficial to treat many of these different forms as the same word (i.e. perform stemming) and only list truly different, but closely related concepts as synonyms?