I specialize in word representations and using them for various tasks. Word Similarity prediction has been used as a basic first evaluation for many years now, with analogies becoming an additional standard task in the past couple of years. But it's worth noting that word representations have a LOT of open parameters (which model? how many dimensions? do I remove stopwords and low frequency words prior? do I use a bag-of-words context or a syntactic context?).
The optimal parameter choices for one task are very frequently not the optimal parameters for another. While there are usually "reasonable defaults" for when you don't want to optimize everything, a solid standardized approach risks vastly overfitting to one task, possibly at the expense of more useful tasks.