As for turning current contests into an exercise in overfitting the "test" set - we already reached that point long ago. Test vs validation scores often diverge wildly in these contests.
Edit - Replying to arnsholt:
Completely true. The huge problem I see, is that all the classic NLP tagging corpuses are created from the very narrow domain of news articles, and a few good corpuses now appearing for biology texts, and that's about it. Want to do, e.g. NER for product reviews or chat logs? - Incredibly bad results. There's a huge corpus problem in NLP today.