1. Get some million tweets from twitter API (each tweet is exactly like what the article describes as a 'document')
2. After ignoring common words (e.g. 'a', 'an', 'the' ... ) in a tweet , start assigning some kind of relevance rank to all the other word pairs found in the tweet e.g. its likely that 'Cricket' and 'Sachin' (Or 'NBA' and <top NBA player) will both increase each other's relevance rank, WRT to each-other.
3. Process all the tweets like this, while maintaining the output of step 2 in a most suitable data structure. You would also need to start dropping word-pairs (to avoid having the problem of storing Million-C-2 words! ) based on some logic/heuristic.
4. If we have a good logic for having reasonably not-big storage (by avoiding the million-c-2 explosion), then what we have is at the end of processing: A simple look up of a million words, where each word has its 'top' max_allowed(k) semantic words.
PS: The most complex piece in the approach is to come out with a solution for dropping word-pairs (in step 3)