From the website, here are some example "documents":
roger heavens trinity cricketers willie sugg early cambridge cricket giles phillips cambridge university cricket
roger hornetttom pridmore 020 7399 4270 collins stewartalan geeves 020 7523 8800 buchanan communicationscharles ryland /isabel podda 020 7466 5000
roger hubbold aig .1 force-field navigation
Now, what is the point in trying to generate any kind of "meaning" from those documents if they consist of completely meaningless gibberish?As I was reading this challenge, I immediately thought of spam filtering / youtube comment classification ("smartness" classification) / etc as a potential useful application of this technology.
For example if each "document" is a youtube comment, then in theory you could write an algorithm to examine each comment and output a "smartness guess" for each. Then you (as in, you personally, by hand) would look at the results and specify your own "smartness rating" for a few comments. Then you'd run an algorithm to look at the difference between your specified "smartness rating" and the "smartness guess". Then, using that difference, it would tweak the settings in the original algorithm until it outputs a "smartness guess" that more closely fits your "smartness rating". If you repeat that process enough times, and your original algorithm has enough modifiers to tweak, then you might wind up with an algorithm that can make a pretty good guess about whether any given youtube comment is retarded or not.
And that's just one example of a practical application for this kind of thing.
That said, if the input "documents" are completely and utterly meaningless, then there does not seem to be any point in trying to build "meaning" from those inputs. (Garbage in, garbage out.)