Relying on vector similarity as opposed to direct keyword matches is an approach that's taken. Utilizing statistical and probabilistic approaches as opposed to "words and rules" remains important.
See: http://cymetica.com/cymetica_about.html
and the following: http://genopharmix.com/biomimetic-cognition/in_silico_cognit...
The corpus is currently just limited Reuters public company profiles & descriptions but I plan to include SEC filings as well.
Traders, investors, hedge funds can engage is quick information arbitrage with it. For example, a stock runs up 20% in minute, you insert a keyword or the symbol related to the stock that ran up, you then get other stocks (a targeted basket) that have sympathetic, symbiotic and parasitic relationships before any research analyst can uncover the connections - rising tide lifts all boats or a lowering tide lowers them. Ref: "Contagious Speculation and a Cure for Cancer: A Non-Event that Made Stock Prices Soar" -http://www0.gsb.columbia.edu/whoswho/getpub.cfm?pub=1555
The system can include tiered consumer/trader/investor subscriptions, licensing and net profit sharing with selected hedge funds. SeekingAlpha, StockTwits, Yahoo Finance etc as revenue partners. We're raising a bit of funding in the meantime while we also use it as a trading tool ourselves.
It's a prototype that's being moved into production this week or next.