http://www.alexkrupp.com/Citevault.html
Basically the economics are insanely good if you're using this as an open source tool to create NYT-style articles, less so if you're Circa.
The people comparing this to the semantic web don't understand Zipf's Law, and just how slowly things actually change. E.g. we only get new data on adult literacy every 10 years. And the last data we have on antibiotic resistance for some bacteria/drugs is from the early 90s, and that's more the rule than the exception. Pretty much every single article about the U.S. is using the same set of a couple thousand facts, and most of those only get updated every ten years or so. There are a few exceptions like with federal arrest data that gets updated yearly, but that's pretty rare.
Basically if you're trying to do this using machine learning or any sort of algorithms, you're completely wasting your time and going down the wrong path. This is way easier to implement well than you think. (But again, not necessarily super profitable unless you own the NYT.)