This is such an over ambitious problem...
1. Webpages are ladden with errors, how do you deal with this? 2. Knowledge does not fit in a graph. It's asymptically a graph, as in: I can define relationships like: This recipe contains carrots. carrots contain sugar => this recipe contains sugar. Cool. But, what about this "sugar-free carrot cake recipe?" Well it still contains carrots, so still contains sugar... Contradiction? => requires human curating... 3. It doesn't even solve a real problem... Look at IBM watson, it probably knows a lot more crap than diffbot, and yet, is a pretty useless piece of software...