A way to process your training data to fit your knowledge graph needs to be unambiguous. No sense it teaching it anything if it becomes full of contradictions.
Finding a dataset for it to learn would be tough as well. Given the breadth of possible questions, you'll need to parse huge encyclopedias and/or wikipedia.
Finally, a way to efficiently query this massive amount of data. If it has to come up with an answer faster than its competitors, it better be able to lookup information pretty damn fast.
I'm going to start with information retreival book I found at stanford but something tells me I need much more than just that!
http://memesteading.com/2011/02/16/ibm-watson-overprovisione...
Apparently the open-source Apache UIMA and Hadoop projects are key parts of Watson's preprocessing and live operation:
https://blogs.apache.org/foundation/entry/apache_innovation_...