The Watson computer is good at playing Jeopardy, but it’s just fact collecting. Inferring a rule set from a corpus of documents, reasoning, and drawing solid conclusions, is a different game. We would be better off with a technology that would go over any given corpus of unstructured plain text, parse, tokenize, normalize, structurize, iterate, and dump the graph in a queryable data store. Natural language processing with a semantic reasoner. Legal documents are indeed a great use case.
I went to see my lawyer yesterday. Over the last three years I paid (i.e. was extorted) about $30k in legal fees. That’s lot of money for thin air. Withal, most of the research and paperwork I did myself: my lawyer copy/pasting and putting his court-accredited sign under it. To the legal scribes I’m but a layman, and may not speak on my own behalf in court (unless I want to jeopardize my cause). Despite the fact that I am particularly good at parsing texts, studied linguistics, did a post-doc in natural language processing, and know how to read laws (that are assumed to be known to and understood by all citizens, in the first place, are they not?). Fact is: lawyers, judges and clerks form a self-sustaining caste that benefits of the de facto (and de jure) monopoly of interpreting the law. They have a lucrative interest in laws and bills that are poorly written, are contradictory, full of ambiguity and logical flaws. They share that interest with retarded lawmakers who produce all that ill-conceived cruft.
Suppose we had our laws written in a formal language, with a well defined regular grammar; suppose we had unit testing for new bills — machines could administer justice, and they would be much better at it than the dunces who went to law school.
There’s simply too little stakes in developing semantic technologies that would do away with human corruption, underdeveloped intelligence, sentiment, and subjective interpretations. Or rather: some industries (especially those which are controlled by the powers that be) have too big a stake in preserving backward human knowledge parsing. If big law firms have an interest in using such technologies, they only have so, as long as they gain a competitive advantage from it (vis-à-vis those which don’t have/use such technologies). Legal corpora, neatly marked-up by hand (or semi-automated, no difference), are certainly a valuable asset, that you wouldn’t want everyone to have cheap access to, not your competitors, and especially not your litigant clients.
If we had instead semantic technologies that would cheaply produce queryable knowledge systems from large, impervious document collections (like our codes and jurisprudence), then, very likely, lots of sectors in our so called “knowledge society” would be disrupted, leaving lots of overpaid “knowledge workers” unemployed, overnight. That’s a threat to the very industries which are supposed to support research and development of such technologies, as potential customers.
That’s different with IBM’s customers, I guess: they do have an interest in disrupting their industries (or rather the industries in which they are newcomers), using tech. And they may do so, only because there’s no monopoly guaranteed by a selfish legal system. Maybe also because present day’s state-of-the-art in semantic technologies and AI offers good enough technologies for these use cases, which are less complicated applications than those that would be needed for use cases wherein more difficult knowledge parsing is required, _and_ are bound by a legal/economic anathema?
Anyhow, I will support any startup that would create such tech with the intention to run the human legalese interpreters out of business. And that’s an exhilarating thought, because if such technology would be produced, it will be equally good at solving problems in all branches of science.