1,595 karma · joined August 13, 2009
The training data for machine translation models is also human-created. Given some fixed amount of human hours, would you rather them be spent annotating text that can train a translation system that can be used for many things, or a system that can just be used for this project? It all depends on the yield that you get per man-hour.
First of all, having a machine-readable database of knowledge(i.e. Wikidata) is no doubt a great thing. It's maintained by a large community of human curators and always growing. However, generating actually useful natural language that rivals the value you get from reading a Wikipedia page from an abstract representation is problematic.
If you look at the walkthrough for how this would work (https://github.com/google/abstracttext/blob/master/eneyj/doc...), this project does not use machine and uses CFG-like production rules to generate natural sentences. Works great for generating toy sentences like "X is a Y".
However, human languages are not programming languages. Many natural languages, like German and Finnish, are so syntactically and morphologically complex that there is no compact ruleset that can describe them. (those that have taken grammar class can relate to the number of exceptions to the ruleset)
Additionally, not every sentence in a typical Wikipedia article can be easily represented in a machine-readable factual format. Plenty of text is opinion, subjective, or describes notions that don't have an proper entity. Of course there are ways that engineer around this, however they will exponential grow the complexity of your ontology, number of properties, and make for a terrible user experience for the annotators.
A much better and direct approach to the stated intention of making the knowledge accessible to more readers is to advance the state of machine translation, which would capture nuance and non-facts present in the original article. Additionally, exploring ML-based ways of NL generation from the dataset this will produce will have academic impact.
As a contrast, take a look at how Diffbot, an AI system that understands the content of the page by using computer vision and NLP techniques on it, interprets the page in question:
https://www.diffbot.com/testdrive/?url=https://www.reddit.co...
It can reliably extract the publication date on each post, without resorting to using site-specific rules. (You can try it on other discussion threads and article pages, that have a visible publication date).
NLP is good enough that we can now explicitly measure how well a system reads text in terms of what knowledge is extracted from it. This task is called Knowledge Base Population, and we've released the first reproducible dataset called KnowledgeNet that measures this task, along with an open source state-of-the-art baseline.
Direct link to the Github repo: https://github.com/diffbot/knowledge-net EMNLP paper: https://www.aclweb.org/anthology/D19-1069.pdf
In the context of that sentence "She was a .. physicist and chemist who conducted .. research on radioactivity.", I think most people would say a physicist or a chemist who conducts research is a researcher. In other contexts, such as in your example, that would be a questionable inference. What you're describing is why natural language understanding is hard--it's context-dependent and not syntactic.
"She did research." http://relex.diffbot.com:8085/?text=She%20did%20research. "She was a researcher." http://relex.diffbot.com:8085/?text=She%20was%20a%20research....
Return different outputs from a state-of-the-art relation extraction system.
The initial example: http://relex.diffbot.com:8085/?text=Marie%20Curie%20was%20bo....
In your example, being a coder is not inferred: http://relex.diffbot.com:8085/?text=He%27s%20a%20hard-surfac....
http://relex.diffbot.com:8085/?text=I%20live%20in%20Japan.%2....
We think that the main way to achieve a practical semantic web is to have AI synthesize a Knowledge Graph from applying CV/NLP techniques to understanding all webpages. More about our project here:
https://www.zdnet.com/article/the-web-as-a-database-the-bigg...
Obvious one is scale, Wikipedia has on order 10M entities and represents the work of thousands of humans whereas the Diffbot KG has 10B entities and is discovering about 120M each day, and is largely limited by the number of machines running the algorithms in the datacenter. The properties and facts indexed about each entity are also a superset because it is not limited to those that would be worthwhile for a human to curate. Lastly, it can be more accurate than facts found in a single source because the automated system utilizes multiple sources of that fact found across the web to estimate a probability of the accuracy of the fact.
The result is that you have a Knowledge Graph that is more useful for work and business because they are the entities you interact with day to day, not the "head" entities that optimize for popularity and the constraints of human curation.
As you point out, the concept "sugar-free" does not equate to there being no "sugar" in the dish. Perhaps it means to folks something closer to that fact that no added sugar was explicit in the recipe. The word "sugar" itself has many meanings (the well-known problem of word Polysemy). So there are several similar but distinct concepts at involved here and an entity could have the value of true for one of these and false for another.
The Diffbot Knowledge Graph is built by applying computer vision and natural language processing techniques to reading all the pages on the web (which can be in any structure and human language) and extracting it into a structured form, without the element of human annotation in the build pipeline.
[1] https://www.recode.net/2017/7/24/16020330/google-digital-mob...
"Examples of Space Data confidential information and trade secrets include, but are not limited to, accumulation of weather data, launch methods, launch timing, balloon types, altitude regulation, business methods, business models, financial information, technology solutions, and unique knowledge and interpretation of weather data, wherein this knowledge and interpretation is not available in public, which for example includes knowledge of the winds between 60,000 and 100,000 foot altitudes. "