Grounding AI in reality with a little help from Data Commons
research.google
research.google
I bought into TBL’s Semantic Web ideas (and I cover ‘lower case’ semantic web topics in a few of my books). I think it is a shame that publicly accessible world knowledge graphs never really took off, but at least Google’s Data Commons is available for free for non-commercial, educational, and research uses.
I have just been looking at the Data Commons data sets, and I think I will add a fun example to my live Common Lisp eBook (and/or my Racket live eBook).
Huh? What's wrong with Wikidata and the Linked Open Data cloud? These seem quite real to me.
I personally use WikiData and DBPedia, and I have for years.
> [...] Trade-offs of the RAG approach: [...] In addition, the effectiveness of grounding depends on the quality of the generated queries to Data Commons.
Like, two layers of duct tape are better than one.
Seems reasonable and I can believe it helps, just also seems like it doesn't do much to improve confidence in the system as a whole. Particularly since they're basically asking it to find things worth checking, then have it write the checking query, and have it interpret the results. When it's the thing that screwed it up in the first place.
I think if you remember that LLMs are not databases, but they do contain a super lossy-compressed version of (it's training) knowledge, this feels less like a hack. If you ask someone, "who won the World Cup in 2000?", they may say "I think it was X, but let me google it first". That person isn't screwed up, using tools isn't a failure.
If the context is a work setting, or somewhere that is data-centric, it totally makes sense to check it. Like a Chat Bot for a store, or company that is helping someone troubleshoot or research. Anything where it really obvious answers that are easy to learn from volumes of data ("what company makes the corolla?"), probably don't need fact checking as often, but why not have the system check its work?
Meanwhile, programming, writing prose, etc are not things you generally fact-check mid-way, and are things that can be "learned" well from statistical volume. Most programmers can get "pretty good" syntax on first try, and any dedicated syntax tool will get to basically 100%, and the same makes sense for an LLM.
'But what bias did we infer with [LLM knowledgebase] background research prior to formulating a hypothesis, and who measured?'
There are various methods of Inference: Inductive, Deductive, and Abductive
What are the limits of Null Hypothesis methods of scientific inquiry?
Uh... they are? They're not better than a properly specc'd fastener installed at appropriately engineered mounting points but still better than one layer of duct tape, let alone none.
In this sense this project offers a remarkable, albeit implicit, endorsement of that broader open data space, as it comes from a major private sector entity and links with the hot LLM technology of the day.
At high level though, this design seems to violate the "bitter lesson" gospel [1].
> 1) AI researchers have often tried to build knowledge into their agents
Which is at the same refreshing (as there is something very incomplete and self-defeating in the "scaling" hypothesis) and hints at the difficulties of meaningfully integrating very heterogeneous sources and representations of information.
[1] http://www.incompleteideas.net/IncIdeas/BitterLesson.html
Or buy an encyclopedia ...