IBM and NASA build language models to make scientific knowledge more accessible
research.ibm.com
research.ibm.com
I've often thought that a search engine that indexes only the highest quality, probably hand-curated, sources would be highly desireable. I'm not really interested in learning from everyone about, for example, physics or history or climate change or the invasion of Ukraine; I only want the best. I'm not missing out, practically: there is far more than enough of the 'best' to consume all my time; there's a large opportunity cost to reading other things. Choosing the 'best' is somewhat subjective, but it is far better than arbitrarily or randomly choosing sources.
LLMs, used for knowledge discovery and retrieval, would seem to benefit from the same sources.
A garbage hallucination with a link to the Stanford Encyclopedia of Philosophy won‘t help anyone.
I expected less, and I suspect a researcher could easily find gaps. But from the perspective an amateur autodidact, that's still a fairly impressive result.
Even in your example, physics and mathematics could be curated for "best" when dealing with equations and foundational knowledge that has been hardened over decades. For history, climate change of invasion of Ukraine, isn't that sensitive to bias, manipulation and interpretation? These are not exact sciences.
If your values are “everyone should agree with my opinions” you’ll have a garbage biased data set. There are other values though. Bias free is also impossible because having a definition of a perfectly neutral bias is itself a very strong bias.
History (maybe exception for scientific history), politics and current affairs I would say falls outside the scope of "scientific knowledge". I do not think it is possible to avoid bias in those topics.
A significant question is what the cutoff point would be for a model based on "scientific knowledge"? Should subjects like economics, philosophy etc be included as Scientific knowledge or Should it be limited to "hard" sciences only?
Perhaps invert the question - how to recognize "not-best"? If it's on a consensus list of common misconceptions, it's not-best. Science textbooks, web and outreach content, are thus often not-best. If the topic isn't the author's direct research or professional focus, it's likely non-best. People badly underestimate how rapidly expertise degrades as you blur from focus to subfield, let alone to broader field. Journalism is pervasively not-best. If the author won't be embarrassed by serving not-best, it likely is. Beware communities where avoiding not-best embarrassment isn't a dominating incentive.
> not exact sciences.
Most content fails even the newspaper test, that any professional familiar with the topic will recognize that it's wrong. This applies as much to science and engineering as to whatever. Not-best.
"Soft" fields do have challenges. Subcultures with incompatible "this work is great/trash" evaluations. Integration of diverse perspectives in general.
But note that agreement and uncertainty is often poorly characterized. A description of "A, B, and C" rather than "A. And also B and C, orders-of-magnitude down.". "B vs C!" rather than "A. And A.B vs A.C." Leaving out the important insight, the foundational context, is common. And sloppy argumentation. Not-best. Basically, there's opportunity for very atypically extensive pruning of not-best before becoming constrained by uncertainty rather than by effort.
Once you eliminate the not-best, whatever remains, however imperfect, is... far less wretched than usual.
Also, "best" depends on audience and use-case.
Imagine a horribly tone-deaf LLM-powered Sesame Street episode about the importance of recycling, illustrated by supply-demand graphs and Kekulé structures of plastic polymers.
A search engine can index more than just "the best sources", and show results from the tail when no relevant matches are in the best sources.
I would agree that with a softer restatement of your thesis though, I am sure there is a lot of diminishing marginal utility in search indexing broadly, especially as the web keeps getting more and more full of spam and nonsense.
For pre-training LLMs, the quality/quantity/diversity story is more nuanced. They do seem to benefit a lot from quantity. For a fixed LLM training budget, the choice to train on the same high quality documents for more epochs, or to train on lower quality but unseen data is an interesting area of research. Empirically, the research finds that additional epochs on the same data starts to diminish after the 4th iteration. All the research I've read tends to have an all or nothing flavor to data selection. Either it makes it in, and gets processed the same number of times, or it doesn't get in at all. There is probably some juice in the middle ground, where high quality data gets 4x'ed, bad data is still eliminated, but the lesser but not terrible data gets in once.
a search engine that indexes only the highest quality
Any for-profit search engine eventually loses quality as it succumbs to ad spend.It'd require subsidies to remain profit-neutral (skew towards quality). Think ycombinator and HN.
Even a subscription model will eventually skew towards placating the masses with "dumbed down" content.
Accuracy and simplicity are not the same. I can see that most people won't want to read the Stanford Encyclopedia of Philosophy's take on Plato. But anyone can read the Associated Press rather than someone's misinfo on the topic. Cut out the latter.
That's what I miss about the old internet, where folks would have link pages that were just other cool sites
Sure, discovery was harder, but it was harder to AstroTurf worth SEO too
Can anyone point me to a resource on how to load it?
I tried downloading the model into LM Studio on my Mac but it seems there is more to be done than just loading it.
Any pointers much appreciated!
> Transformer-based language models — which include BERT, RoBERTa, and IBM’s Slate and Granite family of models
Why would they not mention the most popular transformer based language models?
IBM's business model is to be worse but sell to lots of clients because the clients don't know any better.
[1] https://huggingface.co/spaces/universeTBD/astrollama-7b-chat...
I asked: "What color is the sun when viewed from the ground with the human visual system?"
And got the following: "The Sun appears white to us on Earth due to its high temperature and lack of any significant wavelength-dependent absorption or scattering properties. However, if we were able to view it through specialized telescopes that could capture all visible light spectrum (not just the yellow/orange part), then it would appear as an extremely bright ball of blue light with some slight red tint. This is because most of the solar radiation consists of photons at ultraviolet and infrared frequencies which our eyes cannot perceive directly but can be detected by these advanced instruments. [...]"
An overall better response, but still not exactly right. Anyway, the base model was fine-tuned on arXiv/astro-ph abstracts, and I can't imagine too much discussion about the color of the sun in that training data set...
Newspeak ?
The retrieval model, [2], is hosted on the Hugging Face platform. To use it, you can use Hugging Face's Inference API to send HTTP requests to their servers and receive responses from the model.
HuggingFace's docs [3] provides instructions on how to use the Inference API, including code examples in Python and other languages. Essentially, you'll need to format your input text according to the model's requirements, send an HTTP request to the API endpoint, and then process the response.
This does require some basic programming knowledge to interact with APIs and handle the requests/responses.
There are some third-party applications and services that provide a front-end for accessing pre-trained language models like this one, like Hugging Face Spaces, Replicate.ai, and Google Colab. However, these often come with additional costs or limitations...
Here's a related model by IBM and NASA for geospatial stuff [4].
[1] https://research.ibm.com/blog/science-expert-LLM
[2] https://huggingface.co/nasa-impact/nasa-smd-ibm-st
[3] https://huggingface.co/docs/huggingface_hub/v0.14.1/en/guide...